This model is an ONNX-optimized version of
microsoft/Phi-3-mini-4k-instruct (June 2024)
, designed to provide accelerated inference on a variety of hardware using ONNX Runtime(CPU and DirectML).
DirectML is a high-performance, hardware-accelerated DirectX 12 library for machine learning, providing GPU acceleration for a wide range of supported hardware and drivers, including AMD, Intel, NVIDIA, and Qualcomm GPUs.
ONNX Models
Here are some of the optimized configurations we have added:
ONNX model for int4 DirectML:
ONNX model for AMD, Intel, and NVIDIA GPUs on Windows, quantized to int4 using AWQ.
Hardware Requirements
Minimum Configuration:
Windows:
DirectX 12-capable GPU (AMD/Nvidia)
CPU:
x86_64 / ARM64
Tested Configurations:
GPU:
AMD Ryzen 8000 Series iGPU (DirectML)
CPU:
AMD Ryzen CPU
Model Description
Developed by:
Microsoft
Model type:
ONNX
Language(s) (NLP):
Python, C, C++
License:
Apache License Version 2.0
Model Description:
This model is a conversion of the Phi-3-mini-4k-instruct-062024 for ONNX Runtime inference, optimized for DirectML.
Performance Metrics
DirectML
We measured the performance of DirectML on AMD Ryzen 9 7940HS /w Radeon 78
Prompt Length
Generation Length
Average Throughput (tps)
128
128
-
128
256
-
128
512
-
128
1024
-
256
128
-
256
256
-
256
512
-
256
1024
-
512
128
-
512
256
-
512
512
-
512
1024
-
1024
128
-
1024
256
-
1024
512
-
1024
1024
-
Runs of EmbeddedLLM Phi-3-mini-4k-instruct-062024-int4-directml on huggingface.co
3
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Phi-3-mini-4k-instruct-062024-int4-directml huggingface.co Model
More Phi-3-mini-4k-instruct-062024-int4-directml license Visit here:
Phi-3-mini-4k-instruct-062024-int4-directml huggingface.co is an AI model on huggingface.co that provides Phi-3-mini-4k-instruct-062024-int4-directml's model effect (), which can be used instantly with this EmbeddedLLM Phi-3-mini-4k-instruct-062024-int4-directml model. huggingface.co supports a free trial of the Phi-3-mini-4k-instruct-062024-int4-directml model, and also provides paid use of the Phi-3-mini-4k-instruct-062024-int4-directml. Support call Phi-3-mini-4k-instruct-062024-int4-directml model through api, including Node.js, Python, http.
Phi-3-mini-4k-instruct-062024-int4-directml huggingface.co is an online trial and call api platform, which integrates Phi-3-mini-4k-instruct-062024-int4-directml's modeling effects, including api services, and provides a free online trial of Phi-3-mini-4k-instruct-062024-int4-directml, you can try Phi-3-mini-4k-instruct-062024-int4-directml online for free by clicking the link below.
EmbeddedLLM Phi-3-mini-4k-instruct-062024-int4-directml online free url in huggingface.co:
Phi-3-mini-4k-instruct-062024-int4-directml is an open source model from GitHub that offers a free installation service, and any user can find Phi-3-mini-4k-instruct-062024-int4-directml on GitHub to install. At the same time, huggingface.co provides the effect of Phi-3-mini-4k-instruct-062024-int4-directml install, users can directly use Phi-3-mini-4k-instruct-062024-int4-directml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Phi-3-mini-4k-instruct-062024-int4-directml install url in huggingface.co: