This repository hosts the optimized versions of
microsoft/Phi-3-small-128k-instruct
to accelerate inference with DirectML and ONNX Runtime.
The Phi-3-Small-128K-Instruct is a state-of-the-art, lightweight open model developed by Microsoft, featuring 7B parameters.
Key Features:
Parameter Count: 7B
Tokenizer: Utilizes the tiktoken tokenizer for improved multilingual tokenization, with a vocabulary size of 100,352 tokens.
Context Length: Default context length of 128k tokens.
Attention Mechanism:
Implements grouped-query attention to minimize KV cache footprint, with 4 queries sharing 1 key.
Uses alternative layers of dense attention and a novel blocksparse attention to further optimize on KV cache savings while maintaining long context retrieval performance.
Multilingual Capability: Includes an additional 10% of multilingual data to enhance its performance across different languages.
ONNX Models
Here are some of the optimized configurations we have added:
ONNX model for int4 DirectML:
ONNX model for AMD, Intel, and NVIDIA GPUs on Windows, quantized to int4 using AWQ.
ONNX model for int4 CPU and Mobile:
ONNX model for CPU and mobile using int4 quantization via RTN. There are two versions uploaded to balance latency vs. accuracy. Acc=1 is targeted at improved accuracy, while Acc=4 is for improved performance. For mobile devices, we recommend using the model with acc-level-4.
Usage
Installation and Setup
To use the Phi-3-small-128k-instruct ONNX model on Windows with DirectML, follow these steps:
This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow
Microsoft’s Trademark & Brand Guidelines
. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party’s policies.
Runs of EmbeddedLLM Phi-3-small-128k-instruct-onnx on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Phi-3-small-128k-instruct-onnx huggingface.co Model
More Phi-3-small-128k-instruct-onnx license Visit here:
Phi-3-small-128k-instruct-onnx huggingface.co is an AI model on huggingface.co that provides Phi-3-small-128k-instruct-onnx's model effect (), which can be used instantly with this EmbeddedLLM Phi-3-small-128k-instruct-onnx model. huggingface.co supports a free trial of the Phi-3-small-128k-instruct-onnx model, and also provides paid use of the Phi-3-small-128k-instruct-onnx. Support call Phi-3-small-128k-instruct-onnx model through api, including Node.js, Python, http.
Phi-3-small-128k-instruct-onnx huggingface.co is an online trial and call api platform, which integrates Phi-3-small-128k-instruct-onnx's modeling effects, including api services, and provides a free online trial of Phi-3-small-128k-instruct-onnx, you can try Phi-3-small-128k-instruct-onnx online for free by clicking the link below.
EmbeddedLLM Phi-3-small-128k-instruct-onnx online free url in huggingface.co:
Phi-3-small-128k-instruct-onnx is an open source model from GitHub that offers a free installation service, and any user can find Phi-3-small-128k-instruct-onnx on GitHub to install. At the same time, huggingface.co provides the effect of Phi-3-small-128k-instruct-onnx install, users can directly use Phi-3-small-128k-instruct-onnx installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Phi-3-small-128k-instruct-onnx install url in huggingface.co: