Atomic-Germ / Aquila-mini-35B-A3B-NPU2

huggingface.co
Total runs: 64
24-hour runs: -15
7-day runs: -32
30-day runs: 36
Model's Last Updated: August 26 2026
image-text-to-text

Introduction of Aquila-mini-35B-A3B-NPU2

Model Details of Aquila-mini-35B-A3B-NPU2

Aquilia-mini-NPU2

FastFlowLM Q4NX conversion of XYZAILab/XYZ-Aquila-mini for AMD XDNA NPU inference.

This repository contains a quantized Q4NX port of the model, compiled for the FastFlowLM (FLM) runtime. It is not a GGUF file.

Item Value
Source model XYZAILab/XYZ-Aquila-mini
Source GGUF XYZAILab_XYZ-Aquila-mini-Q4_1.gguf
Weights model.q4nx (21.64 GB)
Modality language
FLM version 1.0.0
Converted 2026-08-11
Source repository

Metadata from the upstream Hugging Face repository:

Item Value
License apache-2.0
Base model ['Qwen/Qwen3.6-35B-A3B']
Library transformers
Model type qwen3_5_moe
Pipeline image-text-to-text
Downloads 2,091
Repo revision 0cad6285baf6f37adf2c4e9696372c0140078fe0
About

Q4NX for FastFlowLM (AMD Ryzen AI XDNA2) quant of https://huggingface.co/XYZAILab/XYZ-Aquila-mini

What is Q4NX?

Q4NX is FastFlowLM's native packed-quantization format - a rearranged Q4_1 layout tuned for the NPU matrix engine's tile sizes and memory access patterns. It is not a GGUF file and it does not run on llama.cpp or Ollama; it is meant exclusively for the FastFlowLM engine on AMD Ryzen AI NPUs.

Requirements
  • FastFlowLM >= 0.9.45 ( flm CLI)
  • AMD Ryzen AI processor with XDNA2 (NPU2) - Strix Point / Ryzen AI 300 series or later
  • XRT NPU stack installed
  • 32 GB of unified system memory (Q4NX weights + activations + KV cache)

Files
File Purpose
model.q4nx Quantized Q4NX weights
config.json FastFlowLM model configuration
tokenizer.json Tokenizer
tokenizer_config.json Special tokens and chat template
chat_template.jinja Chat template (optional)
flm-add.py Installer script - registers this model with FastFlowLM
Install and run

This repository works with flm-add.py , a small installer that copies the model into the FastFlowLM user directory and registers the tag darwin-opus:36b . It never modifies the system FastFlowLM install.

python3 ./flm-add.py Atomic-Germ/Aquila-mini-35B-A3B-NPU2 --family qwen3.6-moe --tag aquila-mini-moe:35b-a3b
FLM_XCLBIN_PATH="$HOME/.config/flm" FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json" flm run aquila-mini-moe:35b-a3b
Kernels

FastFlowLM's NPU kernels (xclbins) are closed source and are not shipped in this repository. flm-add.py links the kernels of the official qwen3.6-moe:35b-a3b model ( Qwen3.6-35B-A3B-NPU2 ), because this model shares the same engine family ( qwen3.6-moe ) and architecture.

Model
  • Registry tag: aquila-mini-moe:35b-a3b
  • Engine family: qwen3.6-moe
  • Kernel source: Qwen3.6-35B-A3B-NPU2
Original model card

See the upstream model card for training details, benchmarks, and upstream usage. This repository only contains the Q4NX conversion for FastFlowLM.

Runs of Atomic-Germ Aquila-mini-35B-A3B-NPU2 on huggingface.co

64
Total runs
-15
24-hour runs
-21
3-day runs
-32
7-day runs
36
30-day runs

More Information About Aquila-mini-35B-A3B-NPU2 huggingface.co Model

More Aquila-mini-35B-A3B-NPU2 license Visit here:

https://choosealicense.com/licenses/apache-2.0

Aquila-mini-35B-A3B-NPU2 huggingface.co

Aquila-mini-35B-A3B-NPU2 huggingface.co is an AI model on huggingface.co that provides Aquila-mini-35B-A3B-NPU2's model effect (), which can be used instantly with this Atomic-Germ Aquila-mini-35B-A3B-NPU2 model. huggingface.co supports a free trial of the Aquila-mini-35B-A3B-NPU2 model, and also provides paid use of the Aquila-mini-35B-A3B-NPU2. Support call Aquila-mini-35B-A3B-NPU2 model through api, including Node.js, Python, http.

Aquila-mini-35B-A3B-NPU2 huggingface.co Url

https://huggingface.co/Atomic-Germ/Aquila-mini-35B-A3B-NPU2

Atomic-Germ Aquila-mini-35B-A3B-NPU2 online free

Aquila-mini-35B-A3B-NPU2 huggingface.co is an online trial and call api platform, which integrates Aquila-mini-35B-A3B-NPU2's modeling effects, including api services, and provides a free online trial of Aquila-mini-35B-A3B-NPU2, you can try Aquila-mini-35B-A3B-NPU2 online for free by clicking the link below.

Atomic-Germ Aquila-mini-35B-A3B-NPU2 online free url in huggingface.co:

https://huggingface.co/Atomic-Germ/Aquila-mini-35B-A3B-NPU2

Aquila-mini-35B-A3B-NPU2 install

Aquila-mini-35B-A3B-NPU2 is an open source model from GitHub that offers a free installation service, and any user can find Aquila-mini-35B-A3B-NPU2 on GitHub to install. At the same time, huggingface.co provides the effect of Aquila-mini-35B-A3B-NPU2 install, users can directly use Aquila-mini-35B-A3B-NPU2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Aquila-mini-35B-A3B-NPU2 install url in huggingface.co:

https://huggingface.co/Atomic-Germ/Aquila-mini-35B-A3B-NPU2

Url of Aquila-mini-35B-A3B-NPU2

Aquila-mini-35B-A3B-NPU2 huggingface.co Url

Provider of Aquila-mini-35B-A3B-NPU2 huggingface.co

Atomic-Germ
ORGANIZATIONS

Other API from Atomic-Germ