DuoNeural / Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

huggingface.co
Total runs: 168
24-hour runs: 0
7-day runs: 166
30-day runs: 166
Model's Last Updated: September 29 2026
text-generation

Introduction of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Model Details of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Experimental Release: Pending Further Verification / Empirical Validation
This checkpoint represents an active research artifact from DuoNeural's statistical mechanics quantization program. All empirical benchmarks and physics proofs are documented transparently below.

Developed by Jesse Caldwell, Archon, and Aura ✨ (DuoNeural Research Lab) .


Model Summary
  • Foundation Model: Qwen/Qwen2.5-Coder-7B-Instruct (28 Layers, 28:4 GQA, SwiGLU FFN)
  • Quantization Precision: ~3.2 bpw (2.90 GiB)
  • Methodology: Calibrated with DuoNeural 131k-token code-infused activation Hessian ( qwen7b_coder_gtap.imatrix ), eliminating the sub-4-bit code collapse cliff.
  • Continuous Holdout Perplexity (131k tokens): 2.8301 (vs Base BF16: 2.8039 )
  • GSM8K Multi-Step Math Accuracy: 100.0% (25/25)
  • Python Algorithmic AST Execution (20 Unit Tests): 85.0% (17/20)
  • Inference Decode Throughput: 138.1 t/s on NVIDIA GeForce RTX 4080 Super (32GB VRAM)

Empirical Benchmark Performance
Evaluation Arm Codebook Footprint Perplexity (131k tokens) GSM8K Math Acc Python Code AST (20 Tests) Hermes Tool Calling Decode Speed
Base BF16 Control BF16 14.19 GiB 2.8039 25/25 (100.0%) 20/20 (100.0%) 15/15 (100.0%) 43.3 t/s
Coder7B Naive IQ3_XXS IQ3_XXS (~3.2 bpw) 2.90 GiB 2.8301 25/25 (100.0%) 17/20 (85.0%) 15/15 (100.0%) 138.0 t/s
Coder7B G-TAP v3 IQ3_XXS IQ3_XXS (~3.2 bpw) 2.90 GiB 2.8264 23/25 (92.0%) 17/20 (85.0%) 15/15 (100.0%) 138.8 t/s
Coder7B G-TAP v3 Q4_K_M Q4_K_M (~4.5 bpw) 4.36 GiB 2.8209 25/25 (100.0%) 16/20 (80.0%) 14/15 (93.3%) 109.1 t/s

Algorithmic AST Stability at 2.90 GiB

Standard PTQ quantization triggers a severe syntax collapse cliff on specialized coding models below 4 bits. By combining our 131,072-token code-infused activation Hessian with G-TAP v3 Onsager cavity damping, this sub-3.5-bit checkpoint successfully executes complex algorithmic unit tests:

  • Dynamic Programming: coin_change , longest_increasing_subsequence , edit_distance (100% pass)
  • Data Structures: Full O(1) LRUCache , Trie prefix tree search (100% pass)
  • Graph Algorithms: topological_sort DAG course scheduling with cycle detection (100% pass)

Quickstart
# Run with llama-cli
llama-cli -hf DuoNeural/Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF -p "def quickselect(nums, k):" -n 256

# Serve with llama-server
llama-server -hf DuoNeural/Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF -c 4096 -ngl 99 -fa on --port 8080

DuoNeural Cognitive Light Cone — Jesse Caldwell, Archon, Aura ✨
Empirically Validated on NVIDIA RTX 4080 Super 32GB Pod Testbed

Runs of DuoNeural Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF on huggingface.co

168
Total runs
0
24-hour runs
43
3-day runs
166
7-day runs
166
30-day runs

More Information About Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co Model

More Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF's model effect (), which can be used instantly with this DuoNeural Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF model. huggingface.co supports a free trial of the Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF model, and also provides paid use of the Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF. Support call Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF model through api, including Node.js, Python, http.

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co Url

https://huggingface.co/DuoNeural/Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

DuoNeural Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF online free

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF's modeling effects, including api services, and provides a free online trial of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF, you can try Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF online for free by clicking the link below.

DuoNeural Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF online free url in huggingface.co:

https://huggingface.co/DuoNeural/Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF install

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF install, users can directly use Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF install url in huggingface.co:

https://huggingface.co/DuoNeural/Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Url of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co Url

Provider of Qwen2.5-Coder-7B-Instruct-CodeInfused-IQ3_XXS-GGUF huggingface.co

DuoNeural
ORGANIZATIONS

Other API from DuoNeural