hy3-gguf — GGUF weights for Tencent Hy3 (tencent/Hy3)
GGUF-format weights for
tencent/Hy3
(
HYV3ForCausalLM
,
model_type: hy_v3
), a
295B-parameter / 21B-active-parameter Mixture-of-Experts model from
Tencent's Hunyuan ("Hy") team
(
tencent/Hy3
).
These files were produced by the
hy3
converter (
hy3-convert
) and are meant to be run with the
hy3
inference
engine
, a from-scratch C/Metal/CUDA implementation.
⚠️ This GGUF does NOT work with llama.cpp
Despite the
.gguf
extension, these files are
only usable by the
hy3
engine
.
llama.cpp
,
ollama
,
LM Studio
,
text-generation-webui
,
koboldcpp
, and any other
llama.cpp-based tool
cannot load these files
. Three independent reasons:
Unknown architecture.
The metadata declares
general.architecture = "hy_v3"
. llama.cpp only knows
hunyuan-moe
,
hunyuan-dense
,
hunyuan_vl
— loading aborts with
unknown model architecture: 'hy_v3'
.
Custom metadata keys.
All hyperparameters use the
hy_v3.*
prefix
(
hy_v3.block_count
,
hy_v3.expert_count
, …), which llama.cpp does not
look up.
Non-fused expert tensors.
Experts are stored
one tensor per expert
(
blk.N.ffn_gate_exps.0.gate_proj.weight
,
…1…
, … — 46080 tensors),
whereas llama.cpp expects experts fused into a single stacked 3D tensor per
layer. This is a fundamentally different on-disk layout.
This is a
custom GGUF
readable only by the
hy3
loader. Do not open
issues against llama.cpp for these files.
git clone https://github.com/yuhai-china/hy3
cd hy3
make # macOS builds the Metal backend automatically# download a GGUF from this repo, then:
./run_metal.sh -m /path/to/hy3_q4k_mixed.gguf -p "The capital of France is" -experts 8
Testing scope:
the
hy3
engine's performance work and benchmarks were
developed and verified
only on macOS / Apple Silicon (Metal backend)
,
measured on an M2 Ultra (~20–27 tok/s decode depending on
-experts
). The
CPU and CUDA backends exist in the source but were not exercised as part of
that work — treat them as untested.
Files / quantization
The mixed-precision GGUF follows this scheme (see
hy3_convert.c
):
Tensor group
Type
Routed experts (
ffn_{gate,up,down}_exps
) — the bulk of the model
sigmoid(router_logits)
; top-8 by
sigmoid + expert_bias
, combined using
unbiased
sigmoid weights, renormalized to sum 1, scaled by
router_scaling_factor = 2.826
The engine supports a runtime
top-k experts
override (
-experts 1..8
) to
trade quality for speed. On a small 13-question code/reasoning eval (greedy,
no-think):
experts=8 → 10/13
,
experts=4 → 7/13
. Default is 8.
Chat template
Hy3 is instruction-tuned and expects the Hunyuan V3 chat format (the
hy3
engine applies it automatically; use
--raw
to bypass). Single user turn,
no-think:
Generation stops on
<|hy_eos:opensource|>
(120025),
<|hy_endofsentence|>
(120001), or
<|hy_EOT|>
(120008).
License & attribution
Weights derive from
tencent/Hy3
; refer
to the upstream repository for the governing model license. This is an
unofficial community conversion, not affiliated with or endorsed by Tencent.
Runs of cloudyu hy3-gguf on huggingface.co
915
Total runs
0
24-hour runs
13
3-day runs
151
7-day runs
512
30-day runs
More Information About hy3-gguf huggingface.co Model
hy3-gguf huggingface.co is an AI model on huggingface.co that provides hy3-gguf's model effect (), which can be used instantly with this cloudyu hy3-gguf model. huggingface.co supports a free trial of the hy3-gguf model, and also provides paid use of the hy3-gguf. Support call hy3-gguf model through api, including Node.js, Python, http.
hy3-gguf huggingface.co is an online trial and call api platform, which integrates hy3-gguf's modeling effects, including api services, and provides a free online trial of hy3-gguf, you can try hy3-gguf online for free by clicking the link below.
cloudyu hy3-gguf online free url in huggingface.co:
hy3-gguf is an open source model from GitHub that offers a free installation service, and any user can find hy3-gguf on GitHub to install. At the same time, huggingface.co provides the effect of hy3-gguf install, users can directly use hy3-gguf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.