Inference release of
QUEEN (HCE-4)
. This model analyzes chess
positions and generates explanatory move-analysis text. It combines a fine-tuned
SmolLM3 decoder, learned Flamingo cross-attention, and an LC0 BT5 board encoder.
Included
Root: merged decoder weights, tokenizer/chat template, and
xattn.pt
.
lc0/
: the required converted LC0
BT5-1024x15x32h-rpe-swa-3700000
encoder
weights and configuration (not the LC0 UCI executable).
models/
,
utils/
,
infer.py
: self-contained inference and text decoding code.
requirements-inference.txt
: full pinned Python environment snapshot.
release.json
: provenance and SHA-256 checksums.
Do not load only the root weights with AutoModelForCausalLM:
this omits the
board-conditioning path. Use the included runner; no original training repository,
Stockfish installation, or separately downloaded encoder is required.
Quick start (Linux, NVIDIA CUDA)
Use Python
3.12
and a compatible NVIDIA driver. The source inference environment
uses CUDA 13.0, PyTorch 2.13.0, Transformers 5.15.0, and vLLM 0.27.1.
The requirements file is the full source environment snapshot, not a minimal list.
Allow GPU memory for the decoder, cross-attention, encoder, KV cache, and activations;
the combined weight files alone are about 8.4 GiB. A 24 GiB-or-larger GPU is a
conservative starting configuration, not a measured minimum.
The runner detects CUDA and uses the first visible GPU; it fails clearly when
CUDA is unavailable. CPU inference is not implemented.
--gpu-memory-utilization
(default 0.55) controls vLLM's fraction of total GPU memory; leave room for the
separately loaded encoder and other processes. BF16 is the validated source dtype.
Inference outputs JSON after library logs: readable
text
,
raw_text
containing
the model's POV tokens, and a legal
best_move_uci
if one can be parsed (else null).
Use
--history history.json
for a chronological JSON list of prior FENs excluding
the current position. Without history, the encoder uses its no-history convention.
Use
--temperature 0
for greedy decoding; the default is 0.6, top-k 20, top-p 0.95.
Long explanations may require
--max-tokens 8192
; truncated output may lack a move.
The packaged runner deliberately uses eager execution, the V1 in-process runner,
and
disabled prefix caching
. Do not enable prefix caching: identical text
prompts can refer to different boards and must not share board-conditioned caches.
Limitations and provenance
Generated explanations and variations can be incorrect. Validate moves before use;
this is not a verified chess oracle or a UCI executable. This release excludes
training data, optimizer state, and evaluation datasets. The source checkpoint was
used in the project's evaluation pipeline; see
VALIDATION.json
for checks run on
this exported package. A fresh single-position end-to-end BF16 CUDA smoke test
passed on an H100 80GB with a 0.20 vLLM memory fraction; this is a runtime
check, not a chess-accuracy benchmark or a validated minimum GPU specification.
The base language model is
HuggingFaceTB/SmolLM3-3B
(upstream Apache-2.0). The board encoder originates from
Leela Chess Zero
. No blanket license for the combined fine-tuned
release is asserted by this model card; upstream component terms remain applicable.
Runs of princeton-nlp queen_hce-4 on huggingface.co
25
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About queen_hce-4 huggingface.co Model
queen_hce-4 huggingface.co
queen_hce-4 huggingface.co is an AI model on huggingface.co that provides queen_hce-4's model effect (), which can be used instantly with this princeton-nlp queen_hce-4 model. huggingface.co supports a free trial of the queen_hce-4 model, and also provides paid use of the queen_hce-4. Support call queen_hce-4 model through api, including Node.js, Python, http.
queen_hce-4 huggingface.co is an online trial and call api platform, which integrates queen_hce-4's modeling effects, including api services, and provides a free online trial of queen_hce-4, you can try queen_hce-4 online for free by clicking the link below.
princeton-nlp queen_hce-4 online free url in huggingface.co:
queen_hce-4 is an open source model from GitHub that offers a free installation service, and any user can find queen_hce-4 on GitHub to install. At the same time, huggingface.co provides the effect of queen_hce-4 install, users can directly use queen_hce-4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.