princeton-nlp / queen_hce-4

huggingface.co
Total runs: 25
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: Octubre 03 2026

Introduction of queen_hce-4

Model Details of queen_hce-4

QUEEN (HCE-4)

Inference release of QUEEN (HCE-4) . This model analyzes chess positions and generates explanatory move-analysis text. It combines a fine-tuned SmolLM3 decoder, learned Flamingo cross-attention, and an LC0 BT5 board encoder.

Included
  • Root: merged decoder weights, tokenizer/chat template, and xattn.pt .
  • lc0/ : the required converted LC0 BT5-1024x15x32h-rpe-swa-3700000 encoder weights and configuration (not the LC0 UCI executable).
  • models/ , utils/ , infer.py : self-contained inference and text decoding code.
  • requirements-inference.txt : full pinned Python environment snapshot.
  • release.json : provenance and SHA-256 checksums.

Do not load only the root weights with AutoModelForCausalLM: this omits the board-conditioning path. Use the included runner; no original training repository, Stockfish installation, or separately downloaded encoder is required.

Quick start (Linux, NVIDIA CUDA)

Use Python 3.12 and a compatible NVIDIA driver. The source inference environment uses CUDA 13.0, PyTorch 2.13.0, Transformers 5.15.0, and vLLM 0.27.1. The requirements file is the full source environment snapshot, not a minimal list. Allow GPU memory for the decoder, cross-attention, encoder, KV cache, and activations; the combined weight files alone are about 8.4 GiB. A 24 GiB-or-larger GPU is a conservative starting configuration, not a measured minimum.

python -m pip install huggingface_hub
hf download princeton-nlp/queen_hce-4 --local-dir queen_hce-4
cd queen_hce-4
python3.12 -m venv .venv
.venv/bin/python -m pip install -r requirements-inference.txt
CUDA_VISIBLE_DEVICES=0 .venv/bin/python infer.py \
  --fen 'rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1' \
  --max-tokens 2048

The runner detects CUDA and uses the first visible GPU; it fails clearly when CUDA is unavailable. CPU inference is not implemented. --gpu-memory-utilization (default 0.55) controls vLLM's fraction of total GPU memory; leave room for the separately loaded encoder and other processes. BF16 is the validated source dtype.

Inference outputs JSON after library logs: readable text , raw_text containing the model's POV tokens, and a legal best_move_uci if one can be parsed (else null). Use --history history.json for a chronological JSON list of prior FENs excluding the current position. Without history, the encoder uses its no-history convention. Use --temperature 0 for greedy decoding; the default is 0.6, top-k 20, top-p 0.95. Long explanations may require --max-tokens 8192 ; truncated output may lack a move.

The packaged runner deliberately uses eager execution, the V1 in-process runner, and disabled prefix caching . Do not enable prefix caching: identical text prompts can refer to different boards and must not share board-conditioned caches.

Limitations and provenance

Generated explanations and variations can be incorrect. Validate moves before use; this is not a verified chess oracle or a UCI executable. This release excludes training data, optimizer state, and evaluation datasets. The source checkpoint was used in the project's evaluation pipeline; see VALIDATION.json for checks run on this exported package. A fresh single-position end-to-end BF16 CUDA smoke test passed on an H100 80GB with a 0.20 vLLM memory fraction; this is a runtime check, not a chess-accuracy benchmark or a validated minimum GPU specification.

The base language model is HuggingFaceTB/SmolLM3-3B (upstream Apache-2.0). The board encoder originates from Leela Chess Zero . No blanket license for the combined fine-tuned release is asserted by this model card; upstream component terms remain applicable.

Runs of princeton-nlp queen_hce-4 on huggingface.co

25
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About queen_hce-4 huggingface.co Model

queen_hce-4 huggingface.co

queen_hce-4 huggingface.co is an AI model on huggingface.co that provides queen_hce-4's model effect (), which can be used instantly with this princeton-nlp queen_hce-4 model. huggingface.co supports a free trial of the queen_hce-4 model, and also provides paid use of the queen_hce-4. Support call queen_hce-4 model through api, including Node.js, Python, http.

princeton-nlp queen_hce-4 online free

queen_hce-4 huggingface.co is an online trial and call api platform, which integrates queen_hce-4's modeling effects, including api services, and provides a free online trial of queen_hce-4, you can try queen_hce-4 online for free by clicking the link below.

princeton-nlp queen_hce-4 online free url in huggingface.co:

https://huggingface.co/princeton-nlp/queen_hce-4

queen_hce-4 install

queen_hce-4 is an open source model from GitHub that offers a free installation service, and any user can find queen_hce-4 on GitHub to install. At the same time, huggingface.co provides the effect of queen_hce-4 install, users can directly use queen_hce-4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

queen_hce-4 install url in huggingface.co:

https://huggingface.co/princeton-nlp/queen_hce-4

Url of queen_hce-4

Provider of queen_hce-4 huggingface.co

princeton-nlp
ORGANIZATIONS

Other API from princeton-nlp