These are research checkpoints from
theta
, a training run of Parallax. Training was stopped at step 6387 (156.2B tokens), before the planned end of the run; the final checkpoint is marked in the table below.
Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several
countries. The hosts have no direct connections to each other and synchronize over the public internet.
Base model.
It is pretrained on web text only. It is not instruction-tuned, chat-tuned or
safety-tuned, and it will continue text rather than follow instructions.
Mid-training.
Every export, the final one included, is a snapshot taken well before the planned ~973B tokens. Quality between exports is not monotonic.
Research artifact.
It is published so the training can be followed and inspected. It is
not intended for production use.
4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token
Expert weights
ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight
Width
d_model 1152; ReLU² expert activation
Tokenizer
Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings
Training context
4096 tokens
Training
Data: FineWeb-Edu, planned as a single pass over about 973B tokens. The published
validation number (
val_mix
) is measured on held-out FineWeb-Edu files that are not in the training stream.
Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Each
routed expert has one owner GPU that folds low-rank updates from the nodes that trained it into its
full-precision master and publishes new ternary versions. The tech report describes the details.
The learning rate was lowered by hand during the run, to 0.5x and later 0.25x of the planned peak,
so later exports do not follow the planned schedule exactly.
Exports
One folder per export under
exports/
, named by training tokens (rounded down to whole billions) and
fleet step. Exports were taken about once an hour; training has stopped and no new exports will be added. Published exports are
never modified or removed.
val_mix
: mean cross-entropy (nats per token, lower is better) on the held-out FineWeb-Edu validation set.
0-shot macro
: mean over 11 tasks (ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, LAMBADA,
OpenBookQA, PIQA, SciQ, SIQA, WinoGrande).
5-shot macro
: the same tasks without LAMBADA (10 tasks). Per task,
acc_norm is used for ARC, HellaSwag, OpenBookQA, PIQA and SciQ, and acc for the rest. Values are percentages.
The scores come from the project's own scorer, which runs the native ternary experts. They track progress
within this run. Compare them with numbers from other evaluation harnesses with care.
A
-
means the export has not been scored yet. The table fills in as scores arrive.
Latest export
exports/156B-tokens_step6387
: 156.2B training tokens, step 6387.
hf download chutesai/parallax-8b-theta --include "exports/156B-tokens_step6387/*" --local-dir parallax-8b-theta
cd parallax-8b-theta/exports/156B-tokens_step6387
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS # every file of the original export, byte for byte
Files and format
The files are in Parallax's
native compact export format
(inference only, no optimizer state). They
are byte-identical to the export the training system produced:
trunk (non-expert) weights in bf16, with their own manifest
packed_experts.tar
the 4096 routed experts (
packed_experts/*.t24p
, packed ternary codes and scales)
SHA256SUMS
sha256 of every file of the original export
export_info.json
step, tokens, time, sizes and digests of this export
The only change from the original export is packaging. The 4096 expert files are stored in one
uncompressed tar to keep the repository's file count manageable. Extract it and check
SHA256SUMS
as shown above. Every upload was checked against the training system's own digests before and
after it was published.
Running it.
Standard
transformers
cannot load this format. A llama.cpp-based runtime for Parallax
and GGUF conversions of these checkpoints are in preparation (link to follow).
Tech report
The Parallax tech report:
https://parallax.chutes.ai/tech-report.pdf
. It is AI-generated from the project's
measurements and logs, and it is a living document that changes as the run progresses. It describes
the previous run (eta). Theta uses the same architecture with changes to the synchronization recipe.
Limitations
This is an early base model trained on educational web text. It can produce incorrect, biased or
nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.
License
MIT.
Table updated 2026-09-28 15:14 UTC.
Runs of chutesai parallax-8b-theta on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About parallax-8b-theta huggingface.co Model
parallax-8b-theta huggingface.co is an AI model on huggingface.co that provides parallax-8b-theta's model effect (), which can be used instantly with this chutesai parallax-8b-theta model. huggingface.co supports a free trial of the parallax-8b-theta model, and also provides paid use of the parallax-8b-theta. Support call parallax-8b-theta model through api, including Node.js, Python, http.
parallax-8b-theta huggingface.co is an online trial and call api platform, which integrates parallax-8b-theta's modeling effects, including api services, and provides a free online trial of parallax-8b-theta, you can try parallax-8b-theta online for free by clicking the link below.
chutesai parallax-8b-theta online free url in huggingface.co:
parallax-8b-theta is an open source model from GitHub that offers a free installation service, and any user can find parallax-8b-theta on GitHub to install. At the same time, huggingface.co provides the effect of parallax-8b-theta install, users can directly use parallax-8b-theta installed effect in huggingface.co for debugging and trial. It also supports api for free installation.