A
56,902,144-parameter
transformer language model trained
from scratch
on
TinyStories
, a corpus of
simple, repetitive children's stories. It is the 50M scale-up in the
tinystories-24m
→
tinystories-50m
lineage.
v2 (2026-09-25):
retrained with a larger
12288-vocab
BPE tokenizer
(was 8192). The 8192-vocab v1 is fully superseded — same repo, same loader,
better weights. v1's held-out val loss was 1.6566; v2's is
1.3837
.
It writes fluent, on-domain children's stories. It is
not
a general
language model — out-of-domain generation degrades, and it should not be used
for anything beyond the story domain it was trained on.
Architecture
Field
Value
Parameters
56,902,144
(exact; verified against the safetensors header)
Data:
TinyStories (ronendagan/TinyStories),
523,389,481 tokens
after
BPE-12288 re-tokenization (2,119,489 stories, ~9.19 tokens/param), with a
2M-token held-out tail for validation.
Optimizer:
AdamW, cosine LR decay with warmup (peak 6e-4), grad clip 1.0.
Batch:
64, seq 512 → 32,768 tokens/step.
Steps:
15,910 (one full epoch). Best checkpoint at step 13,500.
Hardware:
single NVIDIA RTX 5090 (32 GB).
Final val loss:
1.3924;
best val loss 1.3837
(step 13,500). The
shipped weights are the end-of-run checkpoint (val 1.3924), within 0.009 of
the best.
Evaluated numbers
Held-out perplexity (TinyStories val split):
exp(1.3837) ≈
3.99
(best
ckpt). This is the honest primary metric for a narrow-domain model.
General zero-shot log-likelihood accuracy
(the 12288-vocab tokenizer can
read these datasets, so we report them — v1's 8192-vocab tokenizer could not):
Task
Accuracy
n
BLiMP
64.00%
200
ARC-Easy
51.09%
599
PIQA
45.50%
200
HellaSwag
54.83%
600
These are single-shot, zero-shot, no-few-shot, on a 57M model trained on one
narrow domain — treat them as a scale reference, not a competitive result.
Coherence:
seeded generations are fluent, on-domain, with consistent
characters and correct punctuation. Minor artifacts expected at this scale
(occasional garbled quote char, a couple of logical slips).
Files
File
What
model.safetensors
weights (227 MB, 99 tensors, float32)
tokenizer.json
BPE-12288 tokenizer (
tokenizers
format)
config.json
architecture config
load_model.py
self-contained loader +
TinyStoriesGPT
class
Usage
from load_model import load
model, tok = load()
ids = tok.encode("Once upon a time,")
out = model.generate(torch.tensor([ids]).cuda(), 100, temp=0.8, top_k=40)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
What it is and is not
Is:
a small, from-scratch, on-domain story generator. Good for studying
how a ~57M transformer learns a narrow, repetitive domain.
Is not:
a general-purpose LM. Do not expect coherent output on code,
math, or open-domain text. The low perplexity is domain-specific.
Runs of Compactbot tinystories-50m on huggingface.co
766
Total runs
0
24-hour runs
42
3-day runs
135
7-day runs
135
30-day runs
More Information About tinystories-50m huggingface.co Model
tinystories-50m huggingface.co is an AI model on huggingface.co that provides tinystories-50m's model effect (), which can be used instantly with this Compactbot tinystories-50m model. huggingface.co supports a free trial of the tinystories-50m model, and also provides paid use of the tinystories-50m. Support call tinystories-50m model through api, including Node.js, Python, http.
tinystories-50m huggingface.co is an online trial and call api platform, which integrates tinystories-50m's modeling effects, including api services, and provides a free online trial of tinystories-50m, you can try tinystories-50m online for free by clicking the link below.
Compactbot tinystories-50m online free url in huggingface.co:
tinystories-50m is an open source model from GitHub that offers a free installation service, and any user can find tinystories-50m on GitHub to install. At the same time, huggingface.co provides the effect of tinystories-50m install, users can directly use tinystories-50m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.