A decoder-only language model with
4,920,384 parameters
, pretrained from
scratch on
200M tokens
of
HuggingFaceFW/fineweb-edu
(sample-10BT).
The largest model in the LowOnMind family and the third point on its scaling
curve, after
LowOnMind-300k
and
LowOnMind-1M
. All three
share an
identical tokenizer, dataset, token budget (200M) and schedule
shape
, so validation loss, bits-per-character and benchmark results are
directly comparable across the series.
It is also the first model in the family whose benchmark performance is
statistically distinguishable from chance.
Architecture
300k
1M
5M
parameters
296,960
985,152
4,920,384
hidden_size
64
96
192
intermediate_size
136 (2.12x)
256 (2.67x)
512 (2.667x)
num_hidden_layers
6
9
12
heads (q / kv)
4 / 2
6 / 2
12 / 4
head_dim
16
16
16
aspect ratio
10.7
10.7
16.0
embedding share
22.1%
10.0%
4.0%
vocab_size
1024
1024
1024 (same tokenizer)
context
512
512
512
tokens seen
200M
200M
200M
tokens/param
673
203
41
Two deviations from the smaller siblings, both deliberate:
Aspect ratio rises from 10.7 to 16.0.
This is the
normal direction when scaling (GPT-2 small sits at 64). Holding 10.7 at this
budget would require roughly 18 layers of hidden_size=160 with an
implausibly wide MLP.
intermediate/hidden is now exactly 8/3 = 2.667
, the standard SwiGLU ratio
used by Llama. LowOnMind-300k was at 2.12 and LowOnMind-1M at 2.67.
The
vocabulary was deliberately left at 1024
rather than raised to something
more appropriate for this scale. A larger vocabulary would compress better
(1024-token byte-level BPE runs about 2.35 characters per token, so 200M tokens
is only ~470MB of text) and would almost certainly improve absolute results.
Keeping it fixed is what makes the three-model comparison valid — the cost is
that this model spends capacity assembling words from fragments that a
4096-token vocabulary would hand it for free.
Modelling code is otherwise byte-identical to the two smaller siblings: GQA,
SwiGLU, RMSNorm, tied embeddings,
QK-Norm
per head,
precomputed RoPE
with automatic re-expansion, residual projections initialized at
std / sqrt(2 * num_layers)
.
Training
data
HuggingFaceFW/fineweb-edu
, sample-10BT
tokens
200M (6,103 steps x 32,768)
sequence length
512
batch size
64
optimizer
AdamW, betas (0.9, 0.95), wd 0.1
lr
1.2e-03 peak, cosine to 1.2e-04, 250 warmup
grad clip
1.0
precision
float16 + GradScaler
hardware
Tesla T4
wall clock
27 min
At 41 tokens per parameter this run is the closest of the three to the
Chinchilla-optimal ratio of roughly 20 — about 2x above it, against 10x for
LowOnMind-1M and 34x for LowOnMind-300k. Train and validation loss tracked
each other throughout; no overfitting.
Results
metric
300k
1M
5M
validation loss
3.2982
2.9908
2.5828
validation perplexity
27.06
19.90
13.23
bits per character
2.030
1.836
1.586
Perplexity is not comparable across tokenizers, but it is comparable across
these three models because they share one. Bits per character
(loss / ln 2 / 2.35 chars-per-token) is the portable figure.
Deltas: -0.4080 nats from LowOnMind-1M (5.0x the parameters),
-0.7154 nats from LowOnMind-300k (16.6x).
Real-word rate
With a 1024-token byte-level vocabulary, no long word exists as a single token —
the model has to assemble every one of them from fragments. The fraction of
emitted words that are real English words was introduced to measure this.
rate
LowOnMind-1M
98.0%
LowOnMind-5M
96.3%
FineWeb-Edu itself (same lexicon)
98.4%
Measured over 64 unconditional samples (5,398 words), using the same reference
lexicon as LowOnMind-1M: words appearing at least 5 times in a 20k-document
sample of the training corpus.
This number went down, and it should not be read as degraded spelling.
The
drop is statistically real (z = 5.38, not sampling noise), but inspecting the
non-words shows what happened:
illuminator
is an ordinary English word,
phillipsburg
is a US town,
shima
is a common element of Japanese place
names. They are counted as errors only because they fall below the reference
lexicon's frequency-5 threshold. The remainder (
hymenola
,
almanine
,
perleti
,
amiravicis
) skew toward proper-noun and Latinate-technical
morphology rather than the malformed common words the metric was built to catch
— LowOnMind-300k produced things like
landship
and
parsetic
, failures of a
different kind.
The metric has a floor problem as well as a ceiling problem.
As a model
improves it emits rarer real vocabulary — names, places, technical terms — which
a frequency-thresholded lexicon scores as wrong. So the measured rate can fall
while actual quality rises. Comparing against a full dictionary with proper-noun
handling, rather than a corpus-frequency cutoff, would be the fix. The 96.3%
figure is reported as-measured for continuity, but it should not be used to rank
these models.
BananaMind Base Bench 1.1
Evaluated on
BananaMind/BananaMind-Base-Bench-1.1
,
the same 350-item English continuation-likelihood benchmark used across the
family, with identical scoring: context and each of the four continuations
tokenized separately with
add_special_tokens=False
, no BOS, selection by
highest mean conditional token log-probability.
Run validity: dataset SHA-256 matched, full schema validation passed, no context
required truncation against the 512-token window.
Category
300k
1M
5M
z vs chance (5M)
Elo (5M)
language_completion
46.0%
52.0%
62.0%
+6.04
1008
world_knowledge
22.0%
22.0%
38.0%
+2.12
881
context_tracking
14.0%
24.0%
32.0%
+1.14
851
quantitative
32.0%
28.0%
28.0%
+0.49
872
logical_reasoning
24.0%
28.0%
26.0%
+0.16
900
commonsense
34.0%
28.0%
24.0%
-0.16
758
code_completion
14.0%
20.0%
16.0%
-1.47
805
300k
1M
5M
Overall Elo
833
843
863
Chance-level Elo (this grid)
805
805
805
Raw accuracy
26.6%
28.9%
32.3%
95% CI
[22.0, 31.2]
[24.2, 33.6]
[27.4, 37.2]
z vs. chance
+0.69
+1.68
+3.15
significant vs. chance
no
no
yes
Difficulty split: easy 30.8%, medium 33.3%, hard 32.8%.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("DedeProGames/LowOnMind-5M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-5M", trust_remote_code=True)
ids = tok("The ", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=64, use_cache=False)[0]))
trust_remote_code=True
is required — the architecture ships as custom modeling
code in the repository.
use_cache=False
is required: this implementation has no
KV cache and recomputes the full window at each generation step.
Limitations
At ~5M parameters this is still a research artifact, not a usable model. Expect
fluent local syntax and register-appropriate structure, but
no reliable
coherence across a paragraph
, no dependable factual knowledge, and no ability
to track state across a passage. Benchmark accuracy of 32.3% is above chance and
far below usefulness. The 1024-token vocabulary caps absolute quality below what
this parameter count could otherwise reach.
The 512-token context and absent KV cache also make it unsuitable for any real
workload.
Runs of DedeProGames LowOnMind-5M on huggingface.co
536
Total runs
0
24-hour runs
6
3-day runs
536
7-day runs
536
30-day runs
More Information About LowOnMind-5M huggingface.co Model
LowOnMind-5M huggingface.co is an AI model on huggingface.co that provides LowOnMind-5M's model effect (), which can be used instantly with this DedeProGames LowOnMind-5M model. huggingface.co supports a free trial of the LowOnMind-5M model, and also provides paid use of the LowOnMind-5M. Support call LowOnMind-5M model through api, including Node.js, Python, http.
LowOnMind-5M huggingface.co is an online trial and call api platform, which integrates LowOnMind-5M's modeling effects, including api services, and provides a free online trial of LowOnMind-5M, you can try LowOnMind-5M online for free by clicking the link below.
DedeProGames LowOnMind-5M online free url in huggingface.co:
LowOnMind-5M is an open source model from GitHub that offers a free installation service, and any user can find LowOnMind-5M on GitHub to install. At the same time, huggingface.co provides the effect of LowOnMind-5M install, users can directly use LowOnMind-5M installed effect in huggingface.co for debugging and trial. It also supports api for free installation.