A ~9.04M-parameter decoder-only transformer, trained
from scratch
on
HuggingFaceTB/smollm-corpus
— the same training corpus used for
SmolLM
— mixing
fineweb-edu-dedup
(deduplicated educational web text) and
cosmopedia-v2
(synthetic textbooks, articles, and stories) at a 60/40 ratio.
⚠️ Scope of this checkpoint — read before using
This is a
Stage 1 base-pretraining checkpoint
. It is a free-form text
completion model, not an instruction-following or chat model — it has not
seen any instruction-tuning data, and a later, separate Stage 2 is planned for
that.
What it can and can't do
Can
: shift vocabulary and register appropriately across topics — e.g.
producing historical/geopolitical language for history prompts, biology
terminology for science prompts, and basic narrative conventions (named
characters, setting, a hook) for story prompts.
Can't
: maintain factual accuracy or logical consistency across a
passage. Generations tend to drift associatively — correct vocabulary and
plausible local phrasing, but connections between sentences and specific
claims often don't hold up. This is expected behavior for a model this
size on this kind of task, not a bug to be fixed by more training alone.
Can't
solve problems, follow instructions, or hold a conversation —
none of that is in scope for this checkpoint.
Model details
Parameters
~9.04M
Architecture
Llama-style decoder-only, tied embeddings
Hidden size
224
Intermediate size
600
Layers
12
Attention heads
7 (head dim 32)
Vocab size
~8,000 (custom BPE, trained on this corpus)
Context length
512
Training data
Source:
HuggingFaceTB/smollm-corpus
, configs
fineweb-edu-dedup
(60% of
token budget) and
cosmopedia-v2
(40% of token budget) — Cosmopedia was
deliberately over-sampled relative to its natural share of the full corpus,
since its controlled-complexity synthetic text is disproportionately useful
for a model this small (the same reasoning behind why TinyStories-style
corpora help tiny models more than raw web text does).
Quality filters applied to
fineweb-edu-dedup
:
language == "en"
,
language_score >= 0.80
,
int_score >= 3
(all read from that config's
nested
metadata
struct).
cosmopedia-v2
is pre-curated synthetic English text; only a minimum
length filter was applied.
Packed into 512-token blocks:
~688M total tokens
(1,330,218 training
blocks + 13,436 held-out validation blocks, under this model's own 8k-vocab
tokenizer).
Training procedure
Epochs
2
Effective batch size
128 (64 per-device × 2 grad-accum steps)
Total steps
20,786
LR schedule
cosine, 3% warmup
Peak learning rate
3e-4
Precision
bf16
Hardware
1× A100 80GB (Google Colab)
Wall-clock training time
1h 1m
Throughput
~726 samples/sec
Validation split
1% held out (13,436 blocks),
load_best_model_at_end
Loss curve
Training and validation loss tracked closely together for the entire run,
both descending from
4.6 at step 1,000 to a final value of
3.207 (train) /
3.215 (validation)
at step 20,786 — a gap of about 0.01, indicating healthy
generalization with no overfitting. This corresponds to a final perplexity of
**
24.7–24.9** over this model's ~8k-token vocabulary.
Example usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "EphAsad/icarus"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "The history of the Roman Empire"
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.8,
top_p=0.9,
repetition_penalty=1.3, # recommended - avoids short repeated loops,
no_repeat_ngram_size=3, # a known tiny-LM decoding artifact at this scale
)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Known limitations
Factual and logical drift.
Fluent, topic-appropriate vocabulary does
not imply correct or internally consistent content — treat all specific
claims, dates, and figures in its output as unreliable.
No instruction-following.
It completes text; it does not follow
commands or answer questions in a directed way. That's the target of a
planned Stage 2.
Repetition without decoding constraints.
As with other tiny-LM
checkpoints, unconstrained sampling can fall into short repeated loops;
use
repetition_penalty
/
no_repeat_ngram_size
as shown above.
Not benchmarked.
No standardized evaluation has been run against this
checkpoint. Any future evaluation results belong to a separate, explicitly
labeled checkpoint (e.g. a Stage 2 instruction-tuned version), not to this
one.
Intended use
As a base checkpoint for further fine-tuning (instruction tuning, domain
adaptation, etc.).
As a research artifact for studying what a small, from-scratch model can
learn from a curated real+synthetic educational corpus.
Not intended for deployment as a standalone assistant or knowledge source.
Citation
If you use this model, please also cite the underlying corpus and its
component datasets:
Icarus huggingface.co is an AI model on huggingface.co that provides Icarus's model effect (), which can be used instantly with this EphAsad Icarus model. huggingface.co supports a free trial of the Icarus model, and also provides paid use of the Icarus. Support call Icarus model through api, including Node.js, Python, http.
Icarus huggingface.co is an online trial and call api platform, which integrates Icarus's modeling effects, including api services, and provides a free online trial of Icarus, you can try Icarus online for free by clicking the link below.
Icarus is an open source model from GitHub that offers a free installation service, and any user can find Icarus on GitHub to install. At the same time, huggingface.co provides the effect of Icarus install, users can directly use Icarus installed effect in huggingface.co for debugging and trial. It also supports api for free installation.