EphAsad / Icarus

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: July 07 2026
text-generation

Introduction of Icarus

Model Details of Icarus

Icarus — Stage 1 (base pretraining)

A ~9.04M-parameter decoder-only transformer, trained from scratch on HuggingFaceTB/smollm-corpus — the same training corpus used for SmolLM — mixing fineweb-edu-dedup (deduplicated educational web text) and cosmopedia-v2 (synthetic textbooks, articles, and stories) at a 60/40 ratio.

⚠️ Scope of this checkpoint — read before using

This is a Stage 1 base-pretraining checkpoint . It is a free-form text completion model, not an instruction-following or chat model — it has not seen any instruction-tuning data, and a later, separate Stage 2 is planned for that.

What it can and can't do
  • Can : shift vocabulary and register appropriately across topics — e.g. producing historical/geopolitical language for history prompts, biology terminology for science prompts, and basic narrative conventions (named characters, setting, a hook) for story prompts.
  • Can't : maintain factual accuracy or logical consistency across a passage. Generations tend to drift associatively — correct vocabulary and plausible local phrasing, but connections between sentences and specific claims often don't hold up. This is expected behavior for a model this size on this kind of task, not a bug to be fixed by more training alone.
  • Can't solve problems, follow instructions, or hold a conversation — none of that is in scope for this checkpoint.
Model details
Parameters ~9.04M
Architecture Llama-style decoder-only, tied embeddings
Hidden size 224
Intermediate size 600
Layers 12
Attention heads 7 (head dim 32)
Vocab size ~8,000 (custom BPE, trained on this corpus)
Context length 512
Training data
  • Source: HuggingFaceTB/smollm-corpus , configs fineweb-edu-dedup (60% of token budget) and cosmopedia-v2 (40% of token budget) — Cosmopedia was deliberately over-sampled relative to its natural share of the full corpus, since its controlled-complexity synthetic text is disproportionately useful for a model this small (the same reasoning behind why TinyStories-style corpora help tiny models more than raw web text does).
  • Quality filters applied to fineweb-edu-dedup : language == "en" , language_score >= 0.80 , int_score >= 3 (all read from that config's nested metadata struct).
  • cosmopedia-v2 is pre-curated synthetic English text; only a minimum length filter was applied.
  • Packed into 512-token blocks: ~688M total tokens (1,330,218 training blocks + 13,436 held-out validation blocks, under this model's own 8k-vocab tokenizer).
Training procedure
Epochs 2
Effective batch size 128 (64 per-device × 2 grad-accum steps)
Total steps 20,786
LR schedule cosine, 3% warmup
Peak learning rate 3e-4
Precision bf16
Hardware 1× A100 80GB (Google Colab)
Wall-clock training time 1h 1m
Throughput ~726 samples/sec
Validation split 1% held out (13,436 blocks), load_best_model_at_end
Loss curve

Training and validation loss tracked closely together for the entire run, both descending from 4.6 at step 1,000 to a final value of 3.207 (train) / 3.215 (validation) at step 20,786 — a gap of about 0.01, indicating healthy generalization with no overfitting. This corresponds to a final perplexity of ** 24.7–24.9** over this model's ~8k-token vocabulary.

Example usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "EphAsad/icarus"  
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "The history of the Roman Empire"
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(
    **inputs,
    max_new_tokens=100,
    do_sample=True,
    temperature=0.8,
    top_p=0.9,
    repetition_penalty=1.3,   # recommended - avoids short repeated loops,
    no_repeat_ngram_size=3,   # a known tiny-LM decoding artifact at this scale
)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Known limitations
  • Factual and logical drift. Fluent, topic-appropriate vocabulary does not imply correct or internally consistent content — treat all specific claims, dates, and figures in its output as unreliable.
  • No instruction-following. It completes text; it does not follow commands or answer questions in a directed way. That's the target of a planned Stage 2.
  • Repetition without decoding constraints. As with other tiny-LM checkpoints, unconstrained sampling can fall into short repeated loops; use repetition_penalty / no_repeat_ngram_size as shown above.
  • Not benchmarked. No standardized evaluation has been run against this checkpoint. Any future evaluation results belong to a separate, explicitly labeled checkpoint (e.g. a Stage 2 instruction-tuned version), not to this one.
Intended use
  • As a base checkpoint for further fine-tuning (instruction tuning, domain adaptation, etc.).
  • As a research artifact for studying what a small, from-scratch model can learn from a curated real+synthetic educational corpus.
  • Not intended for deployment as a standalone assistant or knowledge source.
Citation

If you use this model, please also cite the underlying corpus and its component datasets:

HuggingFaceTB/smollm-corpus: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus

Runs of EphAsad Icarus on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About Icarus huggingface.co Model

More Icarus license Visit here:

https://choosealicense.com/licenses/mit

Icarus huggingface.co

Icarus huggingface.co is an AI model on huggingface.co that provides Icarus's model effect (), which can be used instantly with this EphAsad Icarus model. huggingface.co supports a free trial of the Icarus model, and also provides paid use of the Icarus. Support call Icarus model through api, including Node.js, Python, http.

EphAsad Icarus online free

Icarus huggingface.co is an online trial and call api platform, which integrates Icarus's modeling effects, including api services, and provides a free online trial of Icarus, you can try Icarus online for free by clicking the link below.

EphAsad Icarus online free url in huggingface.co:

https://huggingface.co/EphAsad/Icarus

Icarus install

Icarus is an open source model from GitHub that offers a free installation service, and any user can find Icarus on GitHub to install. At the same time, huggingface.co provides the effect of Icarus install, users can directly use Icarus installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Icarus install url in huggingface.co:

https://huggingface.co/EphAsad/Icarus

Url of Icarus

Icarus huggingface.co Url

Provider of Icarus huggingface.co

EphAsad
ORGANIZATIONS

Other API from EphAsad

huggingface.co

Total runs: 416
Run Growth: 133
Growth Rate: 32.20%
Updated:June 24 2026
huggingface.co

Total runs: 386
Run Growth: 118
Growth Rate: 31.30%
Updated:June 22 2026
huggingface.co

Total runs: 289
Run Growth: 195
Growth Rate: 67.47%
Updated:June 02 2026
huggingface.co

Total runs: 159
Run Growth: -56
Growth Rate: -38.36%
Updated:June 23 2026
huggingface.co

Total runs: 112
Run Growth: 50
Growth Rate: 44.64%
Updated:June 09 2026
huggingface.co

Total runs: 94
Run Growth: -45
Growth Rate: -46.88%
Updated:June 23 2026
huggingface.co

Total runs: 82
Run Growth: 82
Growth Rate: 100.00%
Updated:September 09 2026
huggingface.co

Total runs: 37
Run Growth: 2
Growth Rate: 5.41%
Updated:May 31 2026
huggingface.co

Total runs: 26
Run Growth: 21
Growth Rate: 80.77%
Updated:December 29 2025
huggingface.co

Total runs: 8
Run Growth: 6
Growth Rate: 75.00%
Updated:December 29 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 13 2026