Emhotob
is a
~51.8M-parameter
Llama-architecture language model
pre-trained
from scratch on ~20 billion Arabic tokens
with a
2048-token
context window. It is
a
proof-of-concept
that a
tiny
model — small enough to run on a CPU — can be
trained from scratch to produce coherent Arabic.
This is a
base (pretrained) model
: a next-token predictor with
no instruction or
chat tuning
. It is the foundation for the fine-tuned variants below.
Architecture & training recipe:
derived from
SupraLabs/Supra-50M-Base
("Project Chimera"), re-run
from scratch on an Arabic corpus
with a custom Arabic tokenizer.
الملخص بالعربية:
«إمحوتب» نموذج لغوي عربي صغير (~51.8 مليون معامل) بمعمارية Llama،
مُدرَّب
من الصفر على ~20 مليار رمز (token) عربي
بطول سياق 2048. نموذج
أساس
(Base)
بدون ضبط تعليمات أو محادثة. الهدف: إثبات إمكانية تدريب نموذج عربي مفيد على نطاق
صغير جدًا يعمل حتى على المعالج (CPU). المعمارية وسكربتات التدريب مشتقة من
SupraLabs/Supra-50M-Base
.
The architecture and training loop follow
SupraLabs/Supra-50M-Base
("Project Chimera — 50M Llama"). The pretraining script (
train.py
) is included in
this repository.
Usage
The tokenizer uses the
TokenizersBackend
class, which requires
transformers>=5.12
.
This is a
base model
— prompt it as a text completer, not a chat assistant:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "oddadmix/50M-2048-Emhotob"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
prompt = "اللغة العربية هي"
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(
**ids, max_new_tokens=128, do_sample=True,
temperature=0.8, top_p=0.9, repetition_penalty=1.2,
)
print(tok.decode(out[0], skip_special_tokens=True))
Intended use.
A from-scratch Arabic base model for research and as a fine-tuning
starting point; a CPU-friendly baseline for tiny Arabic SLM experiments.
Limitations.
As a
base model
it does not follow instructions or hold a
conversation out of the box — fine-tune it first. At ~50M parameters it is a
proof of concept
: expect limited world knowledge, weak reasoning, and repetition
(use a
repetition_penalty
). Training data is web text (
fineweb-edu-ar
), so it
carries that corpus's biases. Not suitable for factual, medical, legal, or financial use.
50M-2048-Emhotob huggingface.co is an AI model on huggingface.co that provides 50M-2048-Emhotob's model effect (), which can be used instantly with this oddadmix 50M-2048-Emhotob model. huggingface.co supports a free trial of the 50M-2048-Emhotob model, and also provides paid use of the 50M-2048-Emhotob. Support call 50M-2048-Emhotob model through api, including Node.js, Python, http.
50M-2048-Emhotob huggingface.co is an online trial and call api platform, which integrates 50M-2048-Emhotob's modeling effects, including api services, and provides a free online trial of 50M-2048-Emhotob, you can try 50M-2048-Emhotob online for free by clicking the link below.
oddadmix 50M-2048-Emhotob online free url in huggingface.co:
50M-2048-Emhotob is an open source model from GitHub that offers a free installation service, and any user can find 50M-2048-Emhotob on GitHub to install. At the same time, huggingface.co provides the effect of 50M-2048-Emhotob install, users can directly use 50M-2048-Emhotob installed effect in huggingface.co for debugging and trial. It also supports api for free installation.