MobileMoE
is a family of on-device Mixture-of-Experts (MoE) language models with sub-billion
active
parameters, designed to push the quality–efficiency Pareto frontier for on-device LLMs, including three model scales (S/M/L): 0.3B/0.5B/0.9B active parameters (1.3B/2.8B/5.3B total), with <3 GB INT4 weight footprints to fit in mobile DRAM. Each scale is released in three variants: a
Base
model (pre-training + mid-training), an
SFT
model (supervised fine-tuning), and a
QAT
model (quantization-aware training). You are currently in the
MobileMoE-S-SFT
repository — the instruction-tuned 0.3B-active model.
MobileMoE establishes a new Pareto frontier for on-device LLMs.
Average benchmark accuracy, computed over 14 benchmarks spanning commonsense, knowledge, science, comprehension, and reasoning, is plotted against (a) per-token inference compute
F
inf
= 2
N
act
(GFLOPs) and (b) total parameters
N
total
(B); in (b), x-axis tick labels show total params (B) | projected INT4 memory (GB). Accuracy is shown for the instruction-tuned models.
Key Features
A new Pareto frontier for on-device LLMs.
Across 14 foundational benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs at 2–4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters.
Scaling-law-derived architecture.
The architecture is derived from an on-device MoE scaling law that jointly optimizes under mobile
memory
and
compute
constraints, identifying an on-device sweet spot: moderate sparsity, with fine-grained experts and shared expert.
Four-stage recipe.
Pre-training → mid-training → instruction fine-tuning → INT4 quantization-aware training, all on open-source datasets.
Model Information
Model:
MobileMoE-S-SFT (instruction-tuned)
Active Parameters:
272M
Total Parameters:
1.3B
Layers:
20
Model Dimension:
768
Attention Heads:
12
KV Heads:
4 (GQA)
Head Dimension:
64
Routed Experts:
60 (fine-grained, FFN hidden dim 384 each)
Active Experts per Token:
4 (top-k sigmoid routing, with normalization)
Shared Expert:
1, always on (FFN hidden dim 1536)
Vocabulary Size:
128,256
Other Features:
QK-Norm, tied input/output embeddings, RoPE (θ = 500,000)
Chat Template:
Yes (end-of-turn token
<|eot|>
)
Input Modality:
Text
Output Modality:
Text
Languages:
English
Training Stages:
Pre-training → mid-training → supervised fine-tuning (SFT)
Context Length:
8,192 tokens
Precision:
BF16
Model Developer:
Meta
Model Release Date:
Aug 2026
License:
MobileMoE is FAIR NC licensed
Results
All results below are for the
instruction-tuned (SFT)
models. We re-evaluated every model under identical settings in non-thinking mode with greedy decoding, using
lm-eval
together with the official
allenai/IFBench
package. Few-shot counts are shown in parentheses after benchmark names; benchmarks without a count are evaluated 0-shot. The MobileMoE results use the exact weights in this repository, which include brief fine-tuning with self-identity beyond the SFT checkpoint in the
technical report
, resulting in a small difference: a foundational-benchmark average of 47.2 here versus 46.7 in the report.
Foundational benchmarks
Capability
Benchmark
Gemma 3 270M
SmolLM2 360M
MobileMoE-S
Active / total params
270M
362M
272M / 1.3B
Commonsense Reasoning
HellaSwag
39.4
56.9
56.1
PIQA
67.1
71.6
74.8
SIQA
39.6
40.6
43.1
WinoGrande
53.0
57.4
59.6
Knowledge
MMLU (5-shot)
26.5
25.9
42.9
NaturalQuestions (5-shot)
2.8
6.4
10.9
TriviaQA (5-shot)
9.1
20.4
30.5
Science
ARC-Challenge (25-shot)
27.7
38.8
46.2
ARC-Easy
50.5
49.1
73.6
OpenBookQA
35.0
36.2
32.6
Reading
BoolQ
56.1
42.5
72.7
DROP (3-shot)
11.0
15.2
33.1
Reasoning
BIG-Bench Hard (3-shot)
31.8
30.5
32.5
GSM8K (8-shot)
5.8
10.0
52.4
Average
32.5
35.8
47.2
Other capabilities
Capability
Benchmark
Gemma 3 270M
SmolLM2 360M
MobileMoE-S
Math
MATH-500 (4-shot)
7.2
3.8
18.8
GSM-Plus (5-shot)
4.3
4.6
28.9
Avg
5.7
4.2
23.8
Code
HumanEval
12.8
0.0
46.3
MBPP (3-shot)
9.8
22.8
27.4
Avg
11.3
11.4
36.9
Instruction Following
IFEval
31.2
40.2
59.5
IFBench
11.2
19.1
14.2
Avg
21.2
29.7
36.8
Training
MobileMoE uses a four-stage recipe. This checkpoint is the output of stage 3 (supervised fine-tuning).
MobileMoE four-stage training recipe:
pre-training (PT) → mid-training (MT) → instruct supervised fine-tuning (SFT) → quantization-aware training (QAT) with INT4 precision.
Pre-training
Mid-training
SFT
QAT
Context length
2,048
8,192
8,192
8,192
Total tokens
~6T
~500B
~126B
~21B
Peak learning rate
4×10
-4
4×10
-5
4×10
-6
4×10
-6
LR schedule
Cosine
Linear
Cosine
Cosine
Token dispatch
drop-and-pad
drop-and-pad
dropless
dropless
How to use
MobileMoE uses a custom architecture (
model_type: mobilemoe
) that is not yet part of upstream
transformers
, so
trust_remote_code=True
is required
. The modeling code ships in this repo (
configuration_mobilemoe.py
,
modeling_mobilemoe.py
).
For batch evaluation we recommend vLLM (≥ 0.10.2) with
enforce_eager=True
.
Known issues.
Loading the tokenizer on
transformers
4.57.6 prints a
fix_mistral_regex=True
warning. Please ignore it and do not set the flag — MobileMoE uses the Llama-3 128k text vocabulary, whose default tokenization is already correct.
Chat
This
instruction-tuned model
includes a chat template. Format prompts with
apply_chat_template
; the template uses
<|eot|>
to mark the end of each turn.
For multi-turn conversations, append each generated reply to
messages
with the
assistant
role. This ensures that each subsequent prompt includes the complete conversation history:
messages = []
for user_message in ["Who are you?", "Why are open-source on-device language models great?"]:
messages.append({"role": "user", "content": user_message})
input_ids = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
input_ids,
attention_mask=torch.ones_like(input_ids),
max_new_tokens=1024,
do_sample=False,
temperature=None,
top_p=None,
pad_token_id=tokenizer.eos_token_id,
)
reply = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True).strip()
messages.append({"role": "assistant", "content": reply})
print(reply)
Citation
@article{chen2026mobilemoe,
title={MobileMoE: Scaling On-Device Mixture of Experts},
author={Chen, Yanbei and Huang, Hanxian and Chang, Ernie and Szwejbka, Jacob and Desai, Digant and Liu, Zechun and Chandra, Vikas and Krishnamoorthi, Raghuraman},
journal={arXiv preprint arXiv:2605.27358},
year={2026}
}
MobileMoE-S-SFT huggingface.co is an AI model on huggingface.co that provides MobileMoE-S-SFT's model effect (), which can be used instantly with this facebook MobileMoE-S-SFT model. huggingface.co supports a free trial of the MobileMoE-S-SFT model, and also provides paid use of the MobileMoE-S-SFT. Support call MobileMoE-S-SFT model through api, including Node.js, Python, http.
MobileMoE-S-SFT huggingface.co is an online trial and call api platform, which integrates MobileMoE-S-SFT's modeling effects, including api services, and provides a free online trial of MobileMoE-S-SFT, you can try MobileMoE-S-SFT online for free by clicking the link below.
facebook MobileMoE-S-SFT online free url in huggingface.co:
MobileMoE-S-SFT is an open source model from GitHub that offers a free installation service, and any user can find MobileMoE-S-SFT on GitHub to install. At the same time, huggingface.co provides the effect of MobileMoE-S-SFT install, users can directly use MobileMoE-S-SFT installed effect in huggingface.co for debugging and trial. It also supports api for free installation.