dmis-lab / OSP-1.4B-100B-Muon-Only

huggingface.co
Total runs: 8
24-hour runs: -2
7-day runs: -8
30-day runs: -16
Model's Last Updated: June 25 2025

Introduction of OSP-1.4B-100B-Muon-Only

Model Details of OSP-1.4B-100B-Muon-Only

Outlier-Safe Pre-Training

arXiv Models code

Introduction

Quantization plays a crucial role in deploying Large Language Models (LLMs) in resource-constrained environments. However, the presence of outlier features significantly hinders low-bit quantization. While many studies address this problem in a post-hoc manner to make use of already pre-trained models, the importance of handling outliers during pre-training is often underestimated.

Our work, Outlier-Safe Pre-Training (OSP) , proposes a practical approach to training models that are robust to outliers from the start, without sacrificing performance or efficiency. Specifically, OSP focuses on the following goals:

  1. 📈 Scaling to production-level training requirements
    Prior methods for quantization-friendly pre-training are often limited to small-scale experiments (e.g., models under 1B parameters or 100B tokens). In contrast, we train a 1.4B-parameter model on 1 trillion tokens, demonstrating that OSP is effective at production scale.

  2. Maintaining computational efficiency comparable to standard training
    A method that prevents outliers but significantly reduces efficiency is unlikely to gain adoption. OSP introduces only a ~2% slowdown while reducing GPU memory usage, making it appealing for those seeking to train quantization-friendly foundation models from scratch.

  3. 🧩 Ensuring full compatibility with existing inference pipelines
    We prioritize compatibility with widely adopted inference frameworks such as vLLM and SGLang. Rather than introducing architectural changes that break compatibility, OSP preserves computational invariance, allowing models to be directly integrated into existing pipelines without additional effort.

Model Checkpoints
Final Models

The models were trained on 1 trillion tokens, following the pre-training recipe of SmolLM . Specifically, training was conducted using the smollm-corpus , a mixture of FineWeb-Edu, Cosmopedia, and Python-Edu.

Ablation Models
Model Optimizer SSNorm EmbProj Ex. Kurt. Had. 4-4-4
Avg. PPL
🤗 OSP-1.4B-100B-Adam Adam 1818.56
26.8
26.9
8e4
3e4
🤗 OSP-1.4B-100B-Muon-Only Muon†
(w/o Adam)
361.35
26.3
33.1
8e5
24.8
🤗 OSP-1.4B-100B-Muon Muon 1575.12
29.0
38.4
1e4
15.8
🤗 OSP-1.4B-100B-Muon-SSNorm Muon 66.69
36.4
38.3
44.2
34.1
🤗 OSP-1.4B-100B-Muon-EmbProj Muon 703.23
30.4
36.2
114.6
22.3
🤗 OSP-1.4B-100B-Muon-SSNorm-EmbProj Muon 0.04
37.5
38.9
19.6
13.5
†Model configuration that disables decoupled embedding optimization by training with Muon optimizer without Adam optimization on embedding layers
Training
Model
  • Architecture: Llama
  • Pretraining tokens: 100 billion tokens
  • Precision: bfloat16
Hardware
Software
Disclaimer

This model family was trained to demonstrate the effectiveness of eliminating outlier occurrences and improving quantization-friendliness. All models are base models, i.e., no instruction tuning or human alignment was applied. These models are not intended for chatting, conversation, or assistant purposes. They may contain toxic or harmful content. Their best use is for evaluating performance degradation on benchmarks after low-bit quantization.

Citation
@article{park2025osp,
      title={Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models}, 
      author={Jungwoo Park and Taewhoo Lee and Chanwoong Yoon and Hyeon Hwang and Jaewoo Kang},
      year={2025},
      eprint={2506.19697},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2506.19697}, 
}

Runs of dmis-lab OSP-1.4B-100B-Muon-Only on huggingface.co

8
Total runs
-2
24-hour runs
-3
3-day runs
-8
7-day runs
-16
30-day runs

More Information About OSP-1.4B-100B-Muon-Only huggingface.co Model

OSP-1.4B-100B-Muon-Only huggingface.co

OSP-1.4B-100B-Muon-Only huggingface.co is an AI model on huggingface.co that provides OSP-1.4B-100B-Muon-Only's model effect (), which can be used instantly with this dmis-lab OSP-1.4B-100B-Muon-Only model. huggingface.co supports a free trial of the OSP-1.4B-100B-Muon-Only model, and also provides paid use of the OSP-1.4B-100B-Muon-Only. Support call OSP-1.4B-100B-Muon-Only model through api, including Node.js, Python, http.

OSP-1.4B-100B-Muon-Only huggingface.co Url

https://huggingface.co/dmis-lab/OSP-1.4B-100B-Muon-Only

dmis-lab OSP-1.4B-100B-Muon-Only online free

OSP-1.4B-100B-Muon-Only huggingface.co is an online trial and call api platform, which integrates OSP-1.4B-100B-Muon-Only's modeling effects, including api services, and provides a free online trial of OSP-1.4B-100B-Muon-Only, you can try OSP-1.4B-100B-Muon-Only online for free by clicking the link below.

dmis-lab OSP-1.4B-100B-Muon-Only online free url in huggingface.co:

https://huggingface.co/dmis-lab/OSP-1.4B-100B-Muon-Only

OSP-1.4B-100B-Muon-Only install

OSP-1.4B-100B-Muon-Only is an open source model from GitHub that offers a free installation service, and any user can find OSP-1.4B-100B-Muon-Only on GitHub to install. At the same time, huggingface.co provides the effect of OSP-1.4B-100B-Muon-Only install, users can directly use OSP-1.4B-100B-Muon-Only installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

OSP-1.4B-100B-Muon-Only install url in huggingface.co:

https://huggingface.co/dmis-lab/OSP-1.4B-100B-Muon-Only

Url of OSP-1.4B-100B-Muon-Only

OSP-1.4B-100B-Muon-Only huggingface.co Url

Provider of OSP-1.4B-100B-Muon-Only huggingface.co

dmis-lab
ORGANIZATIONS

Other API from dmis-lab

huggingface.co

Total runs: 100.5K
Run Growth: -44.8K
Growth Rate: -44.60%
Updated:May 20 2021
huggingface.co

Total runs: 77
Run Growth: -77
Growth Rate: -100.00%
Updated:October 27 2021
huggingface.co

Total runs: 30
Run Growth: 13
Growth Rate: 43.33%
Updated:September 11 2024
huggingface.co

Total runs: 20
Run Growth: 10
Growth Rate: 50.00%
Updated:September 11 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 02 2025