rednote-hilab / dots.llm1.inst-FP8-dynamic

huggingface.co
Total runs: 9
24-hour runs: 0
7-day runs: -15
30-day runs: -20
Model's Last Updated: June 24 2025
text-generation

Introduction of dots.llm1.inst-FP8-dynamic

Model Details of dots.llm1.inst-FP8-dynamic

dots1

🤗 Hugging Face |    📑 Paper
🖥️ Demo |   💬 WeChat (微信) |   📕 rednote

Visit our Hugging Face (click links above), search checkpoints with names starting with dots.llm1 or visit the dots1 collection , and you will find all you need! Enjoy!

News
  • 2025.06.06: We released the dots.llm1 series. Check our report for more details!
1. Introduction

The dots.llm1 model is a large-scale MoE model that activates 14B parameters out of a total of 142B parameters, delivering performance on par with state-of-the-art models. Leveraging our meticulously crafted and efficient data processing pipeline, dots.llm1 achieves performance comparable to Qwen2.5-72B after pretrained on high-quality corpus without synthetic data. To foster further research, we open-source intermediate training checkpoints spanning the entire training process, providing valuable insights into the learning dynamics of large language models.

2. Model Summary

This repo contains the base and instruction-tuned dots.llm1 model . which has the following features:

  • Type: A MoE model with 14B activated and 142B total parameters trained on high-quality corpus.
  • Training Stages: Pretraining and SFT.
  • Architecture: Multi-head Attention with QK-Norm in attention Layer, fine-grained MoE utilizing top-6 out of 128 routed experts, plus 2 shared experts.
  • Number of Layers: 62
  • Number of Attention Heads: 32
  • Supported Languages: English, Chinese
  • Context Length: 32,768 tokens
  • License: MIT

The highlights from dots.llm1 include:

  • Enhanced Data Processing : We propose a scalable and fine-grained three-stage data processing framework designed to generate large-scale, high-quality and diverse data for pretraining.
  • No Synthetic Data during Pretraining : High-quality non-synthetic tokens was used in base model pretraining.
  • Performance and Cost Efficiency : dots.llm1 is an open-source model that activates only 14B parameters at inference, delivering both comprehensive capabilities and high computational efficiency.
  • Infrastructure : We introduce an innovative MoE all-to-all communication and computation overlapping recipe based on interleaved 1F1B pipeline scheduling and an efficient grouped GEMM implementation to boost computational efficiency.
  • Open Accessibility to Model Dynamics : Intermediate model checkpoints are released spanning the entire training process, facilitating future research into the learning dynamics of large language models.
3. dots.llm1.inst.FP8-dynamic

We release the quantized dots.llm1.inst.FP8-dynamic model, which retains approximately 98% of the original performance after quantization.

Docker (vllm)

For convenience, we recommend running vLLM inference using our Docker image rednotehilab/dots1:vllm-openai-v0.9.1 , , which is available on Docker Hub .

python3 -m vllm.entrypoints.openai.api_server \
  --model rednote-hilab/dots.llm1.inst.FP8-dynamic \
  --tensor-parallel-size 4 \
  --pipeline-parallel-size 1 \
  --trust-remote-code \
  --served-model-name dots1
Inference with huggingface

We are working to merge it into Transformers ( PR #38143 ).

Chat Completion
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig

model_name = "rednote-hilab/dots.llm1.inst-FP8-dynamic"
tokenizer = AutoTokenizer.from_pretrained(model_name)

model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype="auto")

messages = [
    {"role": "user", "content": "Write a piece of quicksort code in C++"}
]
input_tensor = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
outputs = model.generate(input_tensor.to(model.device), max_new_tokens=200)

result = tokenizer.decode(outputs[0][input_tensor.shape[1]:], skip_special_tokens=True)
print(result)
Citation

If you find dots.llm1 is useful or want to use in your projects, please kindly cite our paper:

@misc{huo2025dotsllm1technicalreport,
      title={dots.llm1 Technical Report}, 
      author={Bi Huo and Bin Tu and Cheng Qin and Da Zheng and Debing Zhang and Dongjie Zhang and En Li and Fu Guo and Jian Yao and Jie Lou and Junfeng Tian and Li Hu and Ran Zhu and Shengdong Chen and Shuo Liu and Su Guang and Te Wo and Weijun Zhang and Xiaoming Shi and Xinxin Peng and Xing Wu and Yawen Liu and Yuqiu Ji and Ze Wen and Zhenhai Liu and Zichao Li and Zilong Liao},
      year={2025},
      eprint={2506.05767},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2506.05767}, 
}

Runs of rednote-hilab dots.llm1.inst-FP8-dynamic on huggingface.co

9
Total runs
0
24-hour runs
0
3-day runs
-15
7-day runs
-20
30-day runs

More Information About dots.llm1.inst-FP8-dynamic huggingface.co Model

More dots.llm1.inst-FP8-dynamic license Visit here:

https://choosealicense.com/licenses/mit

dots.llm1.inst-FP8-dynamic huggingface.co

dots.llm1.inst-FP8-dynamic huggingface.co is an AI model on huggingface.co that provides dots.llm1.inst-FP8-dynamic's model effect (), which can be used instantly with this rednote-hilab dots.llm1.inst-FP8-dynamic model. huggingface.co supports a free trial of the dots.llm1.inst-FP8-dynamic model, and also provides paid use of the dots.llm1.inst-FP8-dynamic. Support call dots.llm1.inst-FP8-dynamic model through api, including Node.js, Python, http.

dots.llm1.inst-FP8-dynamic huggingface.co Url

https://huggingface.co/rednote-hilab/dots.llm1.inst-FP8-dynamic

rednote-hilab dots.llm1.inst-FP8-dynamic online free

dots.llm1.inst-FP8-dynamic huggingface.co is an online trial and call api platform, which integrates dots.llm1.inst-FP8-dynamic's modeling effects, including api services, and provides a free online trial of dots.llm1.inst-FP8-dynamic, you can try dots.llm1.inst-FP8-dynamic online for free by clicking the link below.

rednote-hilab dots.llm1.inst-FP8-dynamic online free url in huggingface.co:

https://huggingface.co/rednote-hilab/dots.llm1.inst-FP8-dynamic

dots.llm1.inst-FP8-dynamic install

dots.llm1.inst-FP8-dynamic is an open source model from GitHub that offers a free installation service, and any user can find dots.llm1.inst-FP8-dynamic on GitHub to install. At the same time, huggingface.co provides the effect of dots.llm1.inst-FP8-dynamic install, users can directly use dots.llm1.inst-FP8-dynamic installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

dots.llm1.inst-FP8-dynamic install url in huggingface.co:

https://huggingface.co/rednote-hilab/dots.llm1.inst-FP8-dynamic

Url of dots.llm1.inst-FP8-dynamic

dots.llm1.inst-FP8-dynamic huggingface.co Url

Provider of dots.llm1.inst-FP8-dynamic huggingface.co

rednote-hilab
ORGANIZATIONS

Other API from rednote-hilab

huggingface.co

Total runs: 419.1K
Run Growth: 136.1K
Growth Rate: 32.48%
Updated:October 31 2025