InfLLM-V2-Short-Dense-Base
is the foundational base model for the InfLLM-V2 long-context training pipeline.
This model is pre-trained on a large corpus of
short-text data
and utilizes a standard
dense attention
mechanism. It serves as the starting checkpoint for the continued training phase, which unlocks the advanced long-context capabilities seen in the final sparse model.
It is highly performant on short-text tasks and provides a solid foundation for further fine-tuning or continued training.
📌 Role in the InfLLM-V2 Ecosystem
This model is the crucial first step in the InfLLM-V2 training workflow. The entire process is designed to be transparent and reproducible:
The result is the
InfLLM-V2-Long-Sparse-Base
, which is equipped with powerful sparse attention for long-context tasks.
💻 How to Use
As a standard dense-attention model, you can use it directly with the
transformers
library without any special configuration.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
# Set device
device = "cuda"if torch.cuda.is_available() else"cpu"# Load model and tokenizer
model_id = "openbmb/InfLLM-V2-Short-Dense-Base"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id,trust_remote_code=True).to(device,dtype=torch.bfloat16)
# Create a prompt
prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
# Generate text
outputs = model.generate(**inputs, max_new_tokens=10)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(generated_text)
# Expected output: The capital of France is Paris.
Note
: This model is optimized for short sequences. For long-context capabilities, please use the final
InfLLM-V2-Long-Sparse-Base
model.
Citation
If you use our work in your research, please cite our paper:
@misc{zhao2025infllmv2densesparseswitchableattention,
title={InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation},
author={Weilin Zhao and Zihan Zhou and Zhou Su and Chaojun Xiao and Yuxuan Li and Yanghao Li and Yudi Zhang and Weilun Zhao and Zhen Li and Yuxiang Huang and Ao Sun and Xu Han and Zhiyuan Liu},
year={2025},
eprint={2509.24663},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.24663},
}
Runs of openbmb InfLLM-V2-Short-Dense-Base on huggingface.co
122
Total runs
0
24-hour runs
11
3-day runs
100
7-day runs
100
30-day runs
More Information About InfLLM-V2-Short-Dense-Base huggingface.co Model
More InfLLM-V2-Short-Dense-Base license Visit here:
InfLLM-V2-Short-Dense-Base huggingface.co is an AI model on huggingface.co that provides InfLLM-V2-Short-Dense-Base's model effect (), which can be used instantly with this openbmb InfLLM-V2-Short-Dense-Base model. huggingface.co supports a free trial of the InfLLM-V2-Short-Dense-Base model, and also provides paid use of the InfLLM-V2-Short-Dense-Base. Support call InfLLM-V2-Short-Dense-Base model through api, including Node.js, Python, http.
InfLLM-V2-Short-Dense-Base huggingface.co is an online trial and call api platform, which integrates InfLLM-V2-Short-Dense-Base's modeling effects, including api services, and provides a free online trial of InfLLM-V2-Short-Dense-Base, you can try InfLLM-V2-Short-Dense-Base online for free by clicking the link below.
openbmb InfLLM-V2-Short-Dense-Base online free url in huggingface.co:
InfLLM-V2-Short-Dense-Base is an open source model from GitHub that offers a free installation service, and any user can find InfLLM-V2-Short-Dense-Base on GitHub to install. At the same time, huggingface.co provides the effect of InfLLM-V2-Short-Dense-Base install, users can directly use InfLLM-V2-Short-Dense-Base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
InfLLM-V2-Short-Dense-Base install url in huggingface.co: