MOSS-VL-Base-0708 is the foundation checkpoint of the MOSS-VL 0708 release, part of the OpenMOSS ecosystem for open visual understanding.
Built through multimodal pretraining only, this checkpoint serves as a high-capacity offline multimodal base model. It provides strong general-purpose visual-language representations across image and video inputs, and is intended primarily as the base model for supervised fine-tuning, alignment, and domain adaptation.
The 0708 release keeps the MOSS-VL cross-attention design and a 256K text context window while refreshing the data and pretraining recipe for stronger offline multimodal foundations.
Specifically, the pretraining pipeline follows four progressive stages:
Stage 1: Vision-language alignment
Stage 2: Large-scale multimodal pretraining
Stage 3: High-quality multimodal pretraining
Stage 4: Annealing and long-context extension
Highlights
Strong foundation model: provides general visual-language representations for image, video, and text inputs.
Native dynamic resolution: processes images and video frames at their original aspect ratios and resolutions.
Native interleaved image and video inputs: supports mixed image/video/text sequences in a unified pipeline.
Open base checkpoint: designed for continued pretraining, supervised fine-tuning, alignment, and domain adaptation.
Model Architecture
MOSS-VL-Base-0708 adopts a cross-attention-based vision-language architecture that decouples visual encoding from language reasoning. The model processes images, videos, and text in a unified pipeline and uses cross-attention layers to connect language tokens with visual representations.
Key configuration details:
Item
Value
Parameters
11B
Tensor type
BF16
Context length
256K
Vision patch size
16
Temporal patch size
1
Default video FPS
1.0
Default max video frames
256
Absolute Timestamps
For video inputs, MOSS-VL injects absolute timestamps alongside sampled frames. This helps the base model learn event order, duration, pacing, and temporal localization instead of relying only on frame order.
Cross-attention RoPE (XRoPE)
MOSS-VL uses Cross-attention Rotary Position Embedding (XRoPE), which maps text tokens and visual patches into a unified three-dimensional coordinate space defined by Time (t), Height (h), and Width (w). This gives the model a consistent positional representation for image and video understanding.
Model Performance
MOSS-VL-Base-0708 is intended as a pretrained foundation checkpoint for offline multimodal understanding and model adaptation. Detailed benchmark tables for the 0708 release will be maintained in the MOSS-VL project resources.
offline_batch_generate
accepts independent image/video/text queries. Queries in the same batch should share the same
media_kwargs
and
generate_kwargs
.
MOSS-VL-Base-0708 is a pretrained base checkpoint. It is not instruction-tuned, so applied use cases should generally fine-tune or align it before using it as an assistant-style model.
We are continuing to improve OCR and document understanding, extremely long video understanding, mathematical reasoning, code reasoning, RL post-training, and broader task-specific evaluations for future MOSS-VL releases.
MOSS-VL-Base-0708 huggingface.co is an AI model on huggingface.co that provides MOSS-VL-Base-0708's model effect (), which can be used instantly with this OpenMOSS-Team MOSS-VL-Base-0708 model. huggingface.co supports a free trial of the MOSS-VL-Base-0708 model, and also provides paid use of the MOSS-VL-Base-0708. Support call MOSS-VL-Base-0708 model through api, including Node.js, Python, http.
MOSS-VL-Base-0708 huggingface.co is an online trial and call api platform, which integrates MOSS-VL-Base-0708's modeling effects, including api services, and provides a free online trial of MOSS-VL-Base-0708, you can try MOSS-VL-Base-0708 online for free by clicking the link below.
OpenMOSS-Team MOSS-VL-Base-0708 online free url in huggingface.co:
MOSS-VL-Base-0708 is an open source model from GitHub that offers a free installation service, and any user can find MOSS-VL-Base-0708 on GitHub to install. At the same time, huggingface.co provides the effect of MOSS-VL-Base-0708 install, users can directly use MOSS-VL-Base-0708 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.