MOSS-VL-Base-0408 is the foundation checkpoint of the MOSS-VL series, part of the OpenMOSS ecosystem dedicated to advancing visual understanding.
Built through four stages of multimodal pretraining only, this checkpoint serves as a high-capacity offline multimodal base model. It provides strong general-purpose visual-linguistic representations across image and video inputs, and is intended primarily as the base model for downstream supervised fine-tuning, alignment, and domain adaptation.
Specifically, the pretraining pipeline is structured into the following four progressive stages:
Stage 1: Vision-language alignment
Stage 2: Large-scale multimodal pretraining
Stage 3: High-quality multimodal pretraining
Stage 4: Annealing and long-context extension
✨ Highlights
📐
Native Dynamic Resolution
MOSS-VL-Base-0408 natively processes images and video frames at their original aspect ratios and resolutions. By preserving the raw spatial layout, it faithfully captures fine visual details across diverse formats—from high-resolution photographs and dense document scans to ultra-wide screenshots.
🎞️
Native Interleaved Image & Video Inputs
The model accepts arbitrary combinations of images and videos within a single sequence. Through a unified end-to-end pipeline, it seamlessly handles complex mixed-modality prompts, multi-image comparisons, and interleaved visual narratives without requiring modality-specific pre-processing.
🏗 Model Architecture
MOSS-VL-Base-0408
adopts a cross-attention-based architecture that decouples visual encoding from cognitive reasoning. Natively supporting interleaved modalities, it provides a multimodal backbone for image and video understanding.
🧩 Absolute Timestamps
To help the model perceive the pacing and duration of events,
MOSS-VL-Base-0408
injects absolute timestamps alongside sampled video frames, giving the reasoning process an explicit temporal reference even at the pretrained base stage.
🧬 Cross-attention RoPE (XRoPE)
MOSS-VL utilizes Cross-attention Rotary Position Embedding (XRoPE), tailored to its cross-attention-based vision-language architecture. This mechanism maps text tokens and visual features into a unified 3D coordinate space defined by Time (t), Height (h), and Width (w), improving spatial-temporal grounding during multimodal reasoning.
MOSS-VL-Base-0408 is a pretrained base checkpoint, and we are actively improving several core capabilities for future iterations:
📄
Stronger OCR, Especially for Long Documents
— We plan to further improve text recognition, document parsing, and long-document understanding. A key focus is achieving near-lossless information extraction and understanding for extremely long and structurally complex inputs, such as accurately parsing texts, tables, and mathematical layouts from multi-page academic papers (dozens of pages) or dense PDF reports without degrading context or structural integrity.
🎬
Expanded Extremely Long Video Understanding
— We aim to significantly extend the model's capacity for comprehending extremely long videos spanning several hours to dozens of hours. This includes advancing temporal reasoning and cross-frame event tracking for continuous analysis of full-length movies, lengthy meetings, or extended surveillance streams, enabling robust retrieval and understanding over ultra-long visual contexts.
We expect future releases to continue strengthening the base model itself while also enabling stronger downstream aligned variants built on top of it.
MOSS-VL-Base-0408 huggingface.co is an AI model on huggingface.co that provides MOSS-VL-Base-0408's model effect (), which can be used instantly with this OpenMOSS-Team MOSS-VL-Base-0408 model. huggingface.co supports a free trial of the MOSS-VL-Base-0408 model, and also provides paid use of the MOSS-VL-Base-0408. Support call MOSS-VL-Base-0408 model through api, including Node.js, Python, http.
MOSS-VL-Base-0408 huggingface.co is an online trial and call api platform, which integrates MOSS-VL-Base-0408's modeling effects, including api services, and provides a free online trial of MOSS-VL-Base-0408, you can try MOSS-VL-Base-0408 online for free by clicking the link below.
OpenMOSS-Team MOSS-VL-Base-0408 online free url in huggingface.co:
MOSS-VL-Base-0408 is an open source model from GitHub that offers a free installation service, and any user can find MOSS-VL-Base-0408 on GitHub to install. At the same time, huggingface.co provides the effect of MOSS-VL-Base-0408 install, users can directly use MOSS-VL-Base-0408 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.