MOSS-VL-Instruct-0708 is the instruction-tuned checkpoint of the MOSS-VL 0708 release, part of the OpenMOSS ecosystem for open visual understanding.
Built on top of MOSS-VL-Base-0708 through supervised fine-tuning (SFT), this checkpoint is designed as a high-performance offline multimodal model. It supports image understanding, OCR, document parsing, visual reasoning, instruction following, and video understanding, with particular strength in long-form video comprehension, temporal reasoning, action recognition, and fine-grained event localization.
The 0708 release keeps the MOSS-VL cross-attention design and a 256K text context window while refreshing the data and instruction-tuning recipe for stronger offline multimodal usage.
Highlights
Strong video understanding: designed for long videos, temporal reasoning, action recognition, and second-level event localization.
General multimodal perception: supports image understanding, fine-grained visual recognition, OCR, and document analysis.
Reliable instruction following: SFT aligns the base checkpoint with user instructions across image, video, and text tasks.
Open model family: released together with MOSS-VL-Base-0708 for continued pretraining, fine-tuning, and applied research.
Model Architecture
MOSS-VL-Instruct-0708 adopts a cross-attention-based vision-language architecture that decouples visual encoding from language reasoning. The model processes images, videos, and text in a unified pipeline and uses cross-attention layers to connect language tokens with visual representations.
Key configuration details:
Item
Value
Parameters
11B
Tensor type
BF16
Context length
256K
Vision patch size
16
Temporal patch size
1
Default video FPS
1.0
Default max video frames
256
Absolute Timestamps
For video inputs, MOSS-VL injects absolute timestamps alongside sampled frames. This helps the model reason about event order, duration, pacing, and temporal localization instead of relying only on frame order.
Cross-attention RoPE (XRoPE)
MOSS-VL uses Cross-attention Rotary Position Embedding (XRoPE), which maps text tokens and visual patches into a unified three-dimensional coordinate space defined by Time (t), Height (h), and Width (w). This gives the model a consistent positional representation for image and video reasoning.
Model Performance
MOSS-VL-Instruct-0708 is intended for offline multimodal evaluation across visual perception, multimodal reasoning, OCR/document understanding, and video understanding. Detailed benchmark tables for the 0708 release will be maintained in the MOSS-VL project resources.
offline_batch_generate
accepts independent image/video/text queries. Queries in the same batch should share the same
media_kwargs
and
generate_kwargs
.
MOSS-VL-Instruct-0708 is optimized for general offline multimodal understanding. Very dense videos, highly specialized domains, precise small-text OCR, and tasks requiring strict numerical reasoning may still require task-specific prompting, sampling choices, or fine-tuning.
We are continuing to improve mathematical reasoning, code reasoning, RL post-training, and broader task-specific evaluations for future MOSS-VL releases.
MOSS-VL-Instruct-0708 huggingface.co is an AI model on huggingface.co that provides MOSS-VL-Instruct-0708's model effect (), which can be used instantly with this OpenMOSS-Team MOSS-VL-Instruct-0708 model. huggingface.co supports a free trial of the MOSS-VL-Instruct-0708 model, and also provides paid use of the MOSS-VL-Instruct-0708. Support call MOSS-VL-Instruct-0708 model through api, including Node.js, Python, http.
MOSS-VL-Instruct-0708 huggingface.co is an online trial and call api platform, which integrates MOSS-VL-Instruct-0708's modeling effects, including api services, and provides a free online trial of MOSS-VL-Instruct-0708, you can try MOSS-VL-Instruct-0708 online for free by clicking the link below.
OpenMOSS-Team MOSS-VL-Instruct-0708 online free url in huggingface.co:
MOSS-VL-Instruct-0708 is an open source model from GitHub that offers a free installation service, and any user can find MOSS-VL-Instruct-0708 on GitHub to install. At the same time, huggingface.co provides the effect of MOSS-VL-Instruct-0708 install, users can directly use MOSS-VL-Instruct-0708 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
MOSS-VL-Instruct-0708 install url in huggingface.co: