This paper presents
Omni-View
, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation facilitates understanding". Consisting of understanding model, texture module, and geometry module, Omni-View jointly models scene understanding, novel view synthesis, and geometry estimation, enabling synergistic interaction between 3D scene understanding and generation tasks. By design, it leverages the spatiotemporal modeling capabilities of its texture module responsible for appearance synthesis, alongside the explicit geometric constraints provided by its dedicated geometry module, thereby enriching the model’s holistic understanding of 3D scenes. Trained with a two-stage strategy, Omni-View achieves a state-of-the-art score of 55.4 on the VSI-Bench benchmark, outperforming existing specialized 3D understanding models, while simultaneously delivering strong performance in both novel view synthesis and 3D scene generation.
We provide the scripts for evaluating 3D scene understanding, Spatial Reasoning (VSI-bench), and Novel View Synthesis.
Please See
EVAL
for more details.
📊 Benchmarks
1. 3D Scene Understanding
2. VSI-Bench
3. Novel View Synthesis
✍️ Citation
If you find this work useful in your research, please consider citing:
@misc{hu2025omniview,
title={Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images},
author={JiaKui Hu and Shanshan Zhao and Qing-Guo Chen and Xuerui Qiu and Jialun Liu and Zhao Xu and Weihua Luo and Kaifu Zhang and Yanye Lu},
year={2025},
eprint={2511.07222},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.07222},
}
🧡 Acknowledgements
Our implementation is built upon
Bagel
. We appreciate their great work.
📄 License
Copyright (C) 2025 AIDC-AI
Licensed under the Apache License, Version 2.0.
This project contains various third-party components under other open source licenses. You should respect the terms of those licenses.
The component DiT is released under the CC-BY-NC 4.0 License (for non-commercial purposes only).
See the NOTICE file for more information.
Runs of ATH-MaaS Omni-View on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Omni-View huggingface.co Model
Omni-View huggingface.co is an AI model on huggingface.co that provides Omni-View's model effect (), which can be used instantly with this ATH-MaaS Omni-View model. huggingface.co supports a free trial of the Omni-View model, and also provides paid use of the Omni-View. Support call Omni-View model through api, including Node.js, Python, http.
Omni-View huggingface.co is an online trial and call api platform, which integrates Omni-View's modeling effects, including api services, and provides a free online trial of Omni-View, you can try Omni-View online for free by clicking the link below.
ATH-MaaS Omni-View online free url in huggingface.co:
Omni-View is an open source model from GitHub that offers a free installation service, and any user can find Omni-View on GitHub to install. At the same time, huggingface.co provides the effect of Omni-View install, users can directly use Omni-View installed effect in huggingface.co for debugging and trial. It also supports api for free installation.