A Multimodal Autonomous Driving Model: Vision + Language + GPS + Speed.
🚀 Model Overview
DriveFusion-V0.2
is a multimodal model designed for autonomous vehicle applications. Unlike standard Vision-Language models, V0.2 integrates
telemetry data
(GPS and Speed) directly into the transformer architecture to perform dual tasks:
Natural Language Reasoning
: Describing scenes and explaining driving decisions.
Trajectory & Speed Prediction
: Outputting coordinates for future waypoints and target velocity profiles.
Built on the
Qwen2.5-VL
foundation, DriveFusion-V0.2 adds specialized MLP heads to fuse physical context with visual features, enabling a comprehensive "world model" for driving.
🔗 GitHub Repository
Find the full implementation, training scripts, and preprocessing logic here:
DriveFusion-V0.2 huggingface.co is an AI model on huggingface.co that provides DriveFusion-V0.2's model effect (), which can be used instantly with this DriveFusion DriveFusion-V0.2 model. huggingface.co supports a free trial of the DriveFusion-V0.2 model, and also provides paid use of the DriveFusion-V0.2. Support call DriveFusion-V0.2 model through api, including Node.js, Python, http.
DriveFusion-V0.2 huggingface.co is an online trial and call api platform, which integrates DriveFusion-V0.2's modeling effects, including api services, and provides a free online trial of DriveFusion-V0.2, you can try DriveFusion-V0.2 online for free by clicking the link below.
DriveFusion DriveFusion-V0.2 online free url in huggingface.co:
DriveFusion-V0.2 is an open source model from GitHub that offers a free installation service, and any user can find DriveFusion-V0.2 on GitHub to install. At the same time, huggingface.co provides the effect of DriveFusion-V0.2 install, users can directly use DriveFusion-V0.2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.