Simplifying Traffic Anomaly Detection with Video Foundation Models
Svetlana Orlova, Tommie Kerssies, Brun´o B. Englert, Gijs Dubbelman
Eindhoven University of Technology
Recent methods for ego-centric Traffic Anomaly Detection (TAD) often rely on complex multi-stage or multi-representation fusion architectures, yet it remains unclear whether such complexity is necessary. Recent findings in visual perception suggest that foundation models, enabled by advanced pre-training, allow simple yet flexible architectures to outperform specialized designs. Therefore, in this work, we investigate an architecturally simple encoder-only approach using plain Video Vision Transformers (Video ViTs) and study how pre-training enables strong TAD performance. We find that: (i) advanced pre-training enables simple encoder-only models to match or even surpass the performance of specialized state-of-the-art TAD methods, while also being significantly more efficient; (ii) although weakly- and fully-supervised pre-training are advantageous on standard benchmarks, we find them less effective for TAD. Instead, self-supervised Masked Video Modeling (MVM) provides the strongest signal; and (iii) Domain-Adaptive Pre-Training (DAPT) on unlabeled driving videos further improves downstream performance, without requiring anomalous examples. Our findings highlight the importance of pre-training and show that effective, efficient, and scalable TAD models can be built with minimal architectural complexity.
✨ DoTA and DADA-2000 results
Video ViT-based encoder-only models set a new state of the art
on both datasets, while being significantly more efficient than top-performing specialized methods. FPS measured using NVIDIA A100
MIG, 2 1 GPU. † From prior work. ‡ Optimistic estimates using publicly available components of the model. “A→B”: trained on A, tested
on B; D2K: DADA-2000.
If you think this project is helpful, please feel free to like us ❤️ and cite our paper:
@inproceedings{orlova2025simplifying,
title={Simplifying Traffic Anomaly Detection with Video Foundation Models},
author={Orlova, Svetlana and Kerssies, Tommie and Englert, Brun{\'o} B and Dubbelman, Gijs},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2025}
}
@article{orlova2025simplifying,
title={Simplifying Traffic Anomaly Detection with Video Foundation Models},
author={Orlova, Svetlana and Kerssies, Tommie and Englert, Brun{\'o} B and Dubbelman, Gijs},
journal={arXiv preprint arXiv:2507.09338},
year={2025}
}
Runs of tue-mps simple-tad on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About simple-tad huggingface.co Model
simple-tad huggingface.co is an AI model on huggingface.co that provides simple-tad's model effect (), which can be used instantly with this tue-mps simple-tad model. huggingface.co supports a free trial of the simple-tad model, and also provides paid use of the simple-tad. Support call simple-tad model through api, including Node.js, Python, http.
simple-tad huggingface.co is an online trial and call api platform, which integrates simple-tad's modeling effects, including api services, and provides a free online trial of simple-tad, you can try simple-tad online for free by clicking the link below.
tue-mps simple-tad online free url in huggingface.co:
simple-tad is an open source model from GitHub that offers a free installation service, and any user can find simple-tad on GitHub to install. At the same time, huggingface.co provides the effect of simple-tad install, users can directly use simple-tad installed effect in huggingface.co for debugging and trial. It also supports api for free installation.