









daVinci-MagiHuman is an advanced, open-source 15B-parameter AI model developed by Sand.ai and GAIR Lab at Shanghai Jiao Tong University. It is designed to generate high-quality, lip-synced talking videos from a single portrait image and a script or audio file. Unlike traditional methods that combine separate text-to-speech and video pipelines, daVinci-MagiHuman utilizes a unified single-stream Transformer to jointly denoise video and audio tokens simultaneously. Released under the Apache 2.0 license, it allows users to inspect weights, run inference locally, and use the technology for commercial purposes. It is optimized for speed, capable of generating short clips in just seconds on professional-grade hardware like the NVIDIA H100.
To use daVinci-MagiHuman, upload a clear, front-facing portrait photo and provide a script or audio file. Select your desired output resolution (e.g., 256p, 720p, or 1080p) and start the generation process. Once the AI completes the job, you can download your talking video. For local deployment, users can download the model checkpoints from Hugging Face and follow the provided CLI instructions.
Here is the DaVinci MagiHuman support email for customer service: [email protected] . More Contact, visit the contact us page()
DaVinci MagiHuman Company name: Sand.ai and GAIR Lab (Shanghai Jiao Tong University) .
DaVinci MagiHuman Company address: .
More about DaVinci MagiHuman, Please visit the about us page().
DaVinci MagiHuman Github Link: https://github.com/GAIR-NLP/daVinci-MagiHuman

Basic
$19.90/month
1,990 credits (approx. 16 standard or 9 HD video generations/month)
Pro
$31.92/month
3,990 credits, priority processing, batch background removal, and 20% off discount
Max
$47.92/month
5,990 credits, highest priority, dedicated support, and lifetime usage rights
Pay-as-you-go
$1 per 100 credits
Credits never expire, used for one-time top-ups


32.12%
10.59%
Social Listening