Description
daVinci-MagiHuman is a 15-billion-parameter open-source AI model that generates lip-synced talking videos from a single photo. You upload a portrait, add a script or audio, and get a natural talking clip where the audio and video are generated together rather than stitched from separate pipelines. It is developed by Sand.ai and GAIR Lab (Shanghai Jiao Tong University) and released under the Apache 2.0 license, so the weights can be inspected, run locally, and used commercially within the license.
The model uses a single-stream Transformer that jointly denoises video and audio tokens with a reference-image latent, producing unified audio-video output from a face photo plus text or audio. On a single NVIDIA H100 GPU it can generate a short 256p clip in about two seconds of wall time, and published evaluations report strong word-error rates and high human preference against baselines such as Ovi 1.1 and LTX 2.3.
Users can try it through a free online demo, download the checkpoints from Hugging Face, or clone the GitHub repository to self-host and run inference with custom settings and resolutions up to 1080p. It supports multiple languages for lip sync depending on the released training data.
The hosted service offers a free tier with starter and daily check-in credits (15-day temporary storage) plus Basic, Pro, and Max credit plans that add HD generation, priority processing, permanent asset storage, and commercial usage rights.
daVinci-MagiHuman's Core Features
Single-photo talking video generation with lip sync
Unified audio and video generated together in one model pass
15B-parameter single-stream Transformer architecture
Open source under the Apache 2.0 license
Fast inference (~2s for a ~2s 256p clip on an H100)
Multilingual lip sync depending on training data
Self-hosting via Hugging Face and GitHub, or a hosted online demo
Output resolutions up to 1080p
How to use daVinci-MagiHuman?
Upload a portrait: Add a clear, front-facing face photo.
Add a script or audio: Enter text or upload an audio file to drive the speech.
Choose resolution: Select an output resolution such as 256p, 720p, or 1080p.
Generate: Run the model to create the lip-synced clip.
Download: Save the finished talking video, or self-host from Hugging Face or GitHub.
daVinci-MagiHuman's Use Cases
- Talking avatar videos
- Self-hosted generation
- Multilingual dubbing
- Research and benchmarking
- Content creation






