Skip to main content
ToolPotion

daVinci-MagiHuman

daVinci-MagiHuman is an open-source AI model that turns a single portrait photo plus a script or audio into a lip-synced talking video, generating aligned audio and video together in one pass. It can be tried free online or self-hosted.

daVinci-MagiHuman screenshot

Description

daVinci-MagiHuman is a 15-billion-parameter open-source AI model that generates lip-synced talking videos from a single photo. You upload a portrait, add a script or audio, and get a natural talking clip where the audio and video are generated together rather than stitched from separate pipelines. It is developed by Sand.ai and GAIR Lab (Shanghai Jiao Tong University) and released under the Apache 2.0 license, so the weights can be inspected, run locally, and used commercially within the license.

The model uses a single-stream Transformer that jointly denoises video and audio tokens with a reference-image latent, producing unified audio-video output from a face photo plus text or audio. On a single NVIDIA H100 GPU it can generate a short 256p clip in about two seconds of wall time, and published evaluations report strong word-error rates and high human preference against baselines such as Ovi 1.1 and LTX 2.3.

Users can try it through a free online demo, download the checkpoints from Hugging Face, or clone the GitHub repository to self-host and run inference with custom settings and resolutions up to 1080p. It supports multiple languages for lip sync depending on the released training data.

The hosted service offers a free tier with starter and daily check-in credits (15-day temporary storage) plus Basic, Pro, and Max credit plans that add HD generation, priority processing, permanent asset storage, and commercial usage rights.

daVinci-MagiHuman's Core Features

  • Single-photo talking video generation with lip sync

  • Unified audio and video generated together in one model pass

  • 15B-parameter single-stream Transformer architecture

  • Open source under the Apache 2.0 license

  • Fast inference (~2s for a ~2s 256p clip on an H100)

  • Multilingual lip sync depending on training data

  • Self-hosting via Hugging Face and GitHub, or a hosted online demo

  • Output resolutions up to 1080p

How to use daVinci-MagiHuman?

  1. Upload a portrait: Add a clear, front-facing face photo.

  2. Add a script or audio: Enter text or upload an audio file to drive the speech.

  3. Choose resolution: Select an output resolution such as 256p, 720p, or 1080p.

  4. Generate: Run the model to create the lip-synced clip.

  5. Download: Save the finished talking video, or self-host from Hugging Face or GitHub.

daVinci-MagiHuman's Use Cases

  • Talking avatar videos
  • Self-hosted generation
  • Multilingual dubbing
  • Research and benchmarking
  • Content creation

FAQ from daVinci-MagiHuman

daVinci-MagiHuman Reviews

Loading...

Popular AI Tools Like daVinci-MagiHuman

AI Apps

An AI video generation company behind the Magi family of models, including MAGI-2 Preview, a unified audio-video model that generates synchronized dialogue, singing, and cinematic…

AI Video GeneratorsMedia & Entertainment

AI Apps

Digen AI is a free AI video generator that converts images into professional videos with realistic lip-sync, multilingual support, and smart motion. It is aimed at creators who…

AI Video GeneratorsMedia & Entertainment

AI Apps

Gaga AI is an online AI video generator and avatar creator from Sand.ai that turns a single image and audio into cinematic talking-head videos with precise lip sync, natural…

AI Video Generators

Seedance 2.0 is a unified multimodal AI video generator that turns text, images, audio, or video references into cinematic 1080p clips with native lip-sync, physics-accurate…

AI Video GeneratorsMedia & Entertainment

AI Apps

ClipDance is an AI video creation platform that turns text and images into cinematic 1080p videos with native synchronized audio, powered by Seedance 2.0 and other leading models.…

AI Video GeneratorsMedia & Entertainment

LumeFlow AI is an all-in-one AI video and image platform that turns text, images, or videos into high-quality clips using 20+ models, with lip-sync, effects, and image generation.…

AI Video GeneratorsMarketing & Creative Agencies

AI Apps

Monet AI is an all-in-one visual creation platform that unifies 20+ leading video, image, and audio models in one account, letting creators generate and compare AI content without…

AI Video GeneratorsMarketing & Creative Agencies

AI Apps

LipsyncX is an AI lip-sync video generator that turns photos, videos, scripts, and voices into talking photos, dubbed clips, singing videos, and multilingual avatar videos, with…

AI Video GeneratorsMedia & Entertainment