HuMo: Unified Human-Centric Video Generation by ByteDance Research

ComfyUI Wiki

HuMo is a unified human-centric video generation framework from ByteDance Research supporting text, image, and audio-driven human video generation.

H

HuMo

VideoHuman Video GenerationText-to-VideoAudio-DrivenByteDance

Unified human-centric video generation framework by ByteDance Research. Produces high-quality, controllable human videos from multimodal inputs including text, images, and audio. Supports Text-Image, Text-Audio, and Text-Image-Audio generation modes. Available in 1.7B and 17B parameter variants with FP8 quantization option.

DeveloperByteDance Research
Release Date2025-08
ArchitectureDiffusion Transformer
Model Sizes1.7B, 17B (FP16 + FP8)
LicenseApache-2.0
CapabilitiesText-Image, Text-Audio, Text-Image-Audio human video generation
Required Audio EncoderWhisper Large V3 (included)

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…