HuMo: Unified Human-Centric Video Generation by ByteDance Research
ComfyUI Wiki
HuMo is a unified human-centric video generation framework from ByteDance Research supporting text, image, and audio-driven human video generation.
H
HuMo
VideoHuman Video GenerationText-to-VideoAudio-DrivenByteDanceUnified human-centric video generation framework by ByteDance Research. Produces high-quality, controllable human videos from multimodal inputs including text, images, and audio. Supports Text-Image, Text-Audio, and Text-Image-Audio generation modes. Available in 1.7B and 17B parameter variants with FP8 quantization option.
| Developer | ByteDance Research |
| Release Date | 2025-08 |
| Architecture | Diffusion Transformer |
| Model Sizes | 1.7B, 17B (FP16 + FP8) |
| License | Apache-2.0 |
| Capabilities | Text-Image, Text-Audio, Text-Image-Audio human video generation |
| Required Audio Encoder | Whisper Large V3 (included) |
Guides and workflows related to this model series.
No articles found.
Comments
Sign in with GitHub to join the discussion.