- Home
- Models
- Dreamx Creator
- DreamX-Creator 1.0: ComfyUI Setup and Model Guide
DreamX-Creator 1.0: ComfyUI Setup and Model Guide
Run DreamX-Creator in ComfyUI with native nodes: 7B joint audio-video generation, a 1-step 2K refiner, the ~54 GB weight bundle and the generation workflow.
DreamX-Creator 1.0
Video GenerationJoint Audio-Video2KComfyUIDreamX-Creator 1.0 is AMap's open 7B generator that produces synchronised video and audio in a single pass, plus an autoregressive 1-step refiner that super-resolves the result to 2K. Native ComfyUI V3 nodes expose the released generator and refiner.
| Developer | AMap ML (Alibaba) |
| Release Date | 2026-09 |
| Architecture | 7B joint audio-video generator + 5B SR-DiT refiner |
| Base dependencies | Wan2.2-TI2V-5B (video VAE, UMT5-xxl text encoder) |
| Output | Synchronised video + audio, up to 2K after refinement |
| License | Apache-2.0 |
Given a first frame and a text prompt, DreamX-Creator jointly models modality-specialised video and audio streams through gated cross-modal attention: cross-attention weights let each modality attend to the other, with a learned gate controlling how much information flows across. Progressive joint training keeps lip movements, on-screen action and the soundtrack synchronised rather than stitched together after the fact. An autoregressive 1-step refiner (SR-DiT, 5B) then super-resolves the clip to 2K while preserving content, motion and audio-aligned timing, and can also refine external videos.
Installation
DreamX-Creator runs in ComfyUI through native V3 nodes from DreamX Creator T8. Install the node package from ComfyUI Manager (search "DreamX Creator T8"), or clone it into ComfyUI/custom_nodes/ and install its requirements with the same Python that starts ComfyUI.
Download the ComfyUI-ready weight bundle into the shared model directory:
hf download t8star/DreamX-Creator-Comfy --local-dir ComfyUI/models/dreamx_creatorThe loader's model_root=auto searches the repository's checkpoints/ and ComfyUI/models/dreamx_creator/. The bundle is about 54 GB and its root must directly contain the creator/, audio_vae/, refiner/ and wan2.2_ti2v_5b/ folders.
Workflows
The generation graph runs a complete loader, two text encoders, a first-frame audio-video latent, the AV flow shifts (the released preset is 5.0 / 5.0), a multimodal guider, and a FlowMatch sampler on normal scheduling, 20 steps and euler — followed by AV latent split, video and audio VAE decode, and video export. The separate refiner graph performs the 2× pass while preserving audio and FPS. Requested durations are floored to a Wan-compatible 4N+1 frame count.
The node package ships frontend-importable examples (examples/dreamx_creator_ui.json for first-frame to synchronised video and audio, and examples/dreamx_refiner_ui.json for the 2× refinement pass).
Resources
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.