DreamX-Creator 1.0: ComfyUI Setup and Model Guide

ComfyUI Wiki

Run DreamX-Creator in ComfyUI with native nodes: 7B joint audio-video generation, a 1-step 2K refiner, the ~54 GB weight bundle and the generation workflow.

D

DreamX-Creator 1.0

Video GenerationJoint Audio-Video2KComfyUI

DreamX-Creator 1.0 is AMap's open 7B generator that produces synchronised video and audio in a single pass, plus an autoregressive 1-step refiner that super-resolves the result to 2K. Native ComfyUI V3 nodes expose the released generator and refiner.

DeveloperAMap ML (Alibaba)
Release Date2026-09
Architecture7B joint audio-video generator + 5B SR-DiT refiner
Base dependenciesWan2.2-TI2V-5B (video VAE, UMT5-xxl text encoder)
OutputSynchronised video + audio, up to 2K after refinement
LicenseApache-2.0

Given a first frame and a text prompt, DreamX-Creator jointly models modality-specialised video and audio streams through gated cross-modal attention: cross-attention weights let each modality attend to the other, with a learned gate controlling how much information flows across. Progressive joint training keeps lip movements, on-screen action and the soundtrack synchronised rather than stitched together after the fact. An autoregressive 1-step refiner (SR-DiT, 5B) then super-resolves the clip to 2K while preserving content, motion and audio-aligned timing, and can also refine external videos.

Installation

DreamX-Creator runs in ComfyUI through native V3 nodes from DreamX Creator T8. Install the node package from ComfyUI Manager (search "DreamX Creator T8"), or clone it into ComfyUI/custom_nodes/ and install its requirements with the same Python that starts ComfyUI.

Download the ComfyUI-ready weight bundle into the shared model directory:

hf download t8star/DreamX-Creator-Comfy --local-dir ComfyUI/models/dreamx_creator

The loader's model_root=auto searches the repository's checkpoints/ and ComfyUI/models/dreamx_creator/. The bundle is about 54 GB and its root must directly contain the creator/, audio_vae/, refiner/ and wan2.2_ti2v_5b/ folders.

Workflows

The generation graph runs a complete loader, two text encoders, a first-frame audio-video latent, the AV flow shifts (the released preset is 5.0 / 5.0), a multimodal guider, and a FlowMatch sampler on normal scheduling, 20 steps and euler — followed by AV latent split, video and audio VAE decode, and video export. The separate refiner graph performs the 2× pass while preserving audio and FPS. Requested durations are floored to a Wan-compatible 4N+1 frame count.

The node package ships frontend-importable examples (examples/dreamx_creator_ui.json for first-frame to synchronised video and audio, and examples/dreamx_refiner_ui.json for the 2× refinement pass).

Resources

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…