DMAD: 4-Step MiniMax H3 Audio-Video LoRA for ComfyUI

ComfyUI Wikinews

ByteDance and Texas A&M release DMAD, a rank-128 LoRA that cuts MiniMax H3 to four steps for joint audio-video generation, with a Kijai ComfyUI conversion.

DMAD is a distillation method from Texas A&M University and ByteDance that recasts distribution matching as a classification problem. Its first release is a pair of rank-128 LoRAs that turn the 50-step MiniMax H3 teacher into a 4-step generator of 1344x768 video with native stereo audio. The weights, the inference code and the paper (arXiv 2610.02188) landed on 2026-10-02.
50-step MiniMax H3 teacherDMAD 4-step student
The 50-step MiniMax H3 teacherThe same prompt as a 4-step DMAD student

Matched frames from the project page's MiniMax H3 comparison grid. The student keeps the composition and the subject while running in four model evaluations instead of fifty.

What DMAD changes

Few-step distillation for diffusion normally follows DMD: train the student on the difference between the teacher's and the student's scores, which means keeping a second diffusion model fitted to the student's evolving distribution and paying for it in memory and compute. DMAD removes that auxiliary model.

The method recasts distribution matching as classification. Two discriminator heads share one backbone and are trained to separate real and teacher samples from the student's. Linear losses on the discriminator logits then train the student directly, with no score fitting. The paper shows that at the discriminator optimum those losses recover the distribution-matching gradient that DMD relies on, using the identity that links discriminator logits to log-density ratios.

A second component, gap-based reweighting, sets how strongly the teacher supervises each noise level. It reads the real-data head's empirical logit gap between real and teacher samples and adapts the weight from it, instead of using a fixed schedule.

What was released

The release is the MiniMax H3 side of the method. Two rank-128 LoRAs sit on the H3 text-to-audio-video transformer, each about 1.4 GB.

FileDescription
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensorsThe checkpoint of the paper: an EMA of the student at iteration 800 of the main run
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensorsA variant whose critic backbone is fully trained rather than frozen under a LoRA; the card reports a higher AVGen-Bench score

Both adapters carry rank 128 and alpha 128 over the attention projections and the two feed-forward layers of all 50 transformer blocks and the 2 token-refiner blocks, 312 modules in total, in the Diffusers key layout (<module>.lora.down.weight, <module>.lora.up.weight). The safetensors metadata records the rank, alpha and fusion rule.

Inference is text-to-audio-video at four steps, time shift 12 for video and 2 for audio, with no classifier-free guidance because MiniMax H3 is already guidance-distilled. The default output is 1344x768, 124 frames (about 5.2 s at 24 fps) with 32 kHz stereo audio, the setting the students were trained at.

Results

The paper reports the method on three backbones. Only the MiniMax H3 students are published with this release; the SDXL, EDM and Wan2.1 numbers come from the paper's own runs.

BackboneSettingScore
SDXL4 steps, COCO-10K14.47 FID
Wan2.1-T2V-14B4 steps85.15 VBench total
MiniMax-H3-33B4 steps, joint audio-video79.1% overall human preference over DMD2, 84.6% over rCM (ties excluded)

The abstract also reports 1.04 FID for one-step generation on ImageNet-64x64, using a projected discriminator, and states that the few-step students beat the compared few-step methods and the multi-step teachers on those metrics.

Running it in ComfyUI

The official release targets Diffusers rather than nodes: inference.py samples with the re-noise step rule the students were trained under, and run_diffusers_pipeline.py drops them into the official MiniMaxH3ModularPipeline. The base model is the 33B H3 transformer plus its Qwen3-VL-32B text encoder, about 170 GB of components, and the reference recommends one 80 GB GPU with CPU offload while idle.

For ComfyUI, the adapters are ordinary LoRAs once converted. Kijai has published a rank-reduced conversion, minimax_h3_DMAD_4step_full_lora_avg_rank_39_bf16.safetensors on Kijai/MiniMax-H3-experimental, so it loads through the standard H3 LoRA path without a custom node. Community conversions of the original diffusers file are also circulating, with the usual caveat that the two checkpoints and the sampler settings have to match the card.

An otter on a surfboard, generated by the 4-step MiniMax H3 student. Video and audio come from the same four-step pass. Full demo reel on YouTube.

50-step MiniMax H3 teacherDMAD 4-step student
Teacher, 50 stepsDMAD student, 4 steps

A second matched pair from the project page's MiniMax H3 grid.

Limits

  • The published students are trained for 1344x768 at 124 frames with 32 kHz stereo audio, and the card presents that as the operating point rather than one setting among many.
  • The H3 preference figures exclude ties, and the comparisons mix methods the release does not ship, so the head-to-head numbers for SDXL and Wan2.1 are the paper's runs, not the public files.
  • With no official nodes or workflow JSON in the release, ComfyUI support depends on community conversions, and their rank reduction and key layout are not the ones the paper evaluated.
  • The full H3 stack is large (33B transformer plus a 32B text encoder), so four steps cut sampling cost rather than the memory needed to hold the model.

Availability

Project page: yzmblog.github.io/projects/DMAD
Paper: arXiv 2610.02188
Code: Yzmblog/DMAD
Weights: ZhengmingYu/DMAD
Demo video: YouTube
ComfyUI conversion: Kijai/MiniMax-H3-experimental
Base model: MiniMaxAI/MiniMax-H3

Comments

Sign in with GitHub to join the discussion.

Loading comments…
DMAD: 4-Step MiniMax H3 Audio-Video LoRA for ComfyUI | ComfyUI Wiki