PDMD: 4-Step MiniMax H3 Distillation for ComfyUI
UC San Diego and ByteDance release PDMD, a one-line change to DMD that distills MiniMax H3 to four steps, with a Kijai ComfyUI conversion.
![]() | ![]() |
|---|---|
| DMD target | PDMD target |
Both panels decode a real training update target from the same step of the MiniMax H3 run. The project page presents the PDMD target as carrying fewer artifacts, which is what the projection is meant to produce.
What PDMD changes
DMD trains a student on the difference between the critic's score and the student's own score. That update also carries the critic's error, and the error enters every successive student update and accumulates, which shows up as progressive oversaturation and unnatural texture. PDMD keeps the DMD update and subtracts the part of it that points along the student and critic endpoint residual:
r = x0_critic - x0_student
d = d - (d * r).sum() / r.pow(2).sum() * rThe paper shows that at a fixed noisy query this residual is an unbiased estimate of the critic's endpoint error, and that under high-dimensional assumptions the projection removes a constant fraction of the critic error while discarding only a vanishing fraction of the ideal DMD signal. It is a one-line change to DMD with no auxiliary loss, no extra network, no extra model pass and no extra training stage.
The intuition figure from the project page. DMD moves the student endpoint along the update direction; PDMD projects out the component parallel to the endpoint residual.
What was released
The release is the MiniMax H3 side of the method, as three artifacts:
| Artifact | Description |
|---|---|
pdmd_2NFE_lora | 2-step LoRA, published 2026-09-23 |
pdmd_4NFE_full | 4-step full bf16 transformer (MiniMaxH3Transformer3DModel), published 2026-09-27 |
pdmd_4NFE_lora | 4-step LoRA, published 2026-09-29 |
The adapters carry rank 128 and alpha 128 over 312 modules, 624 tensors, of the H3 transformer, in the Diffusers key layout (transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight). Everything else, the VAE, audio VAE, schedulers, text encoder and processor, comes from the base model, and the safetensors metadata records the fusing rule W_base += lora_scale * (lora_B @ lora_A) with lora_scale = 1.0. The card asks for four denoising steps with the base model's released scheduler configuration (shift 12 / 3). The 2-step repository also ships an fp32 copy of its adapter.
Results
The paper reports the method on MiniMax H3 and Wan2.1. On MiniMax H3 it uses VideoGen-Eval with 387 prompts at 544p and a fixed seed:
| Method | NFE | Total | Quality | Dynamic | Semantic | PQ | CE | CU | IS | |--|--:|--:|--:|--:|--:|--:|--:|--:| | MiniMax-H3-33B, teacher | 50 | 82.41 | 82.22 | 66.67 | 83.20 | 6.567 | 4.188 | 6.213 | 5.15 | | MiniMax-H3-33B | 4 | 79.48 | 79.54 | 44.44 | 79.23 | 6.056 | 3.167 | 5.580 | 3.35 | | H3 Turbo LoRA | 4 | 81.57 | 81.29 | 54.78 | 82.69 | 6.406 | 3.917 | 6.015 | 4.52 | | DMD2† | 4 | 82.27 | 82.51 | 59.95 | 81.34 | 6.305 | 3.434 | 5.871 | 4.06 | | DMD† | 4 | 82.76 | 82.68 | 61.76 | 83.05 | 6.381 | 3.801 | 5.931 | 4.79 | | rCM† | 4 | 81.18 | 80.98 | 58.40 | 81.95 | 6.063 | 3.711 | 5.446 | 3.69 | | AnyFlow† | 4 | 81.97 | 81.82 | 64.60 | 82.60 | 6.092 | 3.544 | 5.591 | 4.35 | | PDMD | 4 | 83.17 | 83.25 | 71.83 | 82.86 | 6.530 | 4.062 | 6.180 | 4.98 |
† marks the authors' reimplementations, trained under a shared protocol.
On Wan2.1-T2V-1.3B it uses VBench with 944 prompts at 480p and five seeds each:
| Method | NFE | Total | Quality | Dynamic | Semantic |
|---|---|---|---|---|---|
| Wan2.1-T2V-1.3B, teacher | 50x2 | 83.06 | 85.02 | 68.61 | 75.23 |
| Wan2.1-T2V-1.3B | 4x2 | 68.54 | 73.60 | 18.61 | 48.30 |
| rCM | 4 | 83.13 | 85.14 | 78.06 | 75.10 |
| AnyFlow | 4 | 83.54 | 85.28 | 59.44 | 76.57 |
| DMD† | 4 | 82.70 | 85.06 | 86.39 | 73.24 |
| DMD2† | 4 | 83.44 | 85.84 | 76.67 | 73.84 |
| PDMD | 4 | 83.73 | 85.89 | 89.72 | 75.09 |
The abstract states that on MiniMax H3 joint video and audio generation the 4-step PDMD student scores 83.17 on the VideoGen-Eval visual total, 0.41 points above the strongest distilled baseline, and takes the best result on all six audio metrics among the compared four-step models. On Wan2.1 it reports a VBench total of 83.73 at four steps, 1.03 points above the matched DMD run. The project page adds qualitative comparisons and a user study that favor PDMD over the distilled baselines in visual quality, motion and audio quality.
Running it in ComfyUI
The official release targets Diffusers, not nodes: inference.py and run_a10.py sample with the re-noise rule the students were trained under, and no workflow JSON ships with the release.
For ComfyUI, the adapters are ordinary LoRAs once converted. Kijai has published rank-reduced conversions on Kijai/MiniMax-H3-experimental: minimax_h3_pdmd_2step_lora_avg_rank_38_bf16.safetensors and minimax_h3_pdmd_4step_lora_avg_rank_57_bf16.safetensors. They load through the standard MiniMax H3 LoRA path without a custom node, with the usual caveat that the rank reduction and the sampler settings are not the ones the paper evaluated. A community conversion, Iwannapose/minimax_h3_pdmd_4nfe_comfyui, is also circulating.
The official PDMD trailer. Every frame on the project page runs at four network evaluations or fewer.
Two clips from the project page's MiniMax H3 comparison grid, matched prompt and seed, four steps each:
DMD at 4 steps.
PDMD at 4 steps, same prompt and seed as the clip above.
Limits
- The published MiniMax H3 students target 1344x768 with the base model's text encoder and audio components, so the release is a sampling-speed result rather than a smaller model.
- The ComfyUI conversions reduce the rank 128 adapters to an average rank of 38 or 57, so they are not the exact files the paper evaluated.
- With no official nodes or workflow JSON, ComfyUI support depends on community conversions and their key layout. The release itself is Diffusers-first.
- Several baselines marked with a dagger in the tables above are the authors' own reimplementations trained under a shared protocol, not the numbers those projects publish.
Availability
Project page: pdmd2026.github.io
Paper: arXiv 2609.35768
Code: ZeamoxWang/pdmd
Weights: pdmd2026
ComfyUI conversion: Kijai/MiniMax-H3-experimental
Base model: MiniMaxAI/MiniMax-H3


Comments
Sign in with GitHub to join the discussion.