ComfyUI v0.38.0: Ming-Image, ID-V2V and Faster SeedVR2
ComfyUI v0.38.0 adds native Ming-Image support with an official template, lands Wan ID-V2V, w6a8 quantization, HDR LogC3 and ACEScct color spaces, and speeds up SeedVR2 and YuE2.
A design composition generated by the official Ming-Image 0.1 Design template, which writes RGBA output with transparency (asset from the workflow_templates repository).
Ming-Image in Core
Ming-Image 0.1 Design now runs natively. Core support arrived in Comfy-Org/ComfyUI#16482 with a follow-up detection fix in #16514, after the same change had been carried in the v0.37.3 patch and reverted again in v0.37.4. If the model appeared and then disappeared for you on the stable branch, v0.38.0 is where it stays.
The port adds a dedicated conditioning node, Text Encode Ming Image Edit (TextEncodeMingImageEdit, category model/conditioning/ming image). It tokenizes the prompt together with up to eight reference images (image_1 to image_8) into the Ming text encoder, and the VAE input is optional: without it the images only condition through the vision tower, while with it each reference is appended to the latent sequence as a clean frame. Later references are resized to the first one, and the sampled latent should match that first image's size.
On the model side, Ming-Image is a 30-layer DiT over a 16-channel latent space with its own MingImage latent format, and it runs with a 3.16 shift at the 1024 bucket. The text encoder is the bundled Ling-Mini-2.0 language model, a 20-layer MoE with 256 experts, paired with a Qwen-style vision tower for reference images. Two checkpoints work with the same node graph: the Design model and the Design Layer sibling that decomposes a finished design into editable transparent layers.
The repacked weights are published as Comfy-Org/Ming-Image in bf16 and int8_convrot variants:
π ComfyUI/
βββ π models/
βββ π diffusion_models/
β βββ ming_image_0.1_design_bf16.safetensors
β βββ ming_image_0.1_design_int8_convrot.safetensors
β βββ ming_image_0.1_design_layer_bf16.safetensors
β βββ ming_image_0.1_design_layer_int8_convrot.safetensors
βββ π text_encoders/
β βββ ming_image_0.1_ling_mini_2.0_bf16.safetensors
β βββ ming_image_0.1_ling_mini_2.0_int8_convrot.safetensors
βββ π vae/
βββ ming_image_vae_bf16.safetensorsThe official template image_ming_image_01_design_t2i wires the whole pipeline, including a prompt enhancement stage: a Qwen3.8-27B text encoder (shipped as qwen3.8_27b_w4a8.safetensors) rewrites a short request into a structured Figma-style JSON design brief, and that brief is what the Ming text encoder receives. Sampling is 12 steps of Euler with ModelSamplingFlux at a 1024x2048 latent.
Qwen-Image 2.1 and ControlNet Updates
The Qwen-Image 2.1 path keeps getting work:
- Union Fun ControlNet is now supported (#16519) for alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union, the single ControlNet that covers the individual Fun ControlNet types. The matching H3 Fun-ControlNet-Union 2.0 support from #16471 is in this release as well.
- Tiny VAE support (#16552) adds
taeqi2_1_decoderandtaeqi2_1_encoderfrom madebyollin's taesd, so Qwen-Image 2.1 gets fast latent previews and a cheap TAESD round trip. Drop the files inmodels/vae_approx/. These are 16x models with 64 latent channels and a 4-channel RGBA image side, with per-channel scale and shift instead of the scalar ones older TAESD variants used. - Prefix KV cache placement (#16429) is now decided with the model patcher's free-memory accounting, so the cache can stay in VRAM under dynamic VRAM and smart memory instead of falling back to a full recompute. Transformer block compilation (#16430) adds memory-compiler support, which mainly helps when offload bandwidth is the bottleneck.
- A fp16 activation clamp (#16608) fixes an in-place allocation that broke the memory compiler, and Qwen VL image preprocessing no longer crashes on 4-channel RGBA input (#16548).
Wan ID-V2V Arrives
The ID-V2V support PR, Comfy-Org/ComfyUI#15139, finally merged on 2026-09-27 and ships here. ID-V2V restyles video while preserving identity, and it is built on a Wan 2.1 image-to-video base with VACE control, a combination ComfyUI could not previously load.
Two changes make it work. Model detection recognises the architecture when a VACE model carries the img_emb projection of an I2V base, and Image to Video now takes an optional ref_pad_image input. That input fills the padding frames of the image conditioning with a reference image instead of flat gray, which is the anti-drift padding the model was trained with. ref_pad_image is optional, so existing I2V workflows are unaffected.
Test weights stay on Hugging Face: Kijai/Wan_ID_V2V_comfy publishes wan_2.1_idv2v_int8_convrot.safetensors and a variant with normal depth control, and the PR includes a ready workflow JSON.
Quantization: w6a8
ComfyUI can now load w6a8 checkpoints (#16483, implemented in comfy-kitchen): 6-bit weights with 8-bit activations, so the file is smaller than both the int8 and fp8 repacks. The pruned H3 weights are already published in the new format on Comfy-Org/MiniMax-H3 as minimax_h3_fl2va_pruned_w6a8.safetensors and minimax_h3_ref2va_pruned_w6a8.safetensors, next to the existing bf16, fp8_scaled and int8_convrot files.
Faster SeedVR2 and YuE2
SeedVR2 (#16530) is a substantial rewrite of the upscaler's VAE. Where the model allows it, the VAE now processes video one frame at a time, keeps its causal convolution caches int8-quantized in pinned host memory, and prefetches them back as needed. VRAM use is therefore set by a single frame's working set instead of the whole clip, which matters for long or high-resolution video. It needs the comfy-kitchen kernels from the same release.
Output of the official SeedVR2 image upscale template (asset from the workflow_templates repository).
Two audio paths also got faster:
- YuE2 (#16626) now batches the token transfers from the GPU back to the CPU and runs its sampler under CUDA graphs.
- ACE-Step 1.5 (#16461) gets CUDA graphs and the memory compiler on its autoregressive model.
Two MiniMax H3 fixes ride along: the VAE now blends tiled decodes against already composited neighbours (#16436), and an alignment bug that crashed the H3 VAE when qk_norm_scale was offloaded is fixed (#16485). Upscale Image (Using Model) no longer crashes on RGBA images (#16500).
HDR Color Spaces
Convert Image Color Space (ImageColorSpace) gained two log encodings (#16541): HDR LogC3 (the EI 800 curve with Rec.709 primaries) and HDR ACEScct (AP1 primaries and D60 white, adapted to D65 with Bradford chromatic adaptation). Combined with the existing sRGB, linear Rec.709, HDR HLG and HDR PQ targets, a workflow can now move between the color spaces a VFX or EXR pipeline expects. A follow-up fix (#16565) handles negative values in the HDR output path.
Memory, Attention and Loaders
- Model files can now declare which attention implementation each block should use (#16419), which lets a repack mix attention backends inside one model.
- comfy_attention and
AttentionTensorContainersupport extends to the Lumina family (#16515), a few more model families (#16595) and, with lower memory use, the Flux family (#16488). - The fast disk detection added for dynamic loading in v0.37.0 now covers all model loaders (#16425), and AMD RDNA2 and older architectures were added to the arch list (#16537).
- Ascend NPU devices get asynchronous weight offload streams (#16057).
- The asset system got a round of hardening and performance work: the SQLite write lock is taken up front and file reads moved out of write transactions (#16480, #16486), output rescans list folders instead of checking every file and pause their loops so the UI stays responsive (#16543, #16546), and ComfyUI starts even when the asset packages are missing, reporting what
--enable-assetsneeds (#16580).
Other Changes
- Text Generate takes an optional
system_promptinput and returns the thinking block on its own output (#16442). - The torchaudio dependency was removed, as upstream has stopped shipping stable cu132 builds (#16457).
- MiniMax H3 Fun-ControlNet-Union 2.0 support from #16471 is included, and SheetSage2 ABC creation was aligned with upstream YuE code (#16569).
- Partner nodes: Hunyuan Image 3.5 text-to-image and edit (Tencent, #16462), Arrow 2 with reasoning effort in the SVG nodes (Quiver, #16478), Claude Opus 5.5 (#16479), Recraft V4.1 Flash (#16501), GPT-6 Sol and Luna (OpenAI, #16520), Seedream 5.0 Flash (#16522) and Seedance 2.5 Draft mode (#16529). The deprecated Sora nodes were removed (#16609). Several of these partner additions already shipped in the v0.37.x patch releases.
- workflow templates are at v0.11.70, comfy-kitchen at 0.2.36 and embedded docs at 0.5.13.
Getting v0.38.0
Update through the ComfyUI Manager or your launcher. The Qwen-Image 2.1 ControlNet and tiny VAE support, the new color spaces and ID-V2V all need v0.38.0 or newer. Ming-Image is the one to watch on the patch line: the nodes were present in v0.37.3 and reverted again in v0.37.4, so v0.38.0 is where the support stays.
Comments
Sign in with GitHub to join the discussion.