PixelDiT and PiD: 1.3B VAE-Free Pixel-Space DiT + PiD 4-Step Super Resolution by NVIDIA — nsclv1

ComfyUI Wiki

PixelDiT is NVIDIA's VAE-free 1.3B pixel-space DiT for text-to-image. PiD provides 4-step distilled super-resolution decoders up to 4K for multiple latent spaces.

P

PixelDiT and PiD

Text-to-ImageVAE-freeSuper Resolution1.3B4KDistilled

PixelDiT (CVPR 2026) is a VAE-free pixel-space diffusion transformer with 1.3B parameters, featuring dual-level architecture (Patch-level DiT + Pixel-level DiT) and MM-DiT text-image fusion. PiD (Pixel Diffusion Decoder) reformulates latent-to-pixel decoding as a conditional pixel-space diffusion model, unifying decoding and upsampling into a single generative module with 4-step distilled inference. PiD supports multiple latent spaces including FLUX.1, FLUX.2, SD3, SDXL, and Qwen-Image.

DeveloperNVIDIA
Release Date2026 (CVPR 2026)
ArchitecturePixel-Space DiT (PixelDiT 1.3B) + Pixel Diffusion Decoder (PiD)
Licensensclv1
Output Resolution1024px (PixelDiT), up to 4096px (PiD)
Inference Steps25-50 (PixelDiT), 4 (PiD distilled)
Text EncoderGemma-2-2B-IT

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…