PixelDiT and PiD: 1.3B VAE-Free Pixel-Space DiT + PiD 4-Step Super Resolution by NVIDIA — nsclv1
ComfyUI Wiki
PixelDiT is NVIDIA's VAE-free 1.3B pixel-space DiT for text-to-image. PiD provides 4-step distilled super-resolution decoders up to 4K for multiple latent spaces.
P
PixelDiT and PiD
Text-to-ImageVAE-freeSuper Resolution1.3B4KDistilledPixelDiT (CVPR 2026) is a VAE-free pixel-space diffusion transformer with 1.3B parameters, featuring dual-level architecture (Patch-level DiT + Pixel-level DiT) and MM-DiT text-image fusion. PiD (Pixel Diffusion Decoder) reformulates latent-to-pixel decoding as a conditional pixel-space diffusion model, unifying decoding and upsampling into a single generative module with 4-step distilled inference. PiD supports multiple latent spaces including FLUX.1, FLUX.2, SD3, SDXL, and Qwen-Image.
| Developer | NVIDIA |
| Release Date | 2026 (CVPR 2026) |
| Architecture | Pixel-Space DiT (PixelDiT 1.3B) + Pixel Diffusion Decoder (PiD) |
| License | nsclv1 |
| Output Resolution | 1024px (PixelDiT), up to 4096px (PiD) |
| Inference Steps | 25-50 (PixelDiT), 4 (PiD distilled) |
| Text Encoder | Gemma-2-2B-IT |
Guides and workflows related to this model series.
No articles found.
Comments
Sign in with GitHub to join the discussion.