UltraTex: 2K Multi-View Diffusion for 3D Texturing in ComfyUI

ComfyUI Wikinews

UltraTex from Jilin University and VAST textures 3D meshes at 2048x2048 with background token dropping and block-sparse attention, plus a ComfyUI node pack.

UltraTex is a multi-view diffusion framework for texturing 3D meshes at 2048×2048, from Jilin University, the Shanghai Innovation Institute and collaborators at HKU, Tongji, Fudan and VAST. It keeps the detail of a dense 2K model while cutting training cost by up to 91×, and it has been wrapped for ComfyUI by a community node pack that runs on stock ComfyUI loaders.
UltraTex teaser

Textured assets produced by UltraTex from an untextured mesh plus a reference image.

What UltraTex does

Give UltraTex an untextured mesh and one reference image and it generates six 2048² texture views (front, left, back, right, top, bottom), then bakes them into a UV texture on the mesh. The point of the paper is that density at 2K is mostly wasted.

Object-centric multi-view rendering has two sources of redundancy: background pixels that carry no useful texture information, and sparse interactions between foreground tokens. UltraTex attacks both:

  • Background Token Dropping (BTD) discards background tokens before the DiT backbone, but keeps the original RoPE indices so spatial coordinates and the multi-view geometric prior survive the pruning.
  • Block-Sparse Attention (BSA) applies top-k sparse attention over the remaining foreground sequence instead of dense attention.
  • Foreground-Aware VAE Decoding replaces the undenoised background latent with a canonical in-distribution background and lightly fine-tunes the decoder, so 2K reconstruction stays free of the artifacts the pruning would otherwise introduce.
UltraTex pipeline

The UltraTex pipeline: background token dropping, block-sparse attention and a foreground-aware decoder.

The backbones are FLUX.1-dev and FLUX.2-Klein (4B and 9B), so the texture generator inherits the prompt and image conditioning of a general image model rather than a purpose-built texturing network.

Where the speed comes from

On common samples in the G-buffer TexVerse dataset, the released implementation reports 20.6× to 91.1× training speedup and 22.3× to 74.6× end-to-end inference speedup over the dense baseline at the same 2K output resolution.

UltraTex efficiency

Training and inference speedups over the dense 2K baseline, from the project page.

UltraTex qualitative results

Qualitative texturing results from the paper.

What was released

Running UltraTex in ComfyUI

Community node pack visualbruno/ComfyUI-UltraTex wraps the pipeline as ComfyUI nodes: UltraTex Load LoRA, UltraTex Foreground VAE Decoder, UltraTex Prep (mesh + reference), UltraTex Sampler and UltraTex Bake Texture. Load LoRA exists because the stock Load LoRA node cannot read these checkpoints (PEFT lora_A.default keys for FLUX.2, UNO processor.*_lora keys for FLUX.1), and Bake Texture back-projects the generated views into a UV texture and returns a GLB that connects straight to Preview 3D.

The diffusion model, VAE and text encoder are loaded with the stock ComfyUI loaders, so fp8 and GGUF weights and ComfyUI's memory offloading apply. Prep offers a GPU UV unwrap (comfy_gpu, around 6 s for 500k faces) or a CPU xatlas path (around 4.6 min for the same mesh).

Four example workflows ship with the node pack:

The other two are PBR at 1024px and FLUX.1-dev.

Availability

The official repository ships PyTorch training and inference code with Triton sparse-attention kernels, and the weights live on ModelScope. The ComfyUI path is the third-party node pack above, not an official Comfy-Org release, so installing it means adding the custom node and pulling the UltraTex LoRA and foreground decoder from ModelScope.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
UltraTex: 2K Multi-View Diffusion for 3D Texturing in ComfyUI | ComfyUI Wiki