Veda Sparse Attention Comes to ComfyUI: Faster MiniMax H3

ComfyUI Wikinews

The Veda team ships an official ComfyUI node that overrides MiniMax H3 attention with a distilled sparse predictor, skipping 90% of attention tiles for up to 3x faster video.

<strong>Veda</strong> is a learned sparse-attention method for video diffusion models from ByteDance, HKU and USTC, presented at ICML 2026. Its authors have now packaged it as an official ComfyUI custom node, <strong>Veda Sparse Attention (MiniMax H3)</strong>, which runs MiniMax H3 attention at roughly a tenth of its usual cost without changing a single model weight.
FlashAttention-3 full attention next to Veda at 95% sparsity on the same Waver-T2V-12B clip

From the Veda project page: Waver-T2V-12B at 720p, 241 frames. Full attention on the left, Veda's tile-skipping path on the right.

What Veda does

Rather than compressing or retraining the model, Veda learns which parts of the attention map matter. A distilled lightweight predictor scores attention tiles and keeps roughly the top 10%, and a tile-skipping kernel fetches only the selected K/V tiles. Because it changes how attention is computed rather than the weights themselves, it stays compatible with any LoRA and with fine-tuned or quantized MiniMax H3 checkpoints.

The project page describes the method in three stages: max-pooled full attention supplies a tile-level teacher distribution, a Triplet Pooling estimator with per-head Q/K projections reconstructs that distribution, and head-aware top-k selection emits the tile mask the kernel follows. Per-layer, per-head tile geometry is chosen to fit heterogeneous spatial and temporal heads inside a fixed hardware budget.

On MiniMax H3 the released predictor is trained with the 8-step Turbo LoRA. The repository notes that the checkpoint is modality and step agnostic: it applies to T2VA, FL2VA and R2VA, at any number of sampling steps, with a dedicated R2VA fine-tune announced for a future release.

The ComfyUI node

The node is deliberately small: MODEL in, MODEL out. It installs as an attention override, so it goes on the MODEL wire after the model and any LoRA loaders and just before the guider or sampler. Bypassing it with Ctrl+B renders the same seed with full attention for a direct comparison.

InputDefaultDescription
generated_sparsity90%Sparsity of the generated video's attention. A whole number such as 24 keeps exactly that many 128-token key tiles instead.
reference_sparsity90%The same for reference tokens: first/last frames, guide frames, reference images and videos. 0% gives them full attention.
full_attention_layersempty0-based DiT blocks that keep full attention, e.g. 0, 1, 47-49.
full_attention_stepsempty0-based sampling steps that keep full attention.
verboseoffReport per-phase attention timing, call counts and predictor details after each run.

After a run the node prints what it actually did:

Veda done · Triton INT8 (SM120)
Video: 1344x768 · 5.2 s
Attention computed: 10.9% of full attention (89.1% skipped)

If a kernel cannot run on the current GPU, the node says so and the model falls back to its own attention for that run instead of producing a broken render.

Do not stack the node with ComfyUI's own Model Sparse Attention on H3. That node replaces the attention blocks outright, so Veda would never be called. The Veda node detects the combination and warns about it.

Performance

T2VA at 1344x768, 124 frames (5.2 s), 8-step Turbo LoRA and 90% sparsity, measured on an RTX 5070 12 GB under Windows 11:

Attention pathPer step8 stepsAttention per step
ComfyUI default40.7 s342 s31.1 s
ComfyUI --use-sage-attention24.8 s231 s15.2 s
Veda sparse INT814.0 s130 s4.41 s

That is roughly 2.9x end to end and 7.1x on attention alone, and the gap widens with clip length. A single attention layer at 104k tokens with 90% sparsity takes 528 ms against 11.8 s for full attention.

Veda at 95% sparsity next to the FlashAttention-3 baseline on the same clip

Veda at 95% sparsity. The Veda project reports 5.1x end-to-end acceleration on Waver-T2V-12B at 720p and 241 frames, from 19.4 minutes to 3.8 minutes.

The predictor repository adds end-to-end speedups over its own full-attention baseline, all at 10% attention kept with the Turbo LoRA:

GPUClipAttention speedupEnd-to-end speedup
RTX PRO 6000 Blackwell16:9 · 14.4 s6.79x3.12x
RTX 409016:9 · 5.17 s4.75x1.77x
RTX 409016:9 · 10.1 s6.22x2.42x
RTX 409016:9 · 14.4 s6.31x2.82x

Full attention on the left and Veda on the right, same prompt, seed and Turbo LoRA, generated on a single RTX PRO 6000 Blackwell.

Hardware support

One Triton kernel covers every NVIDIA GPU from SM80 onward on Windows and Linux, and Apple silicon runs through MLX. Cards without a matching kernel are told so, and the model keeps its own attention.

HardwareStatus
RTX 30 / A100 / RTX 40 / L40 (sm80-89)code path ready, not yet verified by the authors
H100 / H200 (sm90)code path ready, not yet verified by the authors
B200 / B300 (sm100 / sm103)code path ready, not yet verified by the authors
RTX 50, RTX PRO 6000 Blackwell (sm120)verified: RTX 5070, Windows 11
DGX Spark / GB10 (sm121)code path ready, not yet verified by the authors
Apple silicon (M series)verified: M3 Pro, macOS 15

Getting started

Veda needs ComfyUI 0.38.0 or newer. Install the node from ComfyUI Manager by searching "Veda", or from the CLI:

comfy node install veda-sparse-attention

Then place the 275 MB predictor at ComfyUI/models/veda/:

hf download Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview \
  minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors \
  --local-dir ComfyUI/models/veda

The node does not download anything by itself. The easier path is to open Workflow -> Browse Templates -> Veda-on-ComfyUI and pick a template, which offers the predictor in ComfyUI's missing-model dialog. Two example workflows ship with the repository:

Availability

The node is on the Comfy Registry as veda-sparse-attention with source at veda-sparse/Veda-on-ComfyUI. The predictor checkpoint lives at Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview, and the training code is at veda-sparse/Miowtion. The released predictor is tagged as a preview, trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5, 10 and 14 seconds with the 8-step Turbo LoRA; other sizes fall back to the nearest trained tile plan and should be compared against full attention.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
Veda Sparse Attention Comes to ComfyUI: Faster MiniMax H3 | ComfyUI Wiki