Veda Sparse Attention Comes to ComfyUI: Faster MiniMax H3
The Veda team ships an official ComfyUI node that overrides MiniMax H3 attention with a distilled sparse predictor, skipping 90% of attention tiles for up to 3x faster video.
From the Veda project page: Waver-T2V-12B at 720p, 241 frames. Full attention on the left, Veda's tile-skipping path on the right.
What Veda does
Rather than compressing or retraining the model, Veda learns which parts of the attention map matter. A distilled lightweight predictor scores attention tiles and keeps roughly the top 10%, and a tile-skipping kernel fetches only the selected K/V tiles. Because it changes how attention is computed rather than the weights themselves, it stays compatible with any LoRA and with fine-tuned or quantized MiniMax H3 checkpoints.
The project page describes the method in three stages: max-pooled full attention supplies a tile-level teacher distribution, a Triplet Pooling estimator with per-head Q/K projections reconstructs that distribution, and head-aware top-k selection emits the tile mask the kernel follows. Per-layer, per-head tile geometry is chosen to fit heterogeneous spatial and temporal heads inside a fixed hardware budget.
On MiniMax H3 the released predictor is trained with the 8-step Turbo LoRA. The repository notes that the checkpoint is modality and step agnostic: it applies to T2VA, FL2VA and R2VA, at any number of sampling steps, with a dedicated R2VA fine-tune announced for a future release.
The ComfyUI node
The node is deliberately small: MODEL in, MODEL out. It installs as an attention override, so it goes on the MODEL wire after the model and any LoRA loaders and just before the guider or sampler. Bypassing it with Ctrl+B renders the same seed with full attention for a direct comparison.
| Input | Default | Description |
|---|---|---|
generated_sparsity | 90% | Sparsity of the generated video's attention. A whole number such as 24 keeps exactly that many 128-token key tiles instead. |
reference_sparsity | 90% | The same for reference tokens: first/last frames, guide frames, reference images and videos. 0% gives them full attention. |
full_attention_layers | empty | 0-based DiT blocks that keep full attention, e.g. 0, 1, 47-49. |
full_attention_steps | empty | 0-based sampling steps that keep full attention. |
verbose | off | Report per-phase attention timing, call counts and predictor details after each run. |
After a run the node prints what it actually did:
Veda done · Triton INT8 (SM120)
Video: 1344x768 · 5.2 s
Attention computed: 10.9% of full attention (89.1% skipped)If a kernel cannot run on the current GPU, the node says so and the model falls back to its own attention for that run instead of producing a broken render.
Performance
T2VA at 1344x768, 124 frames (5.2 s), 8-step Turbo LoRA and 90% sparsity, measured on an RTX 5070 12 GB under Windows 11:
| Attention path | Per step | 8 steps | Attention per step |
|---|---|---|---|
| ComfyUI default | 40.7 s | 342 s | 31.1 s |
ComfyUI --use-sage-attention | 24.8 s | 231 s | 15.2 s |
| Veda sparse INT8 | 14.0 s | 130 s | 4.41 s |
That is roughly 2.9x end to end and 7.1x on attention alone, and the gap widens with clip length. A single attention layer at 104k tokens with 90% sparsity takes 528 ms against 11.8 s for full attention.
Veda at 95% sparsity. The Veda project reports 5.1x end-to-end acceleration on Waver-T2V-12B at 720p and 241 frames, from 19.4 minutes to 3.8 minutes.
The predictor repository adds end-to-end speedups over its own full-attention baseline, all at 10% attention kept with the Turbo LoRA:
| GPU | Clip | Attention speedup | End-to-end speedup |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 16:9 · 14.4 s | 6.79x | 3.12x |
| RTX 4090 | 16:9 · 5.17 s | 4.75x | 1.77x |
| RTX 4090 | 16:9 · 10.1 s | 6.22x | 2.42x |
| RTX 4090 | 16:9 · 14.4 s | 6.31x | 2.82x |
Full attention on the left and Veda on the right, same prompt, seed and Turbo LoRA, generated on a single RTX PRO 6000 Blackwell.
Hardware support
One Triton kernel covers every NVIDIA GPU from SM80 onward on Windows and Linux, and Apple silicon runs through MLX. Cards without a matching kernel are told so, and the model keeps its own attention.
| Hardware | Status |
|---|---|
| RTX 30 / A100 / RTX 40 / L40 (sm80-89) | code path ready, not yet verified by the authors |
| H100 / H200 (sm90) | code path ready, not yet verified by the authors |
| B200 / B300 (sm100 / sm103) | code path ready, not yet verified by the authors |
| RTX 50, RTX PRO 6000 Blackwell (sm120) | verified: RTX 5070, Windows 11 |
| DGX Spark / GB10 (sm121) | code path ready, not yet verified by the authors |
| Apple silicon (M series) | verified: M3 Pro, macOS 15 |
Getting started
Veda needs ComfyUI 0.38.0 or newer. Install the node from ComfyUI Manager by searching "Veda", or from the CLI:
comfy node install veda-sparse-attentionThen place the 275 MB predictor at ComfyUI/models/veda/:
hf download Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview \
minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors \
--local-dir ComfyUI/models/vedaThe node does not download anything by itself. The easier path is to open Workflow -> Browse Templates -> Veda-on-ComfyUI and pick a template, which offers the predictor in ComfyUI's missing-model dialog. Two example workflows ship with the repository:
Availability
The node is on the Comfy Registry as veda-sparse-attention with source at veda-sparse/Veda-on-ComfyUI. The predictor checkpoint lives at Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview, and the training code is at veda-sparse/Miowtion. The released predictor is tagged as a preview, trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5, 10 and 14 seconds with the 8-step Turbo LoRA; other sizes fall back to the nearest trained tile plan and should be compared against full attention.
Comments
Sign in with GitHub to join the discussion.