AuK-Flash: 4-Step Distilled Speech Model by Tencent
AuK-Flash is the distilled variant of Tencent's AuK speech model, running fixed 4-step inference with CFG=0 for fast generation and editing in ComfyUI.
AuK-Flash
TTSDistilledFast InferenceSpeech EditingDistilled variant of Tencent's AuK speech model. Uses fixed 4-step inference with CFG=0 for faster generation and editing while keeping the full task coverage: TTS, voice cloning, speech editing, enhancement, and separation.
| Developer | Tencent Hunyuan + Shanghai Jiao Tong University |
| Release Date | 2026-09-09 |
| Architecture | 1.5B diffusion transformer, distilled (NFE=4) |
| License | MIT |
| Parameters | 1.5B |
| Inference | 4 fixed steps, CFG=0 |
Capabilities
AuK-Flash shares the same task coverage and instruction interface as the AuK base model: zero-shot and instruct TTS, content, acoustic, and paralinguistic editing, speech enhancement, and source separation. The trade-off is speed over maximum quality.
ComfyUI Usage
Load auk_flash.safetensors in AuK Model Loader and set:
nfe_steps=4cfg_strength=0sway_sampling_coef=-1
The same Qwen2.5-Omni-3B encoder and the 30-second source-plus-target limit apply. See the official ComfyUI guide for details.
Related
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.