AuK: Tencent's Open Speech Generation and Editing Model Family
ComfyUI Wiki
AuK is Tencent Hunyuan's open 1.5B speech family: TTS, voice cloning, content and acoustic editing, paralinguistic editing, enhancement, separation, with official ComfyUI nodes.
A
AuK
AudioTTSVoice CloningSpeech EditingSource SeparationOpen 1.5B speech foundation model by Tencent Hunyuan, Shanghai Jiao Tong University, and the Shanghai Innovation Institute. One natural-language instruction interface covers zero-shot and instruct TTS, lyric and content editing, pitch, speed, emotion, and timbre editing, speech enhancement, and source separation. Official ComfyUI nodes and workflow included.
| Developer | Tencent Hunyuan + Shanghai Jiao Tong University |
| Release Date | 2026-09-09 |
| Architecture | 1.5B diffusion transformer + Qwen2.5-Omni-3B MLLM encoder |
| License | MIT |
| Output | Up to 30-second audio per run |
Models
| Model | Description | License | Released |
|---|---|---|---|
| AuK | Base model for high-quality generation, configurable NFE and CFG | MIT | 2026-09-09 |
| AuK-Flash | Distilled variant, fixed 4-step inference with CFG=0 | MIT | 2026-09-09 |
Select a model above for curated weight downloads, tutorials, and related information.
Comments
Sign in with GitHub to join the discussion.