AuK: Tencent's Open Speech Generation and Editing Model Family

ComfyUI Wiki

AuK is Tencent Hunyuan's open 1.5B speech family: TTS, voice cloning, content and acoustic editing, paralinguistic editing, enhancement, separation, with official ComfyUI nodes.

A

AuK

AudioTTSVoice CloningSpeech EditingSource Separation

Open 1.5B speech foundation model by Tencent Hunyuan, Shanghai Jiao Tong University, and the Shanghai Innovation Institute. One natural-language instruction interface covers zero-shot and instruct TTS, lyric and content editing, pitch, speed, emotion, and timbre editing, speech enhancement, and source separation. Official ComfyUI nodes and workflow included.

DeveloperTencent Hunyuan + Shanghai Jiao Tong University
Release Date2026-09-09
Architecture1.5B diffusion transformer + Qwen2.5-Omni-3B MLLM encoder
LicenseMIT
OutputUp to 30-second audio per run

Models

ModelDescriptionLicenseReleased
AuKBase model for high-quality generation, configurable NFE and CFGMIT2026-09-09
AuK-FlashDistilled variant, fixed 4-step inference with CFG=0MIT2026-09-09

Select a model above for curated weight downloads, tutorials, and related information.

Comments

Sign in with GitHub to join the discussion.

Loading comments…