Back to Discover

Tencent open-sources AuK, a 1.5B model for instruction-guided speech generation and editing

Official AuK performance comparison chart across speech generation, editing, enhancement, and separation

Tencent Hunyuan released AuK and distilled AuK-Flash on Sept. 9, 2026, with MIT weights on Hugging Face and ModelScope and a technical report on arXiv (2609.08936). The 1.5B flow backbone unifies zero-shot and instruct TTS, content and acoustic edits, paralinguistic edits, enhancement, and separation behind natural-language instructions. Vendor tables lead several generation and editing benchmarks; independent reproduction is still missing.

Tencent Hunyuan open-sourced AuK on Sept. 9, 2026, as a unified speech generation and editing foundation model. The technical report is arXiv 2609.08936, submitted Sept. 8. Weights and inference code are on Hugging Face under tencent/AuK and tencent/AuK-Flash, mirrored on ModelScope, with training and serving code in the Tencent-Hunyuan/AuK GitHub repository under an MIT license.

Company claim: the diffusion backbone is about 1.5 billion parameters, with 10 dual-stream MMDiT blocks and 20 single-stream DiT blocks, a 1536 hidden size, and 24 attention heads. Semantic conditioning uses a frozen Qwen2.5-Omni-3B multimodal encoder with learnable layer fusion. Acoustic conditioning and reconstruction use a jointly trained speech, general-audio, and music VAE that maps 24 kHz audio to 64-dimensional latents at 50 Hz.

The report says pre-training covers five task families through a shared instruction-plus-optional-audio interface: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. Authors report about 3.03 billion instruction-audio instances and 1.95 million hours of effective supervision. Training starts with generation-only warm-up, then joint generation-editing pre-training, then human-feedback preference optimization for open-ended editing and Flow-GRPO reinforcement learning for generation.

AuK-Flash is a distilled student. Under the report's matched settings it runs a 4-step Euler sampler without classifier-free guidance, against the full model's 32 function evaluations at CFG 2.0, for a claimed 4.5 times wall-clock speedup. Both DiT checkpoints on Hugging Face are about 6.12 GB (auk_base.safetensors and auk_flash.safetensors each measure 6,122,209,092 bytes); the shared VAE file is about 637 MB. The Qwen encoder is downloaded separately.

Company-reported evals: on Seed-TTS-Eval, AuK averages 2.65% recognition error and 0.795 speaker similarity, ahead of listed baselines including Qwen3-TTS and Seed-TTS. On MMAE-Speech with the Prompt Enhancer enabled, AuK posts IFR 48.23, CR 88.11, and EMR 12.44; AuK-Flash leads edit-ratio EMR at 13.85. The paper says results are leading on zero-shot and instruction-controlled generation and general instruction-guided editing, and competitive on signal-level restoration. Those numbers are vendor tables, not third-party leaderboards.

Independent check: Ground Truth confirmed the Sept. 9 MIT open-source drop and noted that performance charts remain author-reported with no independent baseline reproduction yet. OrcaRouter's Sept. 10 analysis, reading the same tables, argues Flash is not a simple quality downgrade and often leads perceptual UTMOS on enhancement sets, while the base model keeps lower word error on Chinese generation and content editing. Neither outlet ran a fresh harness.

Limitations the paper itself flags: robust native understanding of unconstrained editing requests remains incomplete, and the system still benefits from explicit task routing and a Prompt Enhancer. Ground Truth also notes the model card and paper do not describe watermarking or provenance signals on generated audio. Contested institutional context: AuK is listed as the Single Model Track end-to-end baseline for the ICASSP 2027 Audio Editing Challenge, which IEEE lists as a grand challenge; the challenge site says Tencent sponsors the prize pool and several organizers overlap AuK co-authors, so that adoption is not an independent quality audit.

  • @theAIsearch Video

    Discovery roundup covering AuK among other AI news (hook only; not primary evidence).

    Watch on YouTube