Zero-Shot Augmentation
ZS-TTS models synthesize the remaining target-speaker transcripts from a small reference set, supplying broad linguistic coverage for low-resource adaptation.
Interspeech 2026
1Maum AI Inc., Republic of Korea 2Humelo Inc., Republic of Korea
ZeSTA adapts lightweight personalized TTS models with only a small amount of real target-speaker speech by adding zero-shot synthetic speech, explicitly conditioning the model on speech domain, and oversampling scarce real recordings.
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning.
To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality.
ZS-TTS models synthesize the remaining target-speaker transcripts from a small reference set, supplying broad linguistic coverage for low-resource adaptation.
A lightweight domain embedding marks each sample as real or synthetic, allowing the model to retain linguistic benefits while controlling synthetic-domain acoustic bias.
Scarce real recordings are repeated during fine-tuning, reinforcing target-speaker identity without changing the target TTS architecture or inference procedure.
Audio samples compare the reference recording, Real 10% fine-tuning, naive synthetic augmentation baseline, and the proposed ZeSTA setting with domain conditioning and real-data oversampling.
| Source ZS-TTS | Reference | Real 10% | Baseline DC X, OS X |
Proposed DC O, OS O |
|---|---|---|---|---|
| Fish-Speech | ||||
| Fish-Speech | ||||
| CosyVoice 2 | ||||
| CosyVoice 2 |
@misc{choi2026zesta,
title={ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis},
author={Choi, Youngwon and Oh, Jinwoo and Kim, Hwayeon and Kim, Hyeonyu},
year={2026},
eprint={2603.04219},
archivePrefix={arXiv},
primaryClass={cs.SD},
doi={10.48550/arXiv.2603.04219},
url={https://arxiv.org/abs/2603.04219}
}