Corpus design for Uzbek text-to-speech
The data side of an Uzbek text-to-speech system, at Openbank.
The corpus pipeline chains diarization, ECAPA speaker verification, speech-boundary segmentation, selective source separation and 2-of-3 consensus ASR filtering, with an immutable provenance ledger and audit gates a clip has to clear before it can enter the training manifest.
The design is complete: corpus construction, curation thresholds, the evaluation protocol and the manifest gate. It started from a survey of diffusion and flow-matching TTS, multilingual speech corpora including FLEURS, and TTS scaling studies.