Yugo - AI & Speech Data
Speech data collection built for machine learning at scale.
A SaaS platform for collecting high-volume, multilingual speech data from distributed contributors, built for AI and machine-learning training workflows.
- Industry
- Artificial Intelligence · Machine Learning
- Services
- AI ServicesData ServicesProduct EngineeringWeb & Mobile
- Platform
- Mobile · SaaS Platform
Selected implementation details are confidential and have been omitted.
Yugo is a SaaS platform for speech data collection. It gathers the recorded human speech that AI and machine-learning teams need to train and evaluate models: at volume, across multiple languages, from contributors distributed across locations.
The platform supports both scripted prompt recording, where a contributor reads supplied text, and real-time conversation recording, where natural speech between participants is captured with separate audio channels.
KwikTech partnered on the product engineering and mobile experience: the workflows contributors move through, the recording experience on device, and the platform architecture needed to support collection at scale.
The challenge
Speech datasets live or die on consistency. A recording that is unusable (wrong language variant, overlapping speakers on one channel, an incomplete prompt) is not merely a gap in the dataset; it is effort spent producing something that has to be discarded.
The difficulty is that the people producing those recordings are not in a studio. They are distributed contributors, on their own devices, in ordinary environments, often working in a language the reviewing team does not speak. Every requirement the dataset has must be enforced by the product itself, because there is no supervisor standing behind the contributor.
That has to hold while the collection workflow stays simple enough that a large, non-specialist contributor base can complete it reliably, and while the platform handles both scripted prompts and multi-participant conversation.
Data quality at the point of capture
Requirements have to be enforced during recording. Discovering a problem after collection means the effort is already lost.
Two different recording modes
Scripted prompt reading and real-time conversation capture are genuinely different workflows within one product.
Separate audio channels
Conversation recording keeps participants on distinct channels, which the capture experience must handle correctly on device.
Multilingual by default
Collection spans multiple languages, so language is a structural property of the workflow rather than an interface setting.
Distributed, non-specialist contributors
Contributors work independently on their own devices. The product is the only thing guiding them.
The approach
We designed the contributor experience around the constraint that quality has to be earned during capture, not repaired afterwards. The mobile experience guides a contributor through what to record, in which language, and in which mode, so the requirements of the dataset are expressed as the workflow rather than as instructions to remember.
Scripted and conversational collection were treated as distinct workflows sharing one product surface, since reading supplied text and holding a recorded conversation demand different things of the interface, and, for conversation, correct handling of separate audio channels on device.
Multilingual support was built structurally rather than bolted on, so adding a language extends the existing collection workflow instead of duplicating it. The platform architecture was designed for sustained, large-volume collection from a distributed contributor base.
What we delivered
A SaaS speech data collection platform with a mobile recording experience supporting scripted prompts, real-time conversation capture with separate audio channels, and multilingual collection workflows at volume.
Mobile recording experience
The on-device workflow contributors use to produce recordings.
Scripted prompt recording
Guided capture where contributors record supplied text.
Conversation recording
Real-time conversation capture with participants on separate audio channels.
Multilingual collection
Collection workflows spanning multiple languages as a structural feature.
Where the work concentrated. Selected implementation details are confidential.
Mobile product engineering
The contributor-facing recording experience, the surface where data quality is determined.
Audio workflows
Recording workflows covering scripted prompts and real-time conversation with separate audio channels.
Multilingual experiences
Collection workflows structured so language is intrinsic to the product rather than an added option.
Data collection workflows
Guided workflows that hold contributions to dataset requirements at the point of capture.
Scalable product architecture
A platform designed for sustained, high-volume collection from distributed contributors.
Scripted prompt recording
Contributors record supplied text through a guided workflow.
Real-time conversation recording
Natural conversation captured between participants.
Separate audio channels
Conversation participants recorded on distinct channels.
Multilingual workflows
Collection across multiple languages within one platform.
Distributed contribution
Contributors participating from their own devices across locations.
High-volume collection
Built to support speech data collection at the volumes model training requires.
- A mobile-first collection experience for distributed contributors
- Scripted and conversational recording supported within one platform
- Multilingual collection handled as a structural capability
- Conversation capture preserving separate audio channels
- A platform architecture suited to sustained, large-volume collection
The capabilities behind this work.
Building an AI or data product?
Training data is an engineering problem before it is a modelling one. Let's talk about the platform that produces it.