Speech data comes from two kinds of sources. Open datasets, such as those on Hugging Face, Mozilla Common Voice and MLCommons People's Speech, are free to download and work well for research, prototypes and benchmarks. Commercial providers, including Appen, DataForce, Defined.ai, FutureBeeAI, LXT, Pangeanic, Shaip and SpeechData.ai, license ready-made speech datasets, record new data to your spec, or both. Many teams that ship a voice product end up buying from the second group, because they need audio that sounds like their users and a license their legal team will sign off on.
Which provider fits depends less on its size than on two questions. Does the data match what your model will hear in production? And can the provider prove you have the right to train on it? This guide covers what each option offers, puts the commercial vendors side by side, and ends with a checklist to run before you buy.
A note on bias: we run SpeechData.ai, so we're one of the vendors in the table. Every other entry is based only on what each company says on its own website, checked in October 2026. We left out prices and anything we couldn't confirm.
Open speech datasets: free, with license homework
Open data costs nothing to download, but its license decides what you can do with it. These are the three places most teams start.
Hugging Face. Hugging Face is a hosting platform rather than a data provider, and anyone can publish a dataset there. Filtering its hub for audio returned 36,439 datasets in October 2026. Each repository sets its own license in its dataset card, so whether you can train a commercial model on it is decided one dataset at a time.
Mozilla Common Voice. A community-recorded corpus released under CC0, which allows commercial use. Its Scripted Speech release 27.0 covers 295 languages and about 42,600 hours of recordings, of which about 29,300 hours are validated. A separate Spontaneous Speech release (5.0) adds about 540 hours in 80 languages. Downloads now go through the Mozilla Data Collective.
MLCommons People's Speech. More than 30,000 hours of transcribed English speech, licensed for academic and commercial use under CC-BY and CC-BY-SA 4.0. If you use the CC-BY-SA portion, check its share-alike terms with your legal team.
Open data is the right place to build a baseline and an evaluation set. It gets harder when you need spontaneous conversation in one specific language or accent, speaker metadata you can balance on, or a consent record you can hand to procurement. Our list of the 15 best open-source speech datasets covers more of them, with sizes and licenses.
Commercial speech data providers
Commercial providers sell two things. Off-the-shelf datasets already exist, so you can audit a sample and license them quickly. Custom collection means the provider recruits speakers and records to your specification, which takes longer and costs more but gets you exactly the language, accent, device and environment you asked for. Most vendors offer both.
| Provider | Offers | Speech data, as described on its own site |
|---|---|---|
| Appen | Catalog and custom collection | A speech and audio catalog spanning read, conversational and contact-center recordings; custom collection for specific demographics, acoustic environments or domains |
| DataForce (part of TransPerfect) | Custom collection | Scripted or conversational speech, wake words and multi-speaker dialogues, recorded through a contributor app, remotely or in TransPerfect's studios |
| Defined.ai | Marketplace and custom collection | Licensed datasets including scripted monologues, spontaneous dialogue and telephony recordings |
| FutureBeeAI | Catalog and custom collection | Call-center and general conversations, scripted monologues, wake words and commands, in-car speech and TTS datasets; custom projects from single-speaker prompts to multi-person conversations |
| LXT | Catalog and custom collection | Pre-built speech recognition datasets; custom scripted, spontaneous and conversational recordings with metadata captured |
| Pangeanic | Catalog and custom collection | Commercially licensable speech datasets, including dual-channel call-center conversations and meeting recordings; bespoke collection |
| Shaip | Catalog and custom collection | Call-center, general conversation, scripted monologue, wake-word and TTS datasets, shipped with transcriptions, speaker demographics and timestamps |
| SpeechData.ai | Catalog and custom collection | 60 conversational datasets totaling 73,000 hours from 7,250 native speakers, all dual-channel with time-aligned, human-verified transcripts and speaker metadata |
Most of these vendors describe consented or ethically sourced collection on their sites. Whoever you shortlist, us included, ask for the consent wording speakers actually signed. Catalogs also change often, so confirm languages, volumes and license terms directly before you compare quotes.
Off-the-shelf, custom, or both
Off-the-shelf is the faster and cheaper route when a catalog already has your language and speaking style. Custom collection is the answer when nobody has recorded what you need yet: a regional dialect, a wake word, in-car audio, or your own domain vocabulary. Plenty of teams combine the two, licensing a catalog dataset for the baseline and commissioning a smaller custom collection for the gaps.
How to choose a speech data provider
Run every vendor on your shortlist through the same checks.
- Consent and a commercial license. Ask for the consent wording speakers signed and check that the license covers both training and deploying a commercial model. "Publicly available" is not consent.
- Conversational or scripted. Read speech is cheaper and easier to align, but models trained on it struggle with real conversation. Match the speaking style to what your product will hear; conversational vs. scripted speech data explains why the gap is so large.
- Channels. For dialogue, ask for dual-channel audio with one speaker per channel. It gives you clean per-speaker tracks and speaker-turn labels without manual annotation.
- Speaker metadata. Age, gender, region and recording device let you balance a training set and build fair test splits. Check that every speaker has them, not just a sample.
- Transcript verification. Ask how transcripts are checked (human review, a second pass, audited error rates) and which conventions they follow for fillers, overlaps and numbers.
- A free sample you pick. Choose the files yourself, listen to them, check transcripts word by word, and run them through your current model before you sign.
- A custom collection option. If a catalog covers most of what you need, find out whether the same provider can record the rest to the same spec, so you don't end up merging two vendors' conventions.
For the full buying process, including contract clauses and red flags, see our guide on how to buy speech data.
Where SpeechData.ai fits
SpeechData.ai specializes in conversational speech. The catalog has 60 datasets, one per language or dialect, each with 500 to 2,000 hours of spontaneous two-person conversation recorded in-country on dual-channel audio. Every speaker signed commercial-use consent, transcripts are time-aligned and human-verified, and each dataset page publishes its price, from $60 to $95 per hour of audio. Free samples go out on request, and when a language, dialect or domain isn't in the catalog, the same team records it to spec.