Speech Data Providers Compared: How to Choose One in 2026

macro photography of silver and black studio microphone condenser
Photo by Jonathan Velasquez on Unsplash

Speech data comes from two kinds of sources. Open datasets, such as those on Hugging Face, Mozilla Common Voice and MLCommons People's Speech, are free to download and work well for research, prototypes and benchmarks. Commercial providers, including Appen, DataForce, Defined.ai, FutureBeeAI, LXT, Pangeanic, Shaip and SpeechData.ai, license ready-made speech datasets, record new data to your spec, or both. Many teams that ship a voice product end up buying from the second group, because they need audio that sounds like their users and a license their legal team will sign off on.

Which provider fits depends less on its size than on two questions. Does the data match what your model will hear in production? And can the provider prove you have the right to train on it? This guide covers what each option offers, puts the commercial vendors side by side, and ends with a checklist to run before you buy.

A note on bias: we run SpeechData.ai, so we're one of the vendors in the table. Every other entry is based only on what each company says on its own website, checked in October 2026. We left out prices and anything we couldn't confirm.

Open speech datasets: free, with license homework

Open data costs nothing to download, but its license decides what you can do with it. These are the three places most teams start.

Hugging Face. Hugging Face is a hosting platform rather than a data provider, and anyone can publish a dataset there. Filtering its hub for audio returned 36,439 datasets in October 2026. Each repository sets its own license in its dataset card, so whether you can train a commercial model on it is decided one dataset at a time.

Mozilla Common Voice. A community-recorded corpus released under CC0, which allows commercial use. Its Scripted Speech release 27.0 covers 295 languages and about 42,600 hours of recordings, of which about 29,300 hours are validated. A separate Spontaneous Speech release (5.0) adds about 540 hours in 80 languages. Downloads now go through the Mozilla Data Collective.

MLCommons People's Speech. More than 30,000 hours of transcribed English speech, licensed for academic and commercial use under CC-BY and CC-BY-SA 4.0. If you use the CC-BY-SA portion, check its share-alike terms with your legal team.

Open data is the right place to build a baseline and an evaluation set. It gets harder when you need spontaneous conversation in one specific language or accent, speaker metadata you can balance on, or a consent record you can hand to procurement. Our list of the 15 best open-source speech datasets covers more of them, with sizes and licenses.

Commercial speech data providers

Commercial providers sell two things. Off-the-shelf datasets already exist, so you can audit a sample and license them quickly. Custom collection means the provider recruits speakers and records to your specification, which takes longer and costs more but gets you exactly the language, accent, device and environment you asked for. Most vendors offer both.

Provider Offers Speech data, as described on its own site
Appen Catalog and custom collection A speech and audio catalog spanning read, conversational and contact-center recordings; custom collection for specific demographics, acoustic environments or domains
DataForce (part of TransPerfect) Custom collection Scripted or conversational speech, wake words and multi-speaker dialogues, recorded through a contributor app, remotely or in TransPerfect's studios
Defined.ai Marketplace and custom collection Licensed datasets including scripted monologues, spontaneous dialogue and telephony recordings
FutureBeeAI Catalog and custom collection Call-center and general conversations, scripted monologues, wake words and commands, in-car speech and TTS datasets; custom projects from single-speaker prompts to multi-person conversations
LXT Catalog and custom collection Pre-built speech recognition datasets; custom scripted, spontaneous and conversational recordings with metadata captured
Pangeanic Catalog and custom collection Commercially licensable speech datasets, including dual-channel call-center conversations and meeting recordings; bespoke collection
Shaip Catalog and custom collection Call-center, general conversation, scripted monologue, wake-word and TTS datasets, shipped with transcriptions, speaker demographics and timestamps
SpeechData.ai Catalog and custom collection 60 conversational datasets totaling 73,000 hours from 7,250 native speakers, all dual-channel with time-aligned, human-verified transcripts and speaker metadata

Most of these vendors describe consented or ethically sourced collection on their sites. Whoever you shortlist, us included, ask for the consent wording speakers actually signed. Catalogs also change often, so confirm languages, volumes and license terms directly before you compare quotes.

Off-the-shelf, custom, or both

Off-the-shelf is the faster and cheaper route when a catalog already has your language and speaking style. Custom collection is the answer when nobody has recorded what you need yet: a regional dialect, a wake word, in-car audio, or your own domain vocabulary. Plenty of teams combine the two, licensing a catalog dataset for the baseline and commissioning a smaller custom collection for the gaps.

How to choose a speech data provider

Run every vendor on your shortlist through the same checks.

  1. Consent and a commercial license. Ask for the consent wording speakers signed and check that the license covers both training and deploying a commercial model. "Publicly available" is not consent.
  2. Conversational or scripted. Read speech is cheaper and easier to align, but models trained on it struggle with real conversation. Match the speaking style to what your product will hear; conversational vs. scripted speech data explains why the gap is so large.
  3. Channels. For dialogue, ask for dual-channel audio with one speaker per channel. It gives you clean per-speaker tracks and speaker-turn labels without manual annotation.
  4. Speaker metadata. Age, gender, region and recording device let you balance a training set and build fair test splits. Check that every speaker has them, not just a sample.
  5. Transcript verification. Ask how transcripts are checked (human review, a second pass, audited error rates) and which conventions they follow for fillers, overlaps and numbers.
  6. A free sample you pick. Choose the files yourself, listen to them, check transcripts word by word, and run them through your current model before you sign.
  7. A custom collection option. If a catalog covers most of what you need, find out whether the same provider can record the rest to the same spec, so you don't end up merging two vendors' conventions.

For the full buying process, including contract clauses and red flags, see our guide on how to buy speech data.

Where SpeechData.ai fits

SpeechData.ai specializes in conversational speech. The catalog has 60 datasets, one per language or dialect, each with 500 to 2,000 hours of spontaneous two-person conversation recorded in-country on dual-channel audio. Every speaker signed commercial-use consent, transcripts are time-aligned and human-verified, and each dataset page publishes its price, from $60 to $95 per hour of audio. Free samples go out on request, and when a language, dialect or domain isn't in the catalog, the same team records it to spec.

Frequently asked questions

Which companies provide speech datasets for AI training?

Commercial providers include Appen, Defined.ai, FutureBeeAI, LXT, Pangeanic, Shaip and SpeechData.ai, which all offer ready-made speech datasets as well as custom collection, and DataForce (part of TransPerfect), which collects speech data to order. For free data, Hugging Face hosts tens of thousands of audio datasets, and Mozilla Common Voice and MLCommons People's Speech are large open corpora.

Can I use open speech datasets to train a commercial model?

Often, but check each license. Mozilla Common Voice is released under CC0 and MLCommons People's Speech under CC-BY and CC-BY-SA 4.0, and both allow commercial use. Datasets on Hugging Face each carry their own license, set by whoever uploaded them. Most of Common Voice is scripted speech, so teams that need spontaneous conversation, specific accents or documented speaker consent usually license commercial data as well.

What should I ask a speech data provider before buying?

Ask for the consent wording speakers signed, whether the license covers commercial training and deployment, whether the audio is conversational or scripted and how many channels it has, which speaker metadata comes with every recording, how transcripts are verified, and for a free sample you choose yourself. If the catalog doesn't cover everything, ask whether they can collect the rest to the same spec.

Read also

Training a voice model?

Browse 60 conversational speech datasets with transcripts, metadata, and a commercial license. Samples are free on request.