Key takeaways
- Benchmarks decide what gets built.
- A long line of work has shown that ASR error rates are not evenly distributed across the people using them.
- None of that is visible on a leaderboard, and not because the leaderboard is hiding it.
What happened
Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy: All of that makes one number (WER) harder to game. It is still one number.
Speaker concentration and diversity beyond the fields above. Three properties follow, and each is a claim about variance rather than volume. Monsoon ships 18 columns per segment, of which 12 are metadata, where most public ASR test sets ship an identifier, a transcript and a duration. Demographic fields are complete or near-complete; contributors consented to this use. The two languages have different geographic shapes, and the shape is informative.
The Hindi sets concentrate in the Hindi belt, with Uttar Pradesh accounting for roughly 40% of speakers, which is what a Hindi corpus sampled by population should look like. The Indian English sets are much flatter: no state exceeds 13%, and a third of speakers come from outside the eight largest. Public and private halves match closely on both.
Indian state boundaries were drawn along linguistic lines, so district and state carry real accent signal, which is why these fields are released rather than summarised away.
Analysis of this kind has been reported at scale for Indian ASR: district-level error rates spanning roughly 4% to 44%, with underrepresented regions well behind the Hindi belt and the metros, plus disaggregation by audio quality, speaking rate, utterance duration, gender, age and device. Those runs were on a closed benchmark. Monsoon makes the same class of analysis possible on a public leaderboard test set.
Why it matters
A long line of work has shown that ASR error rates are not evenly distributed across the people using them. Racial disparities in automated speech recognition found commercial systems roughly twice as bad for Black speakers as for white speakers, and Quantifying Bias in Automatic Speech Recognition found further differences by gender, age and accent.
None of that is visible on a leaderboard, and not because the leaderboard is hiding it. The test sets it runs on record what was said and almost nothing about who said it. To address this gap, we introduce two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN and Monsoon hi-IN.
Hindi, spoken by more than half a billion people, is the first Indic language on a multilingual tab that currently covers only European languages. Each set is released as a public split, available for self-scoring, and a private split withheld to limit benchmark-specific optimisation. The four splits are speaker-disjoint, comprising 4,888 speakers, with 12 speaker attributes recorded for each.
A test set can only expose a failure mode it varies along. Most benchmarks are built from whatever audio was readily available. Monsoon was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio. Each is a way an aggregate WER can be right on average and wrong for a particular population.
The collection method follows from that. Four splits, two languages, collected through one pipeline. The data is sourced from unscripted dual-channel spontaneous conversations, with clips segmented from a single channel so that each clip carries one speaker. Along with the fields reported in the table, each clip also records occupation, education, marital status, income band, handset brand, current city and years in the current district.
Five clips from the public Indian English split, with the metadata each one carries: 29-year-old woman, West Tripura, Tripura. Student, samsung SM-G781B. 32-year-old woman, Satna, Madhya Pradesh. Unemployed, samsung SM-E146B. 22-year-old man, Rohtas, Bihar. Student, motorola moto g54 5G. 27-year-old woman, Warangal, Telangana. Unemployed, vivo V2247. 57-year-old man, Puducherry. Private job, Xiaomi M2006C3LI. The English sets use standard string references, where the leaderboard's normaliser collapses most spelling variation.
Hindi has far more of it, and no normaliser can resolve it, because the variants are not a fixed mapping between two conventions. The Hindi sets therefore ship a lattice: for each span of the transcript, a list of the spellings that are accepted as correct. Monsoon is small measured in hours and large measured in speakers. That is the design, and it is where most of the value sits.
What to watch
Broad geographic coverage requires recruitment across hundreds of districts rather than longer sessions from fewer speakers, and distributed recruitment at this scale introduces failure modes that a smaller collection does not face: contributors gaming the task, played-back audio submitted as live speech, and inattentive annotation. Each is addressed by an explicit check.
What follows is one example, run on the public Indian English split, to show the kind of evaluation the metadata makes possible. It is not the finding the sets exist to deliver; it is an illustration of what becomes answerable once every clip carries a speaker. 99 WER on this set. 18 points from best to worst, inside what five hours can resolve. Ranked on the corpus, they are the same model. Grouping speakers by region tells a different story.



