Updated 2026-09-16
Word error rate (WER) is the standard metric for automatic speech recognition accuracy — calculated as the percentage of words incorrectly inserted, deleted, or substituted compared to a human reference transcript. Lower WER means more faithful transcription; high WER in voice interviews can mis-score candidates when rubrics depend on spoken content.
WER = (Substitutions + Deletions + Insertions) ÷ Total words in reference transcript. A 10% WER means one in ten words is wrong on average — acceptable for note-taking, risky for automated scoring without review.
Benchmark WER on clean American English podcast data does not predict performance on Indian English, code-mixed Hindi-English, or noisy mobile connections common in bulk hiring screens.
Domain-specific terms — medical, legal, technical jargon — inflate WER unless models are tuned or custom vocabularies are supplied.
ASR word error rate in our glossary covers the same concept in hiring-specific framing. Character error rate (CER) matters for languages without clear word boundaries. Real-time factor and voice AI latency affect whether candidates repeat themselves, indirectly worsening effective WER through truncated answers.
Interview integrity requires ASR quality, turn-taking, and rubric design evaluated together — not WER alone.
No universal threshold — depends on language, noise, and whether scores are auto-applied or human-reviewed. Validate on your own audio samples.
Yes — WER is the standard ASR accuracy metric. Our glossary uses both terms for search discoverability.
Better microphones, quieter prompts, candidate guidance on environment, and custom vocabulary for role terms help — but model choice remains primary.
Head back to Hiring glossary or start now.
· 3 free credits · pay per interview · nothing recurring