Intervues

What is word error rate (WER)?

Updated 2026-09-16

Word error rate (WER) is the standard metric for automatic speech recognition accuracy — calculated as the percentage of words incorrectly inserted, deleted, or substituted compared to a human reference transcript. Lower WER means more faithful transcription; high WER in voice interviews can mis-score candidates when rubrics depend on spoken content.

How WER is calculated

WER = (Substitutions + Deletions + Insertions) ÷ Total words in reference transcript. A 10% WER means one in ten words is wrong on average — acceptable for note-taking, risky for automated scoring without review.

Benchmark WER on clean American English podcast data does not predict performance on Indian English, code-mixed Hindi-English, or noisy mobile connections common in bulk hiring screens.

Domain-specific terms — medical, legal, technical jargon — inflate WER unless models are tuned or custom vocabularies are supplied.

WER and interview fairness

  • Scoring rubrics applied to bad transcripts produce false negatives — qualified candidates marked low on communication.
  • Indian English ASR challenges include accent variation, fast speech, and network compression on WhatsApp or phone channels.
  • Monitor WER by demographic proxy when possible — disparate ASR error rates can contribute to adverse impact.
  • Human review or re-ask flows for low-confidence segments reduce unfair outcomes.
  • Do not publish vendor WER claims as your fairness guarantee — validate on your candidate population.

Related metrics

ASR word error rate in our glossary covers the same concept in hiring-specific framing. Character error rate (CER) matters for languages without clear word boundaries. Real-time factor and voice AI latency affect whether candidates repeat themselves, indirectly worsening effective WER through truncated answers.

Interview integrity requires ASR quality, turn-taking, and rubric design evaluated together — not WER alone.

Frequently asked

What WER is good enough for hiring?

No universal threshold — depends on language, noise, and whether scores are auto-applied or human-reviewed. Validate on your own audio samples.

Is WER the same as ASR word error rate?

Yes — WER is the standard ASR accuracy metric. Our glossary uses both terms for search discoverability.

Can WER be improved without changing models?

Better microphones, quieter prompts, candidate guidance on environment, and custom vocabulary for role terms help — but model choice remains primary.

Ready to practise?

Head back to Hiring glossary or start now.

· 3 free credits · pay per interview · nothing recurring

Start practising