Speech Scoring Glossary
These are the terms behind pronunciation-scoring APIs — the algorithms, the processing steps, and the evaluation vocabulary that vendors use in their documentation and that determine what a scoring product can and cannot actually do. Each entry states what the thing is in one paragraph, then explains how it works and where it breaks.
The terms
Goodness of Pronunciation (GOP)
GOP score
Goodness of Pronunciation (GOP) is an algorithm that scores how well a speaker produced a specific expected phoneme, by comparing how likely an acoustic model thinks that phoneme is over the relevant stretch of audio against how likely the model thinks every other phoneme is.
Read the definitionForced alignment
Forced alignment is the process of taking audio together with a transcript of what was said, and determining the exact start and end time of each word and each phoneme within that audio.
Read the definitionMispronunciation detection
MDD — mispronunciation detection and diagnosis
Mispronunciation detection is the task of automatically identifying which specific sounds in a spoken utterance were produced incorrectly.
Read the definitionPhoneme scoring
per-phoneme pronunciation scoring
Phoneme scoring is pronunciation assessment that returns a separate score for each individual speech sound in a word, rather than a single score for the word or the sentence as a whole.
Read the definitionHow to use this glossary
If you are choosing between pronunciation-scoring vendors, read Goodness of Pronunciation first — almost every commercial engine in this category is a GOP variant, so its limitations are the industry's limitations rather than any one product's. Then read mispronunciation detection, which explains why a single published accuracy percentage is close to uninterpretable without the base rate of errors in the test set.
Everything here is written to be read on its own. Where a claim comes from our own evaluation runs, the page says so and gives the sample size.