How ArticScore Scores Individual Speech Sounds
Automatic pronunciation scoring does not behave the same way on every sound. These pages take the sounds that matter most in speech therapy — /r/, /s/, /l/ and “th” — and set out, for each one, why it is difficult for children, which of its common errors a phoneme scorer can detect, which it structurally cannot, and what ArticScore actually returns for a word containing it.
The sounds
/ɹ/
ARPAbet R, ER, AXR · alveolar approximant (rhotic)
The /r/ sound, written /ɹ/ in IPA, is an alveolar approximant and the latest-acquired consonant in American English.
Scoring notes for /r//s/
ARPAbet S · voiceless alveolar fricative
The /s/ sound is a voiceless alveolar fricative, produced by channelling a narrow stream of air along a groove in the tongue toward the alveolar ridge.
Scoring notes for /s//l/
ARPAbet L · voiced alveolar lateral approximant
The /l/ sound is a voiced alveolar lateral approximant: the tongue tip contacts the ridge behind the upper teeth while air flows around the sides of the tongue.
Scoring notes for /l//θ/ and /ð/
ARPAbet TH, DH · interdental fricatives, voiceless and voiced
"Th" is not one sound but two separate phonemes that happen to share a spelling: the voiceless /θ/ in "think" and the voiced /ð/ in "this".
Scoring notes for /th/Error type predicts performance better than sound identity
The useful way to think about these pages is not “which sounds is the engine good at” but “which errors is it able to represent”. A substitution replaces one phoneme with another phoneme the model already knows, so the engine can score the target low and name what it heard instead. An omission leaves no acoustic evidence where a sound was expected, which is a strong signal. A distortion is neither: it is a badly formed version of the right sound, it is not any other symbol in the inventory, and there is nothing for it to lose against.