Use case · ESL, ELL and English-learning products

Pronunciation scoring API for ESL and English-learning apps

ArticScore takes a recording of a learner saying a word you chose and returns a score for every sound in it, what a slipped sound most likely came out as, and whether the recording was usable at all. For a product that teaches English to children, that is the difference between “try again” and “your TH came out as an S”.

Where the calibration comes from

The public corpus we calibrate and benchmark on is speechocean762: Mandarin-speaking children reading English, rated sound by sound by trained listeners. p_error and confidence are calibrated on it, and the quality checks were validated on synthetic failures built from it. For a speech-therapy product that is a limitation, because those children are learning English rather than working on a speech sound disorder. For an English-learning product for children, it is closer to your users.

It is still one first language, read speech, and our own measurement. Run recordings from your own learners through the engine before you rely on the default thresholds.

Activity, request, fields to read

One endpoint, POST /api/v1/score. What changes between activities is the input you send and the fields you read back.

POST /api/v1/score: audio plus a text, a pronunciation, or choices
ActivityWhat to sendFields to read
Say a word, get feedback per soundtext = "think"words[].phonemes[]: phone, score, tier, sound_most_like, error_type, tip
Minimal pairs: which one did they say?choices = ["ship", "sheep"] (2–4 candidates; the first is the target)choices.options[] (p and score per candidate), best, said_target
A sound on its own, a syllable, or a made-up wordphones = "TH IY1" (ARPAbet), or IPA with phone_alphabet = "ipa"The same per-sound rows, scored against the pronunciation you sent
Word stress in longer wordsinclude = "syllables,stress"words[].syllables (a score per syllable) and words[].stress (which syllable stood out against the one the dictionary stresses)
Don't score silence or the wrong wordinclude = "quality"quality.usable and quality.reason (no_speech, wrong_word, extra_speech, cut_off, incomplete): ask for another try instead of showing a score
Show the likely alternativesinclude = "alternatives,confidence"phonemes[].alternatives (up to three candidate sounds), p_error and confidence
British English learnersdialect = "en-gb"Scored against a non-rhotic British reference instead of American English
A minimal pair, scored with choices (illustrative values)
POST /api/v1/score
Authorization: Bearer <your key>
Content-Type: multipart/form-data

audio   = <file>
choices = ["ship", "sheep"]   # no text: the first choice is the target

# response (excerpt): the scores describe "ship"; choices says "sheep" was heard
"text": "ship",
"choices": {
  "options": [
    { "text": "ship",  "phones": ["SH", "IH1", "P"], "p": 0.22, "score": 51.0, "is_target": true },
    { "text": "sheep", "phones": ["SH", "IY1", "P"], "p": 0.78, "score": 86.5, "is_target": false }
  ],
  "best": 1,
  "said_target": false,
  "p_target": 0.22
}

Every field is defined in the API reference, and the first working call is in the quickstart. Stress scoring is new and we have not published an accuracy figure for it yet; the same goes for British English, which has been evaluated on adult British speech only.

How it differs from general language-learning engines

General pronunciation APIs for language learning are mature and well engineered. Many support several languages, publish self-serve pricing, and return fluency, prosody and proficiency scores. If your product needs any of those, they are the category to evaluate, and we would rather say so here than in a pilot. Where ArticScore is different is in what it is built for:

  • Fine-tuned on children's speech. The acoustic model is adult-pretrained with its top layers fine-tuned on child speech, which lowered false flags on correct productions against our own pre-fine-tuning baseline. If your learners are children, that is the population the tuning was for.
  • Named substitutions at the sound level. A flagged sound comes back with what was most likely produced instead (sound_most_like) and an error type: substitution, omission or unclear. For a learner, “/θ/ came out as /s/” is the actionable part.
  • Inputs built for practice design. choices answers “ship or sheep?” directly, and phones scores an isolated sound, a CV syllable or a non-word that no dictionary spells.
  • Quality checks before a score. include=quality says when a recording has no speech, a different word, extra speech or a cut-off, so a child is asked to try again instead of being shown a meaningless number.
  • No audio retention. Audio sent to the API is scored and deleted, never kept and never used for training. See children's privacy.
  • Published evidence, including what goes against us. Our accuracy measurements, sample sizes and confidence intervals are on the benchmark page. We do not publish measurements of other vendors and do not claim to be more accurate than any of them.

Compare engines on recordings from your own learners, scored against labels from someone qualified. Our guide to choosing a pronunciation scoring API sets out how.

What it does not do

Stated up front
Not forWhat people mean by itWhy not
Fluency, rhythm and intonationSpeaking rate, pauses, sentence melody, prosody scoresNot exposed. Our engine's fluency and prosody measures did not meet our own validation bar, so they ship nowhere, including the API
Proficiency bandsCEFR, IELTS, TOEFL, PTE or TOEIC style levels; placementThere is no proficiency or band output and nothing in the engine is calibrated to those scales
Other languagesScoring Spanish, Mandarin or a learner's first languageEnglish only: American English by default, British English with dialect=en-gb
Open conversationFree speech, role-play, transcription of whatever was saidEvery request scores a known target: a text, a pronunciation, or a short list of choices
Accent ratingsA score for how native-like someone soundsNot what the output means, and not something we would help build. The scores describe sounds against a reference, not a person

A practice loop that survives being wrong

  • Correct on flag only. Keep review for a teacher or a later round, and treat clean as no signal, never as proof a sound is mastered.
  • Check quality before showing a number. If quality.usable is false, ask for another try.
  • Never gate progress on one score. A false flag should cost the learner nothing: no lost streak, no lost points.
  • Show the sound, not a grade. “Your TH sounded like S” with a tip helps a learner. A percentage next to their name does not.

Used every day in our own app

The same engine scores every attempt in SpeechTherapyMagic, including the practice we offer for children learning English. Trying it there as a user is the quickest way to judge whether this kind of feedback fits your product.

Building for English learners?

Create an account and request access from the developer portal. Access is approved case by case; once approved you create your own keys and see every request you send.

Rather talk it through first?

Tell us about your learners, your use case and your expected volume, and we will come back to you.

Request access

ESL and English-learning apps: frequently asked questions

Can ArticScore be used in an ESL or English-learning app for children?
Yes, when the app knows what the learner was asked to say. Every request scores a known target: a word or phrase, an explicit pronunciation, or a list of two to four choices such as a minimal pair. The response scores each sound, names what a flagged sound most likely came out as, and can check whether the recording was usable at all. It does not transcribe open conversation and does not produce fluency or proficiency scores.
What data is the confidence calibration based on?
p_error and confidence are calibrated on speechocean762, a public corpus of Mandarin-speaking children reading English, rated by trained listeners. The quality checks were validated on synthetic failures built from the same recordings. That is one first language and read speech, so validate on recordings from your own learners before you rely on the defaults.
Does it support British English?
Yes, with dialect=en-gb, which scores against a non-rhotic British reference. It has been evaluated on adult British speech only, because no British child speech exists under a licence we can use. American English is the default and the dialect we have evaluated on children. Other varieties of English are not supported as a reference.
Does it score fluency, intonation or give CEFR/IELTS levels?
No. The engine's fluency and prosody measures did not meet our own validation bar, so they are not exposed anywhere, and there is no proficiency or band output. If a single proficiency number is what your product needs, established language-learning APIs are built for that and are the category to evaluate.
Will it penalise a learner's accent?
It compares each sound with the reference pronunciation, so a sound produced the way the learner's first language makes it can score lower even when the word was perfectly understandable. That is why we recommend correcting only on the flag tier, framing feedback as practice towards being understood, and never presenting the output as an accent rating.
How do we get access, and what does it cost?
Create a SpeechTherapyMagic account and request access from the developer portal. Access is approved case by case; an approved account creates its own keys and sees its own request log. There is no public price list yet: pricing is arranged directly, based on your use case and volume.