Pronunciation scoring API for ESL and English-learning apps
ArticScore takes a recording of a learner saying a word you chose and returns a score for every sound in it, what a slipped sound most likely came out as, and whether the recording was usable at all. For a product that teaches English to children, that is the difference between “try again” and “your TH came out as an S”.
Where the calibration comes from
The public corpus we calibrate and benchmark on is speechocean762: Mandarin-speaking children reading English, rated sound by sound by trained listeners. p_error and confidence are calibrated on it, and the quality checks were validated on synthetic failures built from it. For a speech-therapy product that is a limitation, because those children are learning English rather than working on a speech sound disorder. For an English-learning product for children, it is closer to your users.
It is still one first language, read speech, and our own measurement. Run recordings from your own learners through the engine before you rely on the default thresholds.
Activity, request, fields to read
One endpoint, POST /api/v1/score. What changes between activities is the input you send and the fields you read back.
| Activity | What to send | Fields to read |
|---|---|---|
| Say a word, get feedback per sound | text = "think" | words[].phonemes[]: phone, score, tier, sound_most_like, error_type, tip |
| Minimal pairs: which one did they say? | choices = ["ship", "sheep"] (2–4 candidates; the first is the target) | choices.options[] (p and score per candidate), best, said_target |
| A sound on its own, a syllable, or a made-up word | phones = "TH IY1" (ARPAbet), or IPA with phone_alphabet = "ipa" | The same per-sound rows, scored against the pronunciation you sent |
| Word stress in longer words | include = "syllables,stress" | words[].syllables (a score per syllable) and words[].stress (which syllable stood out against the one the dictionary stresses) |
| Don't score silence or the wrong word | include = "quality" | quality.usable and quality.reason (no_speech, wrong_word, extra_speech, cut_off, incomplete): ask for another try instead of showing a score |
| Show the likely alternatives | include = "alternatives,confidence" | phonemes[].alternatives (up to three candidate sounds), p_error and confidence |
| British English learners | dialect = "en-gb" | Scored against a non-rhotic British reference instead of American English |
POST /api/v1/score
Authorization: Bearer <your key>
Content-Type: multipart/form-data
audio = <file>
choices = ["ship", "sheep"] # no text: the first choice is the target
# response (excerpt): the scores describe "ship"; choices says "sheep" was heard
"text": "ship",
"choices": {
"options": [
{ "text": "ship", "phones": ["SH", "IH1", "P"], "p": 0.22, "score": 51.0, "is_target": true },
{ "text": "sheep", "phones": ["SH", "IY1", "P"], "p": 0.78, "score": 86.5, "is_target": false }
],
"best": 1,
"said_target": false,
"p_target": 0.22
}Every field is defined in the API reference, and the first working call is in the quickstart. Stress scoring is new and we have not published an accuracy figure for it yet; the same goes for British English, which has been evaluated on adult British speech only.
How it differs from general language-learning engines
General pronunciation APIs for language learning are mature and well engineered. Many support several languages, publish self-serve pricing, and return fluency, prosody and proficiency scores. If your product needs any of those, they are the category to evaluate, and we would rather say so here than in a pilot. Where ArticScore is different is in what it is built for:
- Fine-tuned on children's speech. The acoustic model is adult-pretrained with its top layers fine-tuned on child speech, which lowered false flags on correct productions against our own pre-fine-tuning baseline. If your learners are children, that is the population the tuning was for.
- Named substitutions at the sound level. A flagged sound comes back with what was most likely produced instead (
sound_most_like) and an error type: substitution, omission or unclear. For a learner, “/θ/ came out as /s/” is the actionable part. - Inputs built for practice design.
choicesanswers “ship or sheep?” directly, andphonesscores an isolated sound, a CV syllable or a non-word that no dictionary spells. - Quality checks before a score.
include=qualitysays when a recording has no speech, a different word, extra speech or a cut-off, so a child is asked to try again instead of being shown a meaningless number. - No audio retention. Audio sent to the API is scored and deleted, never kept and never used for training. See children's privacy.
- Published evidence, including what goes against us. Our accuracy measurements, sample sizes and confidence intervals are on the benchmark page. We do not publish measurements of other vendors and do not claim to be more accurate than any of them.
Compare engines on recordings from your own learners, scored against labels from someone qualified. Our guide to choosing a pronunciation scoring API sets out how.
What it does not do
| Not for | What people mean by it | Why not |
|---|---|---|
| Fluency, rhythm and intonation | Speaking rate, pauses, sentence melody, prosody scores | Not exposed. Our engine's fluency and prosody measures did not meet our own validation bar, so they ship nowhere, including the API |
| Proficiency bands | CEFR, IELTS, TOEFL, PTE or TOEIC style levels; placement | There is no proficiency or band output and nothing in the engine is calibrated to those scales |
| Other languages | Scoring Spanish, Mandarin or a learner's first language | English only: American English by default, British English with dialect=en-gb |
| Open conversation | Free speech, role-play, transcription of whatever was said | Every request scores a known target: a text, a pronunciation, or a short list of choices |
| Accent ratings | A score for how native-like someone sounds | Not what the output means, and not something we would help build. The scores describe sounds against a reference, not a person |
A practice loop that survives being wrong
- Correct on
flagonly. Keepreviewfor a teacher or a later round, and treatcleanas no signal, never as proof a sound is mastered. - Check quality before showing a number. If
quality.usableis false, ask for another try. - Never gate progress on one score. A false flag should cost the learner nothing: no lost streak, no lost points.
- Show the sound, not a grade. “Your TH sounded like S” with a tip helps a learner. A percentage next to their name does not.
Used every day in our own app
The same engine scores every attempt in SpeechTherapyMagic, including the practice we offer for children learning English. Trying it there as a user is the quickest way to judge whether this kind of feedback fits your product.
Building for English learners?
Create an account and request access from the developer portal. Access is approved case by case; once approved you create your own keys and see every request you send.
Rather talk it through first?
Tell us about your learners, your use case and your expected volume, and we will come back to you.
Request access