Use cases

What ArticScore is built for — and what it is not

ArticScore scores a recording against a word or phrase you already know the speaker was asked to say, and reports one row per sound. That single constraint — a known target, in American or British English, scored sound by sound — decides everything this API is good at and everything it is the wrong tool for. Both lists are on this page, and the second one is longer on purpose.

The things it is built for

  • Articulation practice with immediate, specific feedback. A child says the target word; the response names the sound that went wrong and the sound that came out instead, in time for it to mean something. That is a different interaction from “try again”.
  • A clinician's listening queue. Every attempt scored and ranked worst-first, so a person with forty minutes spends them on the recordings most likely to matter. This is the shape our own measurements support most strongly.
  • Home practice that can be counted. Practice between sessions has historically been invisible. Per-attempt rows turn it into a number a clinician can open on a Monday.
  • Decoding and phonics against a known text. Phonics objectives are written one grapheme-phoneme correspondence at a time, which is the resolution the response reports at.
  • English practice for children learning English. Sound-level feedback on a known word, minimal-pair choices and usable-recording checks, for ESL and English-learning products. The children in our public calibration corpus are English learners themselves.
  • Progress on one sound, in one position, over weeks. Per-phoneme rows filter to “initial /r/” and aggregate. A word-level score cannot be taken apart again.

Three of these have a page of their own, written for the people building them:

Use case, coverage, and where it is the wrong tool

The third column is the one worth reading. Every use case ArticScore serves has an adjacent one it does not, and the boundary is usually a single design decision away.

Fit, and the boundary next to it
Use caseWhat it coversWhere ArticScore is the wrong tool
Paediatric articulation practicePer-sound feedback in a child's practice loop; per-attempt records for the adultAny flow that clears a sound, screens a child, or reaches a conclusion without a person
Clinician triage and reviewRanking a week of recordings worst-first so limited listening time lands wellDeciding which recordings a clinician may skip. The engine's clean tier is not a clearance
Teletherapy homeworkAssigned word lists scored at home, with the attempts visible to the therapistUnattended assessment. Nothing here is a standardized measure
Decoding, phonics and word-list readingPer-sound accuracy against a target text you supply, with millisecond spansReading rate, words correct per minute, or any fluency composite - there is no fluency object
Adult English pronunciation practiceSound-level feedback on a known target word or phrase, in American English or British English (dialect=en-gb, evaluated on adult British speech)Proficiency scoring or banding of any kind - CEFR, IELTS, PTE, TOEIC. No such output exists
Browser and mobile practice gamesA single POST per attempt, answered cross-origin, typically in about a secondOpen-ended conversation, free retell, or anything without a known target text

What it is not for

These come up often enough that leaving them ambiguous would waste your time and ours. None of them is a roadmap item we are being coy about; each is a thing the API does not do and, in two cases, a thing we would decline to support if it did.

Declined outright
Not forWhat people mean by itWhy not
Adult language proficiency and bandingCEFR, IELTS, PTE and TOEIC style scoring; placement and certificationThere is no proficiency, fluency or band output in the response, and nothing in the engine is calibrated to any of those scales. Our own utterance-level correlation is the weakest column we publish - if a single whole-utterance proficiency number is your product, this is not where the design effort went
Call-centre, aviation and occupational screeningGating a person's job, licence or placement on a pronunciation scoreHigh-stakes gating of adults is outside everything we have measured, and our own evaluation concluded that systems in this category should not run unsupervised at the operating points we measured. We will not support this use
Spanish, French, or any language but EnglishScoring pronunciation in another languageThe API scores English only: American English by default, or British English with dialect=en-gb. American English is evaluated on children; British English on adult British speech only, as no licensed British child speech exists. Other varieties of English are accepted but not evaluated
Oral reading fluency and words correct per minuteReading rate, hesitation, self-correction, WCPM, fluency compositesThere is no fluency object in the response. You get per-phoneme scores and [start_ms, end_ms] spans; composing those into a fluency metric is your product's work and its design is a pedagogical decision. And because the engine is not an open-vocabulary recogniser, a word substituted from outside the passage is not detected
Writing, essays, roleplay and open conversationScoring written text, dialogue turns, or unscripted speechOne endpoint, two inputs: audio and the target text it should contain. There is no transcription output, no text scoring and no dialogue state. Without a known target there is nothing to align against
AAC and dysarthric speechScoring the speech of AAC users, or of speakers with dysarthria or apraxiaWe have not evaluated the engine on these speakers and would not recommend it for them. It scores against a closed set of phonemes, and productions that are not cleanly any phoneme - which is much of what these populations produce - have no symbol to be scored as

Every one of these is shaped the same way

A therapy app, a phonics game and a teletherapy homework flow look different to their users and identical to the API. Four steps, in the same order, every time.

  • Record against a known target. Your product decides what the speaker should say and captures the audio. Everything downstream depends on that pairing being correct — a mislabelled target scores a correct production as wrong.
  • Score one attempt per call. A single POST with the audio and the text. Roughly a second, so the feedback can land while the speaker still remembers the attempt.
  • Branch on the tier. flag is the only tier that should ever interrupt a child. review goes silently to a queue. clean is silence — no signal, not a verdict.
  • Put a person at the end. Every use case above ends with a human reading, listening or deciding. Where a design has no such person, the design is outside what our measurements support.

Record → score → tier → human

The last step is not a disclaimer bolted on the end. It is the architecture our own measurements support: the engine is good at ordering a list and specific about what it heard, and it is not reliable enough for any individual verdict to stand on its own. Build the ranking, show the explanation, and let a person decide. A product designed that way is useful today and can be described accurately to a clinician. One that decides unsupervised cannot.

Not sure which of these you are building?

ArticScore is in limited early access, rolled out case by case. Tell us your use case and expected volume - if it is one of the things on the second list, we would rather say so now than in a pilot.

Request access

Use case questions

What is ArticScore actually built for?
Products that record a child saying a known word and need to know which sound went wrong. The two shapes it serves best are a practice loop that gives a child immediate, specific feedback, and a triage layer that ranks a week of recordings so a clinician's limited listening time lands on the ones most likely to matter. Everything it is good at has an adult reading the output somewhere.
Can ArticScore score CEFR, IELTS, PTE or TOEIC style proficiency?
No. There is no proficiency or band output in the response and nothing in the engine is calibrated to those scales. It returns per-phoneme scores against a target you supply. If a single whole-utterance proficiency number is your product's main output, established language-learning and test-preparation APIs are the category to shop in - and our own published utterance-level correlation is the weakest column on our benchmark, which is the honest reason we point elsewhere.
Does ArticScore support Spanish, French or other languages?
No. The API scores English only: American English by default, or British English with dialect=en-gb. American English has been evaluated on children; British English on adult British speech only, because no British child speech exists under a licence we can use. Other varieties of English are accepted but not evaluated, and patterns that are standard in them - th-fronting, dark-/l/ vocalisation - register as mismatches against either reference.
Can I measure oral reading fluency or words correct per minute?
Not from the response alone, and we would rather say that plainly. There is no fluency object. What you get is a score for each phoneme in the passage you supplied plus a [start_ms, end_ms] span for each one. A words-correct-per-minute or hesitation metric is something your product composes on top of that, and its definition is a pedagogical decision we would not want to make for you. The engine also does not transcribe, so a word substituted from outside the passage will not be detected.
Is ArticScore suitable for AAC users or speakers with dysarthria?
We have not evaluated it on those speakers and we would not recommend it for them. Scoring works by comparing a production against a closed set of phonemes, so a production that is not cleanly any phoneme has no symbol to be scored as - which is much of what these populations produce, and the same structural reason distortions are largely invisible to phoneme-based scoring generally.
Can ArticScore be used to screen a person for a job or a licence?
No, and we will not support it. Gating a person's employment, licensing or placement on an automated pronunciation score is outside everything we have measured, and our own evaluation concluded that systems in this category should not run unsupervised at the operating points we measured. ArticScore is a practice and screening-support tool, not an assessment instrument.