What ArticScore is built for — and what it is not
ArticScore scores a recording against a word or phrase you already know the speaker was asked to say, and reports one row per sound. That single constraint — a known target, in American or British English, scored sound by sound — decides everything this API is good at and everything it is the wrong tool for. Both lists are on this page, and the second one is longer on purpose.
The things it is built for
- Articulation practice with immediate, specific feedback. A child says the target word; the response names the sound that went wrong and the sound that came out instead, in time for it to mean something. That is a different interaction from “try again”.
- A clinician's listening queue. Every attempt scored and ranked worst-first, so a person with forty minutes spends them on the recordings most likely to matter. This is the shape our own measurements support most strongly.
- Home practice that can be counted. Practice between sessions has historically been invisible. Per-attempt rows turn it into a number a clinician can open on a Monday.
- Decoding and phonics against a known text. Phonics objectives are written one grapheme-phoneme correspondence at a time, which is the resolution the response reports at.
- English practice for children learning English. Sound-level feedback on a known word, minimal-pair choices and usable-recording checks, for ESL and English-learning products. The children in our public calibration corpus are English learners themselves.
- Progress on one sound, in one position, over weeks. Per-phoneme rows filter to “initial /r/” and aggregate. A word-level score cannot be taken apart again.
Three of these have a page of their own, written for the people building them:
For speech therapy apps
What changes when feedback is per-phoneme, how to build a clinician worklist from the three tiers, and the limits that decide your architecture.
For reading and literacy apps
How sound-level scoring maps onto phonics, and why a literacy product must never score a speech difference as a reading deficit.
For ESL and English-learning apps
Minimal-pair choices, named substitutions and quality checks for children learning English, and why an accent is not an error.
Use case, coverage, and where it is the wrong tool
The third column is the one worth reading. Every use case ArticScore serves has an adjacent one it does not, and the boundary is usually a single design decision away.
| Use case | What it covers | Where ArticScore is the wrong tool |
|---|---|---|
| Paediatric articulation practice | Per-sound feedback in a child's practice loop; per-attempt records for the adult | Any flow that clears a sound, screens a child, or reaches a conclusion without a person |
| Clinician triage and review | Ranking a week of recordings worst-first so limited listening time lands well | Deciding which recordings a clinician may skip. The engine's clean tier is not a clearance |
| Teletherapy homework | Assigned word lists scored at home, with the attempts visible to the therapist | Unattended assessment. Nothing here is a standardized measure |
| Decoding, phonics and word-list reading | Per-sound accuracy against a target text you supply, with millisecond spans | Reading rate, words correct per minute, or any fluency composite - there is no fluency object |
| Adult English pronunciation practice | Sound-level feedback on a known target word or phrase, in American English or British English (dialect=en-gb, evaluated on adult British speech) | Proficiency scoring or banding of any kind - CEFR, IELTS, PTE, TOEIC. No such output exists |
| Browser and mobile practice games | A single POST per attempt, answered cross-origin, typically in about a second | Open-ended conversation, free retell, or anything without a known target text |
What it is not for
These come up often enough that leaving them ambiguous would waste your time and ours. None of them is a roadmap item we are being coy about; each is a thing the API does not do and, in two cases, a thing we would decline to support if it did.
| Not for | What people mean by it | Why not |
|---|---|---|
| Adult language proficiency and banding | CEFR, IELTS, PTE and TOEIC style scoring; placement and certification | There is no proficiency, fluency or band output in the response, and nothing in the engine is calibrated to any of those scales. Our own utterance-level correlation is the weakest column we publish - if a single whole-utterance proficiency number is your product, this is not where the design effort went |
| Call-centre, aviation and occupational screening | Gating a person's job, licence or placement on a pronunciation score | High-stakes gating of adults is outside everything we have measured, and our own evaluation concluded that systems in this category should not run unsupervised at the operating points we measured. We will not support this use |
| Spanish, French, or any language but English | Scoring pronunciation in another language | The API scores English only: American English by default, or British English with dialect=en-gb. American English is evaluated on children; British English on adult British speech only, as no licensed British child speech exists. Other varieties of English are accepted but not evaluated |
| Oral reading fluency and words correct per minute | Reading rate, hesitation, self-correction, WCPM, fluency composites | There is no fluency object in the response. You get per-phoneme scores and [start_ms, end_ms] spans; composing those into a fluency metric is your product's work and its design is a pedagogical decision. And because the engine is not an open-vocabulary recogniser, a word substituted from outside the passage is not detected |
| Writing, essays, roleplay and open conversation | Scoring written text, dialogue turns, or unscripted speech | One endpoint, two inputs: audio and the target text it should contain. There is no transcription output, no text scoring and no dialogue state. Without a known target there is nothing to align against |
| AAC and dysarthric speech | Scoring the speech of AAC users, or of speakers with dysarthria or apraxia | We have not evaluated the engine on these speakers and would not recommend it for them. It scores against a closed set of phonemes, and productions that are not cleanly any phoneme - which is much of what these populations produce - have no symbol to be scored as |
Every one of these is shaped the same way
A therapy app, a phonics game and a teletherapy homework flow look different to their users and identical to the API. Four steps, in the same order, every time.
- Record against a known target. Your product decides what the speaker should say and captures the audio. Everything downstream depends on that pairing being correct — a mislabelled target scores a correct production as wrong.
- Score one attempt per call. A single POST with the audio and the text. Roughly a second, so the feedback can land while the speaker still remembers the attempt.
- Branch on the tier.
flagis the only tier that should ever interrupt a child.reviewgoes silently to a queue.cleanis silence — no signal, not a verdict. - Put a person at the end. Every use case above ends with a human reading, listening or deciding. Where a design has no such person, the design is outside what our measurements support.
Record → score → tier → human
The last step is not a disclaimer bolted on the end. It is the architecture our own measurements support: the engine is good at ordering a list and specific about what it heard, and it is not reliable enough for any individual verdict to stand on its own. Build the ranking, show the explanation, and let a person decide. A product designed that way is useful today and can be described accurately to a clinician. One that decides unsupervised cannot.
Not sure which of these you are building?
ArticScore is in limited early access, rolled out case by case. Tell us your use case and expected volume - if it is one of the things on the second list, we would rather say so now than in a pilot.
Request access