Quickstart

Your first scored recording

One POST, two fields: a recording, and the word the speaker was asked to say. What comes back is a row for every sound in that word — a score, the phoneme the engine believes was produced instead, a clinical error type and a confidence tier. This page takes you from a key to a scored recording of your own, in curl, Python and Node, and then shows you how to record the audio in a browser.

Before you start

  • A key, issued by arrangement. ArticScore is in limited early access; there is no self-service signup. Keys are issued individually so that integrators hear about schema changes before they happen. Request one through our contact form, mentioning ArticScore, your use case and expected volume.
  • The key goes in an Authorization: Bearer header. This API is key-only by design: a browser session on our own site does not authenticate it, and a key must carry the api:score ability to reach the scoring endpoint.
  • Two endpoints need no key at all. GET /api/v1/info reports the upload ceiling, the accepted formats and the rate limits in force; GET /api/v1/sounds lists every phoneme the engine scores with its ARPAbet and IPA symbols. Both are callable before a key exists.
  • A recording of your own, and the word that was said in it. You send audio plus the target text. You do not need to send phonemes - the expected phoneme sequence is derived from the text and aligned against the audio (for a non-word or a sound on its own, an optional phones field takes them directly).

The one call you can make right now, with no key, reports the limits your integration should enforce on its own side rather than hard-coding them:

Discovery — no key required
curl https://api.speechtherapymagic.com/api/v1/info

Making the call

curl
curl -X POST https://api.speechtherapymagic.com/api/v1/score \
  -H "Authorization: Bearer $ARTICSCORE_API_KEY" \
  -F "audio=@my-recording.wav" \
  -F "text=rabbit"
  • audio is a file field, sent as multipart form data. WAV, MP3, M4A, OGG, WebM and FLAC all decode. Up to 5 MB; a few seconds of speech is a few hundred kilobytes.
  • text is what the speaker was asked to say, up to 500 characters. It is scored as General American English unless you send dialect=en-gb, which scores it against a British English (non-rhotic) reference.
  • There is no engine or model parameter. Two fields go up for a standard call. Everything else - include, dialect, phones, choices - is optional and described in the API reference.
  • Every response carries a request_id. Log it. Quoting it in a support request lets us find the exact scoring run without you sending the recording again.

Reading the response

A child saying “rabbit” as “wabbit”. Three of the five phonemes are trimmed here for length; a real response carries a row for every sound in the target text, flagged or not.

200 OK
{
  "request_id": "9f2b1c84-1f2a-4d5e-9a77-0c3b6e5d4a10",
  "text": "rabbit",
  "overall_score": 61.4,
  "overall_status": "fair",
  "overall_stars": 0,
  "overall_label": "Keep practicing!",
  "words": [
    {
      "word": "rabbit",
      "score": 61.4,
      "status": "fair",
      "stars": 0,
      "label": "Keep practicing!",
      "phonemes": [
        {
          "phone": "R",
          "phone_ipa": "ɹ",
          "score": 41.2,
          "sound_most_like": "W",
          "sound_most_like_ipa": "w",
          "extent": [120, 260],
          "status": "poor",
          "tip": "Try curling your tongue back - it's like a quiet growl!",
          "flagged": true,
          "error_type": "substitution",
          "error_reason": "span+margin",
          "tier": "flag"
        },
        {
          "phone": "AE",
          "phone_ipa": "æ",
          "score": 88.1,
          "sound_most_like": "AE",
          "sound_most_like_ipa": "æ",
          "extent": [260, 390],
          "status": "good",
          "flagged": false,
          "error_type": null,
          "error_reason": null,
          "tier": "clean"
        }
        // ... B, IH, T
      ],
      "syllables": [],
      "tips": [
        { "expected": "R", "got": "W", "tip": "Try curling your tongue back - it's like a quiet growl!" }
      ]
    }
  ],
  "tips": [
    { "expected": "R", "got": "W", "tip": "Try curling your tongue back - it's like a quiet growl!" }
  ],
  "passed": false
}

What to read first

  • words[].phonemes[] is the point of the API. One row per sound in the target text, in order, whether or not it was flagged. Everything above it - overall_score, overall_stars, overall_label - is a roll-up you can show a child, not the layer to build product logic on.
  • tier is the field to branch on. flag, review or clean, per phoneme.
  • sound_most_like is what makes the output an observation rather than a number. It is the engine's best candidate for what was actually produced, and it is what lets you say “the R came out as a W” instead of “try again”.
  • extent is [start_ms, end_ms] - whole milliseconds from the start of the recording, so you can replay or highlight one sound. It is empty when no timing was reported for that phoneme.
  • tip is written for a five-year-old. Render it verbatim where you show it at all; it is present when a phoneme scored below 80 and the engine named a different sound.

Field-by-field definitions for everything above, including the full phoneme inventory and each error_type value, are in the API reference.

Branch on the tier, not on the score

The most common mistake in a first integration is picking a score threshold. The score saturates near the top, so a number that looks decisive today is one whose meaning moves if the calibration does. The tier exists so you do not have to make that choice: it collapses the engine's signals into the three decisions a product actually makes.

One tier per phoneme, on every scored sound
tierWhat it meansWhat your first integration should do
flagThe engine's strongest available signal that something went wrongShow it, gently, or put it at the top of a human's list
reviewThe engine saw something and is not confidentQueue it for a human. Never correct a child on this alone
cleanNo signal was detectedSilence. Not a green tick - see the caveat below

One implementation note. tier, flagged, error_type and tip are not marked required on every phoneme in the specification, while phone, score, sound_most_like, extent and status are. Read the optional ones defensively — phone.tier ?? null rather than phone.tier — and treat an absent tier exactly the way you treat clean: as no signal.

The same call in Python and Node

Python — requests
import os
import requests

with open("my-recording.wav", "rb") as audio:
    response = requests.post(
        "https://api.speechtherapymagic.com/api/v1/score",
        headers={"Authorization": f"Bearer {os.environ['ARTICSCORE_API_KEY']}"},
        files={"audio": ("my-recording.wav", audio, "audio/wav")},
        data={"text": "rabbit"},
        timeout=30,
    )

response.raise_for_status()
result = response.json()

for word in result["words"]:
    for phone in word["phonemes"]:
        # tier is not a required field on every phoneme - read it defensively
        # and treat a missing tier the way you treat "clean": as no signal.
        if phone.get("tier") == "flag":
            print(
                phone["phone"],
                "heard as",
                phone["sound_most_like"],
                phone.get("error_type"),
                phone.get("tip", ""),
            )
Node — fetch, no dependencies
import { readFile } from "node:fs/promises";

const form = new FormData();
// Do not set Content-Type yourself. fetch writes the multipart boundary, and
// overriding the header is the most common reason a first call 422s.
form.append(
  "audio",
  new Blob([await readFile("my-recording.wav")], { type: "audio/wav" }),
  "my-recording.wav",
);
form.append("text", "rabbit");

const response = await fetch("https://api.speechtherapymagic.com/api/v1/score", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.ARTICSCORE_API_KEY}` },
  body: form,
  signal: AbortSignal.timeout(30_000),
});

if (!response.ok) {
  const body = await response.json().catch(() => ({}));
  throw new Error(`${response.status} ${body.error ?? ""} ${body.message ?? ""}`);
}

const result = await response.json();
const flagged = result.words
  .flatMap((word) => word.phonemes)
  .filter((phone) => phone.tier === "flag");

console.log(result.request_id, flagged);

Both set a 30-second client timeout. Scoring normally returns in about a second, so a request still open at thirty is a request that is not coming back; time it out, keep the audio, and let the user retry rather than leaving a child staring at a spinner.

Recording audio in a browser

MediaRecorder is enough. You do not need to convert anything: WebM from Chrome and MP4 from iOS Safari both decode. The sample below is the shape our own product ships, including the two workarounds it took to make mobile behave.

MediaRecorder — record, then upload
// Mono, and let the browser's own cleanup run - except on Apple devices,
// where Safari's noiseSuppression and autoGainControl can hand back audio so
// quiet it scores as silence. This is the constraint set our own product ships.
const isApple =
  /iPad|iPhone|iPod/.test(navigator.userAgent) ||
  (navigator.userAgent.includes("Mac") && navigator.maxTouchPoints > 1);

const stream = await navigator.mediaDevices.getUserMedia({
  audio: isApple
    ? { channelCount: 1 }
    : {
        channelCount: 1,
        echoCancellation: true,
        noiseSuppression: true,
        autoGainControl: true,
      },
});

// iOS Safari records mp4/aac; Chrome records webm/opus. Both decode.
const preferred = isApple
  ? ["audio/mp4;codecs=aac", "audio/mp4", "audio/webm;codecs=opus"]
  : ["audio/webm;codecs=opus", "audio/mp4;codecs=aac", "audio/webm"];
const mimeType = preferred.find((type) => MediaRecorder.isTypeSupported(type)) ?? "";

const recorder = mimeType
  ? new MediaRecorder(stream, { mimeType })
  : new MediaRecorder(stream);

const chunks = [];
recorder.ondataavailable = (event) => {
  if (event.data.size > 0) chunks.push(event.data);
};

recorder.onstop = async () => {
  stream.getTracks().forEach((track) => track.stop());

  // Use the recorder's own mimeType. Hardcoding "audio/webm" mislabels iOS
  // Safari's mp4/aac and the upload is rejected instead of scored.
  const blob = new Blob(chunks, { type: recorder.mimeType || "" });

  const form = new FormData();
  form.append("audio", blob, "recording");
  form.append("text", "rabbit");

  // Posting to your own server, which holds the API key. See the note below.
  await fetch("/api/score-proxy", { method: "POST", body: form });
};

recorder.start();
setTimeout(() => recorder.state === "recording" && recorder.stop(), 8000);
  • Use recorder.mimeType when you build the Blob. Hardcoding audio/webm mislabels iOS Safari's mp4/aac output, and the upload is rejected rather than scored. This one costs teams a day.
  • On Apple devices, drop the audio constraints. Safari's noiseSuppression and autoGainControl can hand back audio so quiet it is effectively silent. Ask for channelCount: 1 and nothing else there; keep the browser cleanup on everywhere else.
  • Discard takes under about half a second. A mis-tap produces a near-empty file, and an empty upload is refused with a 422 rather than scored. Catching it in the browser saves a round trip and a confusing message.
  • Cap the recording length. Eight seconds is plenty for a word or a short phrase, and it keeps you far away from the upload ceiling.
  • Chrome's WebM has no duration in its metadata. If you want the clip length, decode it with AudioContext.decodeAudioData and read the buffer, rather than trusting an <audio> element's duration.

What the microphone is doing matters more than any of this. Mic distance, room noise and a tablet speaker replaying the prompt into its own microphone all move scores, and background noise is the largest controllable source of scoring error we see. That is a page of its own: audio requirements.

Common first-call failures

Almost every first integration hits one of these, and each has a different cause and a different fix.

Errors on POST /api/v1/score
StatusWhat it meansWhat to do
401No usable key, or you authenticated as a browser session rather than a keySend Authorization: Bearer <key>. error is api_key_required when a session was presented
403The key is valid but does not carry the api:score abilityA property of the key, not the request. Retrying will not help - ask us to re-issue it
422Validation. audio missing, empty or over 5 MB, or text missing or over 500 charactersRead errors, which names the field. An empty upload answers 'That recording came through empty. Please try again.'
413Rejected by the web server above 6 MB, before the API saw itThe only non-JSON error here. You are far past the ceiling - check you are not uploading uncompressed audio
429A rate limit. error is rate_limit_exceeded (per minute) or daily_quota_exceeded (per day)Back off for retry_after_seconds. Nothing was scored and nothing was charged against anything else
502Well-formed and authorised; the engine did not answererror is scoring_failed. Retrying after a short pause is reasonable. Quote request_id if it persists

Two more that are worth knowing before you see them. Setting Content-Type yourself on a multipart request breaks the boundary and produces a 422 that looks like a missing file — let your HTTP client write that header. And a recording that contains no speech does not have its own error code: it comes back scored, at the floor, usually with omissions. The full list, including that one, is on errors and troubleshooting.

Where to go next

  • Read every field. The API reference documents the whole response, including the complete phoneme inventory and what each error_type value means.
  • Fix your recording pipeline before you tune anything else. Audio requirements covers microphones, mobile browsers, and recording children specifically. It moves scores more than any threshold you will pick.
  • Handle the errors properly. Errors and troubleshooting has the full status list, the two failures that are really recording bugs, and what to put in a support request.
  • Decide what you are building. Use cases says what the engine is shaped for - and, at more length, what it is not.
  • Read the measurements before you design the interface. The benchmark publishes what we measured, on what sample, with its confidence intervals and the results that go against us.
  • Sort out consent before you record a child. Children's privacy and COPPA covers what is required of you, not of us.

Ready to make the first call?

ArticScore is in limited early access, rolled out case by case. Tell us your use case and expected volume and we will come back to you with a key and the current specification.

Request access

Quickstart questions

How do I get an ArticScore API key?
By asking. ArticScore is in limited early access and keys are issued individually rather than through self-service signup, so that we can talk through what a use case needs and tell integrators about schema changes before they happen. Request one through our contact form at https://speechtherapymagic.com/contact, mentioning ArticScore, your use case and expected volume. The two discovery endpoints, GET /api/v1/info and GET /api/v1/sounds, need no key and can be called before you have one.
What is the minimum request to score a recording?
A POST to https://api.speechtherapymagic.com/api/v1/score with an Authorization: Bearer header and two multipart fields: audio, the recording, up to 5 MB in WAV, MP3, M4A, OGG, WebM or FLAC; and text, the word or phrase the speaker was asked to say, up to 500 characters. Everything else is optional: include for extra blocks, dialect (en-us by default, or en-gb), and phones or choices for targets a dictionary does not spell. There is no engine parameter.
Do I send phonemes or a phonetic transcription?
No. You send the target text in ordinary spelling. The expected phoneme sequence is derived from it and aligned against the audio, and phonemes only appear on the way out - in ARPAbet and IPA, one row per sound.
Should I branch on the score or on the tier?
On the tier. The score saturates near the top, so a raw threshold you pick today is a threshold that quietly changes meaning if the calibration moves. The tier collapses the engine's signals into the three decisions a product actually makes: show it, route it to a human, or do nothing. Note that tier is not marked required on every phoneme in the specification - read it defensively and treat an absent tier the way you treat clean.
Can I call the ArticScore API directly from a browser?
Technically yes - the scoring endpoint answers cross-origin requests from any origin, which is what lets a browser-based game call it. But a key in client-side JavaScript is readable by anyone who opens the network tab, and rate limits are enforced per key, so one extracted key is your whole allowance. For anything past a prototype, record in the browser and post the audio to your own server, which holds the key.
Is there a sample recording I can test with?
Not a downloadable one. We have not resolved the licensing to publish child speech audio, and we would rather ship nothing than ship a clip we do not have clear rights to. Record yourself saying the target word instead - the first call works the same either way, and testing on audio from your own recording pipeline tells you more about your integration than our clip would.