What to record, and how
Your recording pipeline moves these scores more than our model does. Background noise is the largest controllable source of scoring error in real deployments, and microphone distance is a close second — both of them decided entirely on your side of the API. This page is the part of the integration worth the most engineering attention, and it is usually the part that gets the least.
Formats, and the one we evaluate on
Six containers decode. Send whatever your recorder produced — converting on the client buys nothing and is one more thing to get wrong on a phone.
| Format | Status | Notes |
|---|---|---|
| WAV | Accepted | 16 kHz mono is the format our evaluation audio is in, and the most predictable choice |
| MP3 | Accepted | Fine. Lossy, but not at a bitrate you would use for speech |
| M4A / MP4 | Accepted | What iOS Safari's MediaRecorder produces. Send it as recorded |
| OGG | Accepted | Usually Opus. Fine |
| WebM | Accepted | What Chrome's MediaRecorder produces. Send it as recorded |
| FLAC | Accepted | Lossless and smaller than WAV, if your pipeline already produces it |
Uncompressed WAV at 16 kHz, mono, is the format our evaluation audio is in, and the one that will behave most predictably. That is a statement about what we measured on, not a requirement: a WebM from Chrome and an MP4 from iOS Safari are both scored, and neither is a second-class input. If you control the pipeline end to end and have a free choice, choose the WAV.
The accepted formats and the current size ceiling are also reported by GET /api/v1/info, which needs no API key. Read them at startup rather than hard-coding them, and you turn a wasted round trip into an immediate, friendlier message.
Sample rate, bit depth, channels
| Property | What we evaluate on | What to do |
|---|---|---|
| Sample rate | 16 kHz | Record at 16 kHz or above. Never upsample from below it - the information is already gone |
| Bit depth | 16-bit PCM | Only a question for WAV. Compressed containers carry their own |
| Channels | Mono | Ask for channelCount: 1. Stereo doubles the bytes and adds nothing |
| Length | A word or a short phrase | A few seconds. text is capped at 500 characters and the upload at 5 MB |
| File size | A few hundred KB | 5 MB is the ceiling the API enforces; 6 MB is refused by the web server as a 413 |
The sample rate is the only row here that can actually hurt you. Recording above 16 kHz costs bytes and nothing else. Recording belowit band-limits the signal, and the high-frequency energy that separates /s/ from /f/ from “th” is the first thing a low rate throws away — telephone-band audio at 8 kHz cuts off around 4 kHz, which is precisely where those sounds live. Upsampling afterwards does not bring them back; it just makes a bigger file with the same missing information.
Recording in a browser
MediaRecorder is enough, and no conversion step is needed. Two details in the snippet below are workarounds we shipped after they cost us time, and both are mobile.
// One channel. On Apple devices, ask for nothing else: Safari's
// noiseSuppression and autoGainControl can hand back audio so quiet it is
// effectively silent. This is the constraint set our own product ships.
const isApple =
/iPad|iPhone|iPod/.test(navigator.userAgent) ||
(navigator.userAgent.includes("Mac") && navigator.maxTouchPoints > 1);
const stream = await navigator.mediaDevices.getUserMedia({
audio: isApple
? { channelCount: 1 }
: {
channelCount: 1,
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
},
});
// iOS Safari records mp4/aac natively; Chrome records webm/opus. Both decode,
// so pick whichever the browser actually supports rather than converting.
const preferred = isApple
? ["audio/mp4;codecs=aac", "audio/mp4", "audio/webm;codecs=opus"]
: ["audio/webm;codecs=opus", "audio/mp4;codecs=aac", "audio/webm"];
const mimeType = preferred.find((type) => MediaRecorder.isTypeSupported(type)) ?? "";
const recorder = mimeType
? new MediaRecorder(stream, { mimeType })
: new MediaRecorder(stream);
// And when you build the blob, use the recorder's own mimeType. Hardcoding
// "audio/webm" mislabels iOS Safari's mp4 and the upload is rejected.
const blob = new Blob(chunks, { type: recorder.mimeType || "" });getUserMedianeeds a secure context and a user gesture. HTTPS or localhost, and a real tap - a recorder that starts on page load will be refused on mobile.- Build the blob from
recorder.mimeType. Never a hardcoded string. iOS Safari records mp4/aac, and labelling that asaudio/webmgets the upload rejected instead of scored. - Cap the take and discard the very short ones. Eight seconds is generous for a word or a phrase; under about half a second is a mis-tap, and an empty upload is refused with a 422 rather than scored.
- Stop the tracks when you are done.
stream.getTracks().forEach((t) => t.stop()), or the recording indicator stays lit and the microphone stays open.
iOS Safari
- Ask for
channelCount: 1and nothing else. Safari'snoiseSuppressionandautoGainControlcan return audio so quiet it is effectively silent. This is the one place where turning the browser's cleanup off is the right call. - Detect iPadOS properly. It sends a desktop Macintosh user agent, so a userAgent test alone will treat an iPad as a Mac. Check
navigator.maxTouchPoints > 1as well. - Prefer
audio/mp4;codecs=aac. It is the native container, and it decodes on our side exactly like everything else. navigator.permissions.querydoes not cover the microphone in Safari. It throws or returns nothing useful, so your permission UI needs a path that does not depend on it.
Chrome mobile and Android
audio/webm;codecs=opusis what you will get, and it is fine. Send it as recorded.- Chrome's WebM carries no duration in its metadata. An
<audio>element will reportInfinityor nothing. If you need the clip length, decode the blob withAudioContext.decodeAudioDataand read the buffer. - Keep the browser's echo cancellation, noise suppression and gain control on. Unlike Safari, they behave, and they earn their keep on a device held at arm's length.
- Test on a cheap device. Microphone quality across the Android range is much wider than across Apple's, and a mid-range tablet in a classroom is a common deployment.
Recording children specifically
Everything above applies to any speaker. What follows is the part that is different because the speaker is five, and it is where most of the avoidable error in a paediatric product comes from.
| The problem | What it looks like in practice | What to do |
|---|---|---|
| Microphone distance | A child leans in, turns away mid-word, or holds a tablet at arm's length | Fix the distance rather than trusting the child to hold one. A wired headset or earbud mic sits at a constant distance and is the single biggest quality win available to you |
| The prompt played aloud | A tablet speaker plays the target word and its own microphone records the tail of it | Gate the recorder until playback has finished, then add a short gap. Do not rely on echo cancellation - on Apple devices you have probably disabled that constraint set already |
| Other voices in the room | A sibling, a parent modelling the word, a classroom | Push-to-talk rather than voice activation, short takes, and a close mic. The engine aligns the target against whatever is in the audio - it does not know which voice is the child's |
| Room noise | Kitchens, cars, classrooms, a television two rooms away | Capture a one-second baseline before the prompt, store it with the attempt, and warn above a threshold you calibrate on your own devices |
| Volume that moves | The same child whispers one word and shouts the next | Keep the browser's automatic gain control on everywhere except Apple devices, and keep takes short so one loud word cannot dominate a long clip |
| Trailing off and false starts | A child says 'um', starts over, or stops halfway | Cap the recording length, discard takes under about half a second, and let the child retry cheaply. A retake costs one call; a bad score costs their confidence |
The prompt playing back into the microphone
This one deserves its own paragraph because it is silent, common, and produces exactly the kind of wrong answer that is hard to trace. A tablet plays the target word aloud to model it, the recorder opens immediately, and the tail of the model production lands in the recording. The engine is then handed audio that contains an adult saying the word correctly, aligned against the same target — and the child's own attempt competes with it. Gate the recorder on the playback ended event, add a short gap after it, and check by listening to stored recordings rather than by reading the code. On Apple devices you have probably disabled the constraint set that would otherwise have partly covered for you.
Capture a room-noise baseline
A second of silence before the prompt, reduced to one number and stored with the attempt, is cheap and pays for itself the first time somebody asks why a child's scores dropped for a week.
// A one-second baseline before the prompt, so a bad score has an explanation.
const ctx = new AudioContext();
const source = ctx.createMediaStreamSource(stream);
const analyser = ctx.createAnalyser();
analyser.fftSize = 2048;
source.connect(analyser);
const buffer = new Float32Array(analyser.fftSize);
const samples = [];
const sample = () => {
analyser.getFloatTimeDomainData(buffer);
let sum = 0;
for (const value of buffer) sum += value * value;
samples.push(Math.sqrt(sum / buffer.length));
};
const timer = setInterval(sample, 50);
setTimeout(() => {
clearInterval(timer);
const rms = samples.reduce((a, b) => a + b, 0) / samples.length;
// Store it with the attempt. Calibrate the threshold on your own devices -
// a number that means "too loud" on a headset does not mean it on a tablet.
onNoiseFloor(rms);
}, 1000);Calibrate the threshold on your own devices. There is no universal number: a level that means “too loud” through a headset does not mean it through a tablet array two feet away. What matters is that the value is recorded, so a run of bad scores can be explained rather than argued about.
What noise does to a score
The engine is given a target text and has to place every expected sound somewhere in the audio. Noise blurs the evidence it does that with. The sounds that disappear first are the quietest ones — the fricatives and stop releases, /s/, /f/, “th”, the release of a final /t/ — which are also, unhelpfully, among the sounds a paediatric product cares most about.
- It does not fail loudly. A noisy recording is still scored. What you get is not an error but a plausible-looking response with false flags on sounds the child produced correctly.
- A false flag costs more than a miss in a child-facing loop. A missed error is one repetition that did not land. A false flag tells a child who was right that they were wrong, which is the thing a practice product cannot afford.
- The numbers we publish come from clean audio. Our benchmark is read-aloud recording in reasonable conditions. Expect worse in a noisy classroom, and validate on your own users' audio before you trust any threshold.
- Aggressive cleanup is not a free fix either. On Apple devices, the browser's own noise suppression and automatic gain control produced audio quiet enough to be useless in our own product. More processing is not automatically more signal.
Pre-launch checklist
Ten things to confirm on real devices before real children use the product. None of them is about the API.
- Every target device records something we can decode. Test the real phones and tablets, not just your laptop. Log
recorder.mimeTypefrom production and check it is what you expect. - The blob is built from
recorder.mimeType. Not a hardcoded string. This is the single most common mobile bug in this pipeline. - iOS Safari has been tested on a real device. Not the simulator, not desktop Safari with a responsive viewport.
- Empty and near-empty takes are caught client-side. Zero bytes, and anything under about half a second.
- The recorder cannot start while prompt audio is playing. Verified by listening to a stored recording, not by reading the code.
- A noise-floor baseline is captured and stored with each attempt. You will want it the first time someone asks why a child's scores dropped for a week.
- The target text sent is captured at record time. Not read from the UI when the upload fires - that is how a prompt advance becomes a mystery score.
- Uploads are capped before they are sent. Check the byte length against the ceiling
GET /api/v1/inforeports rather than hard-coding it. - The key is on your server, not in the page. Rate limits are per key; one extracted key is your whole allowance.
- You have listened to fifty real recordings from real users. Not read the scores - listened. Half an hour of this finds things no metric will.
Unsure whether your audio is good enough?
Tell us how you record - device, browser, microphone, room - when you request access. It is the part of an integration we can help with most concretely, and the part that changes results the most.
Request access