JavaScript Web Speech API Table
| Piece | What it does | Field note |
|---|---|---|
speechSynthesis | The TTS engine | speak() queues utterances; cancel()/pause()/resume() - universal support, zero permission |
SpeechSynthesisUtterance | The request object | text + voice + rate/pitch/volume multipliers - one object per chunk, queue serializes |
getVoices() | The async list | empty until voiceschanged fires in Chrome - cache on the event, pick by lang prefix |
onboundary | The progress ticker | word/char offsets for karaoke highlighting - engine support varies, degrade to a progress bar |
SpeechRecognition | The ASR half | webkit prefix in Chrome/Edge/Safari 14.1+; continuous + interimResults; Firefox never shipped it |
onresult / onerror | The transcript stream | final + interim results; errors arrive as events (not-allowed, no-speech, network) |
the mic gate | The permission ask | start() inside a user gesture; Chromium transcribes server-side - dies offline, bytes leave the device |
support matrix | The honest map | synthesis everywhere since ~2016; recognition Chromium + Safari only - detect it, keep a text field fallback |
The Web Speech API is two APIs wearing one name, and their fortunes could not be more different. SpeechSynthesis (text to speech) is universal: every major browser and every phone speaks with speechSynthesis.speak(new SpeechSynthesisUtterance(text)) - no permission, no network dependency you control, voices from the OS. SpeechRecognition (speech to text) is the scarce half: Chrome and Edge expose it (as webkitSpeechRecognition), Safari 14.1+ joined with the same prefix, Firefox never shipped it - and in Chromium the audio is transcribed server-side, so it dies offline and guarantees nothing about where bytes travel.
Bottom line: the two halves have opposite failure modes. Synthesis is reliable but stateful - the voice list arrives empty until voiceschanged fires, long utterances get truncated mid-sentence by some engines, and iOS requires the first speak() inside a user gesture. Recognition is capricious but simple to start - one tap, one mic permission, onresult hands you transcripts - and the engineering risk is not support (feature-detect is one line) but the network hop: audio leaves the device in Chrome, so never promise privacy the implementation does not have.
The design asymmetry is the guide: use synthesis freely - read-aloud buttons, form readback, accessibility polish - it just works everywhere. Gate recognition behind a feature check and a clear fallback (a text field), and tell users where their audio goes. The API hands you the flags (window.SpeechRecognition || window.webkitSpeechRecognition, speechSynthesis.getVoices()) - honest handling of the empty voice list and the absent recognizer is your job.
How to use
- Speak text: const u = new SpeechSynthesisUtterance('Hello'); u.rate = 1; speechSynthesis.speak(u) - works on click, on load (desktop), no permission. Rate 1 is that voice's normal pace, pitch 1 its normal pitch; 0.1-10 and 0-2 are multipliers, not absolutes.
- Get voices the async-safe way: call speechSynthesis.getVoices() once AND listen for voiceschanged - Chrome populates the list late, so the array is empty on first paint. Cache voices when the event fires, pick by lang prefix (en-US), and let the user override.
- Chunk long text: split paragraphs into sentence-sized utterances and queue them - speak() serializes naturally. Some engines clip a single long utterance; boundaries (onboundary gives word/char offsets) let you build karaoke highlighting.
- Add recognition with a feature gate: const SR = window.SpeechRecognition || window.webkitSpeechRecognition; if (!SR) show the fallback. Set continuous and interimResults, call start() in a click handler, and handle onerror (not-allowed, no-speech, network) explicitly - recognition errors are events, not exceptions.
- Respect the gates: recognition.start() triggers the mic permission prompt - never fire it on load. On iOS, call speak() inside the tap handler the first time (the autoplay policy applies to speech). Stop synthesis with cancel() on navigation or you get a page that keeps talking.
Frequently asked questions
Why is getVoices() returning an empty array?
Because the voice list is populated asynchronously. Chrome loads its voice inventory after page load and fires voiceschanged when ready - the classic bug is reading getVoices() in an initializer and caching an empty array forever. The fix is an event-driven cache: read it once immediately (some engines are synchronous), and also listen for voiceschanged and rebuild your voice picker then. Safari and Firefox tend to be synchronous, so the bug hides in testing on one engine and breaks in Chrome. Corollary: voice objects go stale across the event - always re-pick from the fresh list rather than storing a voice reference from the empty first call.
Does speech recognition work in Firefox, or offline?
No and no, in practice. Firefox has never shipped SpeechRecognition; Chrome, Edge and Safari 14.1+ have it (Safari also uses the webkit prefix). And the Chromium implementation transcribes server-side - your audio is recorded, sent to Google's servers, and the transcript comes back - so it requires a network connection and has real privacy weight: an HTTPS page can capture speech, but where the bytes travel is decided by the browser vendor, not by you. For offline or private-by-design transcription, the honest options are native apps or local models (Whisper-class) running outside the browser sandbox. Feature-detect, disclose, and keep the text-input fallback one tap away.
Why does my long text stop or get cut off mid-read?
Two known culprits. First, per-utterance limits: several engines (notably Chrome's remote voice path) truncate very long single utterances - the fix is chunking: split on sentence boundaries, speak each chunk, and chain via the onend event or just call speak() for each in order (the queue serializes). Second, lifecycle: mobile browsers suspend synthesis when the tab hides or the page unloads, and iOS stops speech after screen lock - a read-aloud feature needs visibilitychange handling (pause on hide, resume on show) and a visible stop button. cancel() clears the queue; pause()/resume() exist for mid-sentence control.
Is speech synthesis free to use, or does it need permission?
Free and universal - synthesis has no permission gate in any browser, works from file:// and in every modern engine including mobile Safari and Firefox (each ships its own voice inventory from the OS). That asymmetry with recognition is by design: speaking TO the user is harmless; listening requires the mic permission prompt and a secure context. The practical design rule: read-aloud is a safe default feature (add it anywhere text is long), while voice input is a progressive enhancement - detect it, explain it, and never make it the only path. One more trap: speechSynthesis.speak() on iOS must be called during a user gesture the first time, or the audio is silently dropped - the same autoplay policy that gates video gates speech there.