JavaScript WebCodecs API Table
| Piece | What it does | Field note |
|---|---|---|
VideoEncoder / VideoDecoder | The codec pair | configure() then encode()/decode() frame by frame - the browser media pipeline with the lid off |
EncodedVideoChunk | The packet | type key/delta + timestamp in MICROSECONDS + byteLength - decoder input equals encoder output |
VideoFrame | The raw pixel | copyTo() for bytes; close() every frame immediately - native/GPU memory, a 1080p60 loop OOMs in seconds |
isConfigSupported() | The capability probe | codec string carries profile+level (avc1.42001E) - always probe before configure, hardware varies |
key frame discipline | The decoder entry | deltas mean nothing without their key - start runs at key chunks, force GOP via encode(frame, {keyFrame:true}) |
AudioEncoder / AudioData | The audio sibling | same shape, f32-planar data - opus/mp3/aac depending on platform licensing |
encodeQueueSize + dequeue | The backpressure | encoder slower than your source - drop frames (realtime) or await dequeue (offline) or leak memory |
support truth | The matrix | Chromium 94+/Safari 16.4+/Firefox 130+ but per-codec per-platform - isConfigSupported is the only honest check, never UA |
WebCodecs is the browser's media pipeline with the lid off: instead of a <video> element or MediaRecorder deciding everything, you hand VideoDecoder encoded chunks and get VideoFrame pixels back (or feed VideoEncoder frames and get EncodedVideoChunks out) - frame by frame, with real timestamps, configurable codecs, and no DOM. This is the machinery screen recorders, video editors, WebRTC-ish pipelines and cloud gaming clients are built from, and it is the difference between recording what a element plays and processing media yourself.
Bottom line: with the lid off comes the plumbing. WebCodecs encodes and decodes - it does not mux, demux, or make files. Feeding an MP4 means you (or a library like mp4box.js) extract the chunks first; producing a file means writing the container yourself. The trade is total control: exact frame timing (microseconds everywhere - the classic sync bug is mixing in millisecond rAF timestamps), forced key frames, per-frame encode decisions, and access to the GPU-backed hardware codecs the browser itself uses.
The engineering risks are memory and flow, not support. VideoFrame objects hold native/GPU memory - close() every frame the moment you finish with it, or a 1080p60 loop OOMs the renderer in seconds. And the encoder is slower than your frame source: encodeQueueSize grows unless you respect it (drop, or await the dequeue event). Support is now broad - Chromium since 94, Safari since 16.4, Firefox since 130 - but per-codec reality varies by platform and hardware, so isConfigSupported() before every configure() is the only honest capability check.
How to use
- Probe then configure: const s = await VideoEncoder.isConfigSupported({codec: 'avc1.42001E', width: 1280, height: 720, bitrate: 2_500_000}); if (s.supported) enc.configure(s.config); - the codec string carries profile and level (42001E = baseline 3.1), which is why generic strings like 'h264' do not exist here.
- Feed the encoder: await enc.encode(frame, {keyFrame: isKey}); - frame is a VideoFrame from canvas, drawImage source, or a decoder. Force key frames yourself every N frames (your GOP length) - the encoder will not guess your streaming or seek needs.
- Collect the output: new VideoEncoder({output: (chunk, meta) => sink.push(chunk)}) - chunks are EncodedVideoChunk objects with type ('key' or 'delta'), timestamp (microseconds) and byteLength; the decoder's input is exactly this shape, which is what makes transcode pipelines symmetric.
- Decode symmetrically: dec.configure({codec: ...}); dec.decode(chunk) - deltas are only valid after their key frame landed; feed out of order and frames drop silently. drain with await enc.flush() before closing.
- Respect the queue: watch enc.encodeQueueSize and listen for the dequeue event - when the encoder falls behind, drop frames (a video call) or buffer deliberately (an offline render); ignoring the queue converts backpressure into memory pressure. And call frame.close() on every frame, every time.
Frequently asked questions
How is WebCodecs different from MediaRecorder?
Abstraction depth. MediaRecorder hands you a muxed webm/mp4 blob of whatever a MediaStream plays - simple, opaque, real-time-only, and its bitrate and keyframe cadence are the browser's secrets. WebCodecs gives you the same codecs at the frame level: you choose the codec string and bitrate, force key frames, read exact timestamps, decide what each frame becomes - but you also own demuxing the input and muxing the output, because the API produces and consumes raw chunk streams, not files. The honest decision rule: recording a stream for download - MediaRecorder; processing, transcode, streaming at controlled GOP, or frame-accurate capture (canvas animation, game, screen compositor) - WebCodecs, plus a container library for the file wrapper.
Why does my decoder drop frames or error with 'a key frame is required'?
Because H.264-style codecs are differential: delta chunks describe only what changed since the previous frame, so a decoder that starts mid-GOP has nothing to anchor to and throws or drops until the next key frame arrives. The fix is stream discipline: begin every decode run at a key chunk (type === 'key'), never reorder, and if you join a live stream late you must wait for - or request via a keyframe signal - a fresh key frame. When encoding, remember you control GOP length: encode(frame, {keyFrame: true}) every N frames; shorter GOPs cost bitrate (key frames run 5-10x delta size) but make seeks and late-joins instant. The error message is the spec being helpful - it means your chunk ordering, not your codec string.
What does the dequeue event and encodeQueueSize actually control?
Backpressure. encode(frame) is asynchronous under the hood: frames queue while the hardware encoder works, and encodeQueueSize tells you how deep that queue is. If you encode a 60fps canvas without looking, an encoder that can only do 30fps turns your frame queue into a memory leak; the queue growing is the signal to shed load. The patterns: real-time paths (calls, screen share) drop the oldest frame when encodeQueueSize exceeds a threshold - latency beats completeness; offline paths (exports) await the dequeue event (or flush()) between frames and just run slower. The same discipline applies on decode with decodeQueueSize. This is the API teaching you what <video> used to hide.
Which codecs can I actually use, and how do I check?
Never by user agent - by probing. Codec support is a matrix of engine (Chromium 94+, Safari 16.4+, Firefox 130+), platform, and hardware: H.264 (avc1....) is near-universal, VP8/VP9 broad, AV1 wherever hardware or software decode exists, and audio brings opus, mp3, aac (platform-dependent licensing). The strings carry profile and level on purpose - avc1.42001E (baseline 3.1) plays everywhere; a high-profile string may fail on old hardware. So the pattern is: await VideoEncoder.isConfigSupported({codec, width, height, bitrate, framerate}) and use its returned config (it may be adjusted); treat configure() without the probe as a crash you shipped. For hardware acceleration questions, hardwareAcceleration: 'prefer-hardware' asks and the probe answers honestly.