JavaScript Screen Capture API Table

PieceWhat it doesField note
getDisplayMedia({video})The askuser gesture + the browser source picker - the picker IS the permission, re-shown every call
displaySurfaceThe hintprefers browser/window/monitor; the user overrides - read track.getSettings().displaySurface for the truth
MediaStreamThe resultordinary stream capital: video element, MediaRecorder, WebRTC replaceTrack, WebCodecs all consume it unchanged
selfBrowserSurfaceThe mirror guardexclude kills the infinite echo of a tab capturing its own preview; surfaceSwitching lets users change mid-share
track.onendedThe stop signalbrowser stop-sharing button, closed tab, unplugged monitor - the single reliable end-of-share event; build UI on it
audio: trueThe screen soundtab audio (browser picker) or system audio (monitor, Chromium/Windows); mic is a separate getUserMedia + WebAudio mix
captureControllerThe focus dialconditional focus (Chromium): decide whether starting a share steals focus from the shared surface
support matrixThe platform truthall engines, HTTPS only; Chromium richest (system audio, switching), Safari/Firefox core; iframes need allow="display-capture"
Reference: the MDN Screen Capture API. Consent lives in the browser source picker, the user holds the stop button, and the MediaStream comes back shaped like a camera - which is why everything downstream (recorder, peer connection, encoder) needs no new plumbing.
Bottom line: anchor the whole UI on the video track ended event - it fires for stop-sharing, closed tabs and unplugged monitors alike, and nothing else does. Default selfBrowserSurface: exclude so your own preview never feeds back into itself, keep screen audio and mic audio as separate sources mixed in WebAudio, and never cache a screen permission (there is no persistent one to cache - the picker re-runs and the user can always walk).
Related tools: the getUserMedia table (the camera/mic sibling), the WebRTC table (shipping the capture peer-to-peer), the WebCodecs table (frame-level encoding of the capture), the visibility table (what pauses when you share), and the permissions table (the camera permission this API deliberately skips).

The Screen Capture API (getDisplayMedia) is how a page asks the user to share a screen, window, or tab: navigator.mediaDevices.getDisplayMedia({video: true}) opens the browser's own source picker - the user picks one surface, and you receive a MediaStream shaped exactly like a camera stream. That familiar shape is the design win: everything that consumes camera input (a video element, MediaRecorder, WebRTC peer connection, WebCodecs encoder) consumes screen capture unchanged.

Bottom line: with screen capture, the browser runs the consent flow, not you. There is no persistent permission to script - every call shows the picker, the user choice IS the grant, and the user can end it at any second from browser chrome (the stop-sharing button). So the professional pattern anchors everything on the track's 'ended' event: that event is the single reliable signal the share is over, whether because the user clicked stop, closed the tab being shared, or the source disappeared.

The traps are audio and the mirror. Screen audio is not microphone audio: getDisplayMedia({audio: true}) gets tab or system sound, and a voiceover needs a separate getUserMedia mic call mixed in WebAudio. And capturing the tab your preview plays in creates an infinite mirror - the feedback echo that ruined a thousand screen shares - which is why selfBrowserSurface: 'exclude' exists and should be your default, along with systemAudio hints and conditional focus (captureController) where supported.

How to use

  1. Ask with hints: const stream = await navigator.mediaDevices.getDisplayMedia({video: {displaySurface: 'browser'}, audio: false}); - displaySurface only PREFERS tab/window/monitor; the user overrides freely and the actual choice comes back on track.getSettings().displaySurface.
  2. Apply the privacy dials: selfBrowserSurface: 'exclude' (no capturing your own tab - kills the mirror loop), preferCurrentTab: true for share-your-tab flows, surfaceSwitching: 'include' to let the user change source mid-share, monitorTypeSurfaces to keep HR tabs out of the monitor option.
  3. Wire the end signal: stream.getVideoTracks()[0].addEventListener('ended', () => stopUi()) - the browser stop button, closing the shared tab, or unplugging a monitor all land here; there is no other reliable end-of-share event.
  4. Add audio honestly: audio: true captures tab sound (browser picker) or system sound (monitor picker, Chromium); a microphone is a SEPARATE getUserMedia call - mix both in WebAudio before MediaRecorder if the recording needs voiceover.
  5. Ship it downstream: preview in a muted <video>, record with MediaRecorder or encode frames with WebCodecs, or replace the camera track on an RTCPeerConnection with sender.replaceTrack - the stream is ordinary MediaStream capital everywhere.

Frequently asked questions

Why is there no permission prompt like the camera has?

Because the source picker IS the permission - and it re-runs every time. A camera grant makes sense: the device list is stable and the page can justify persistent access. Screens are not stable (monitors come and go) and the blast radius is huge (whatever happens to be on the monitor), so the spec made consent per-session: the user picks the exact surface, sees it highlighted, and can stop with one browser-chrome click at any moment. Consequences for design: never cache a 'has screen permission' state (there is no such thing), expect the picker on every start(), and build your UI around the ended event rather than around a promised duration. For iframes, the parent must delegate via allow="display-capture" - the policy replaces the top-level user's implicit trust.

How do I record system audio or make a screen recording with voiceover?

Split the two audio problems. Screen audio: getDisplayMedia({audio: true}) - when the user shares a browser TAB you get that tab's audio; on Windows/Chromium, sharing a whole MONITOR can also expose a system-audio option in the picker (Chrome's checkbox), and Firefox/Safari largely do not do system audio at all - so feature-test, do not assume. Microphone audio never comes from getDisplayMedia: call getUserMedia({audio: true}) separately and mix the two streams with a WebAudio graph (createMediaStreamSource from each, one destination, hand that to MediaRecorder). The recorder itself does not care where tracks came from - a combined MediaStream with the display video track plus mixed audio track records exactly like a camera recording. Watch audio track 'ended' too: users can mute tab audio independently of the video.

What is the infinite mirror problem and how do I prevent it?

Share a tab that displays its own preview and light enters a loop: frame N appears in frame N+delta, smaller and brighter, until the screen fills with the recursive echo and the audio (if any) shrieks. It happens whenever the captured surface can show the capture's own output - your share button's tab previewing the capture, or a video-call grid showing every participant including the sharer. The API-level fix: selfBrowserSurface: 'exclude' removes your own tab from the picker; for call grids, preferCurrentTab plus hiding the self-view during capture, or a one-frame delay check. There is also a UX variant with no API fix: a user shares a window showing your site. You cannot prevent that one - but you can warn in the preview lag notice ('viewers see this ~1s late') which kills the most common confusion, and keep your own surface excluded so the loop at least never starts from your page.

Which browsers support screen capture, and where are the sharp edges?

Core getDisplayMedia with a picker works in all three engines - Chromium (richest: tab/window/monitor, system audio on Windows, surface switching, conditional focus via captureController), Firefox (tab/window/monitor, solid core, no system audio), Safari 13+ (screen and window; historically no tab capture on macOS until recent versions, no system audio). The edges that bite: iframe embedding needs allow="display-capture" or the call rejects; insecure contexts never get the API (HTTPS mandatory); pointer capture and keyboard events do not follow the remote surface; and the user can switch surfaces mid-share (surfaceSwitching) - so re-read track.getSettings() when settings change instead of caching displaySurface at start. Content protection (DRM/black windows) shows as black frames - that is the OS refusing, not your code, and it is unfixable from JavaScript.

Related tools