JavaScript WebXR Device API Table

PieceWhat it doesField note
navigator.xr.isSessionSupportedThe probecapability-gate the Enter VR button; check immersive-vr and immersive-ar separately, support is not symmetric
requestSession(mode)The session askclick-gated + secure context; declare reference spaces in optionalFeatures or requestReferenceSpace rejects
requestReferenceSpaceThe coordinate systemlocal (seated) / local-floor (standing) / bounded-floor (guardian polygon) / unbounded - locomotion design inherits it
session.requestAnimationFrameThe XR loopthe headset clock (72-120Hz), signature (time, XRFrame) - a slow frame is a wobbling world, not a dropped pixel
frame.getViewerPoseThe predicted posewhere the head WILL be at scanout, per view per eye - XRRigidTransform position + quaternion to your camera
session.inputSourcesThe controllersunified controllers/hands/taps with targetRay + grip spaces; selectstart/squeeze events, gamepad hides inside
XRWebGLLayerThe render contractmakeXRCompatible + updateRenderState({baseLayer}) - WebGL is the only path; the framebuffer belongs to the compositor
support matrixThe platform truthQuest browser is the real target, Android AR varies, desktop rare, iOS near-zero - probe and keep a flat fallback
Reference: the MDN WebXR Device API. The immersive web inverts the usual contract: the session owns the clock, the framebuffer and the exit - you render predicted poses on the device schedule and design for a session that ends the moment the headset comes off.
Bottom line: motion sickness is a performance bug - keep render work under one frame budget on the device clock and push logic to workers. Gate the UI on isSessionSupported, declare optionalFeatures up front (late reference-space requests are the classic rejection), and build interactions on select/squeeze abstractions so controllers, hands and taps all work. The platform is headset-first: probe honestly and always ship the flat-screen version of the content.
Related tools: the WebGL table (the only render path), the gamepad table (what input sources hide), the fullscreen table (the sibling gesture-gated mode), the device motion table (the 3-DoF fallback), and the getUserMedia table (the camera half of AR passthrough ideas).

WebXR is the browser's door into VR and AR headsets: navigator.xr.isSessionSupported('immersive-vr') tells you whether hardware exists, requestSession('immersive-vr') (from a user gesture) starts a session that hands your code headset and controller poses at the device's own frame rate, and you render each eye into a WebGL framebuffer the compositor scans out to the display. Immersive AR is the same API with camera passthrough instead of a headset screen - the session mode string is the only fork.

Bottom line: WebXR inverts the usual web rendering contract. You do not own the frame loop, the canvas, or the display - the session does. Your renders happen inside session.requestAnimationFrame callbacks (the headset's clock, 72-120Hz, where a slow frame is not dropped pixels but a wobbling world - motion sickness is your performance bug), the pixels go to an XRWebGLLayer the compositor owns, and the session can end at any moment without your consent (the user took the headset off). Build the UI to be re-enterable and the loop to be finished in under a frame.

The honest part is the platform: the immersive web lives mostly on standalone headsets (Quest's browser is the real target), Android Chrome does AR with varying quality, desktop needs tethered hardware and is rare, and iOS Safari ships navigator.xr for almost nobody. That makes capability-gated UI the professional default: probe first, show the Enter VR button only when supported, and always keep a flat-screen fallback of the same content - WebXR is an enhancement path, not a foundation.

How to use

  1. Probe before you promise: const ok = await navigator.xr?.isSessionSupported('immersive-vr'); show the button only when true - and check 'immersive-ar' separately, support is not symmetric. Cross-origin embeds also need allow="xr-spatial-tracking".
  2. Start from a click: button.onclick = async () => { const session = await navigator.xr.requestSession('immersive-vr', {optionalFeatures: ['local-floor', 'bounded-floor']}); ... } - gesture-gated, secure-context only; request the reference space features as options so permission is granted up front.
  3. Bind the render contract: await gl.makeXRCompatible(); session.updateRenderState({baseLayer: new XRWebGLLayer(session, gl)}); - WebGL (2 preferred) is the only render path; the layer's framebuffer is what the headset shows, so size your viewport from it each frame.
  4. Walk the reference space: const refSpace = await session.requestReferenceSpace('local-floor'); - 'local' is head-locked origin, 'local-floor' adds the ground, 'bounded-floor' exposes guardian bounds; the type decides whether you need teleport locomotion or can trust roomscale.
  5. Render on the headset clock: session.requestAnimationFrame((t, frame) => { const pose = frame.getViewerPose(refSpace); for (const view of pose.views) { /* render view into baseLayer viewport */ } session.requestAnimationFrame(loop); }) - one callback per frame covering both eyes, then re-arm; controllers come from session.inputSources with selectstart/selectend and squeeze events.

Frequently asked questions

Why does requestSession fail even though isSessionSupported said true?

Because support and permission are different gates. The classic causes: missing user gesture (requestSession is click-gated like fullscreen - async chains that lose user activation get rejected), non-secure context (the whole API requires HTTPS), the headset being off or asleep (the browser talks to the runtime, and runtimes die), permissions policy in embeds (allow="xr-spatial-tracking" missing), or requesting a reference space the session options did not declare - requestReferenceSpace('local-floor') fails unless 'local-floor' was in optionalFeatures at requestSession. Also note the probe checks capability, not availability: isSessionSupported true on a desktop with no headset connected means the runtime COULD do it, and the session request is where reality answers. Debug order: gesture, HTTPS, feature flags, runtime status.

What are reference spaces and which should I pick?

They define where the world's origin sits and how far the user may travel - the coordinate system decision everything else inherits. 'local' anchors to the headset's startup pose (seated experiences, cockpit apps); 'local-floor' adds a detected or assumed floor under the user (the default for standing and roomscale-ish apps); 'bounded-floor' exposes the guardian/play-area polygon so you can design within it; 'unbounded' is for larger-area hardware with tracking that follows the user across rooms; and viewer/gaze spaces exist for non-immersive inline AR. The design coupling: pick 'local-floor' when you can - teleport locomotion plus floor anchoring works on every headset - and treat 'bounded-floor' as a bonus that enables physical-walk design. Request the types as optionalFeatures up front; asking for the space later is the rejection source, not a fallback.

How does the XR render loop differ from the normal rAF loop?

It is the headset's clock, not the tab's: session.requestAnimationFrame callbacks run at device rate (72Hz on Quest, 90-120Hz elsewhere) aligned to the compositor's schedule, with a different signature - (time, XRFrame) - where the XRFrame is your one-shot access to poses for that exact predicted moment. Three rules follow. Do not render the XR scene from window.requestAnimationFrame - the session owns the framebuffer and the timing. Do not run heavy logic inline - a frame that overruns its budget does not tear, it wobbles the world (motion sickness); push non-render work to a worker and read results as they land. And poses are predictions: getViewerPose gives where the head WILL be at scanout - use it for rendering, and getBoundingClientRect-style thinking about 'current' positions does not apply. End the loop by simply not re-arming after the session's end event.

How do controllers, hands and taps arrive in WebXR?

As input sources on the session, unified. session.inputSources is a live array of XRInputSource objects - controllers, tracked hands, screen taps - each with handedness, targetRaySpace and gripSpace (two poses: pointing vs holding) and, for controllers, a gamepad object under .gamepad using the xr-standard mapping. Interaction is event-driven, not polling: selectstart/selectend is the primary trigger (trigger, tap, or hand pinch - the abstraction is deliberate), squeeze events cover grip buttons. The design guidance: build on select/squeeze and the target ray, not on button indices, and your app works across controllers, hands and eye-trackers without branching. Haptics, where supported, hide in the gamepad's vibrationActuator - the same pulse API as regular Gamepad - so the gamepad table's patterns transfer directly.

Related tools