JavaScript Regex Groups Table

PieceWhat it capturesField note
(\d+)Numbered group - $1 or match[1]Numbering opens at the first opening paren
(?<year>\d{4})Named groupmatch.groups.year - self-documenting extraction
(?:...)Non-capturing groupGroups without consuming a $N slot
(?<=\$)\d+Lookbehind - preceded byMatches the digits after a literal $ without capturing it
\d+(?=px)Lookahead - followed byMatches numbers whose unit is px, unit not captured
(?<i>word)Case-insensitive groupi flag scoped to one group
\1 / $1BackreferenceMatch the SAME text group 1 captured - doubles, quotes
Reference: the MDN regular expressions guide. Named groups (?<name>...) are the upgrade that pays immediately: match.groups.year reads in code and in review, while match[3] needs a comment. The two lookaround forms split cleanly: lookahead (?=...) looks AFTER the match, lookbehind (?<=...) BEFORE it - neither consumes characters, which is why they compose for delimited extraction. Bottom line: numbered groups in throwaway scripts, named groups in anything a colleague will read, lookarounds for context without capture. Related tools: string methods table (match/replace/replaceAll), optional chaining table (match.groups access), and extract numbers for the no-code version.

Groups turn a regex from a matcher into an extractor: parentheses capture the parts you actually want out of a match. The table below is the working seven - numbered and named captures, the non-capturing form, both lookarounds, and backreferences - with the numbering rule that trips everyone once.

Bottom line: numbering opens at the FIRST opening parenthesis and counts left to right - groups nest, and the outer group numbers before its inner ones. Named groups (?<year>\d{4}) are the upgrade that pays immediately: match.groups.year reads in code and in review, while match[3] needs a comment to explain what it holds.

The honest part: lookarounds look but never touch. (?<=\$)\d+ matches digits PRECEDED by a dollar sign without the dollar joining the match; \d+(?=px) matches numbers whose unit is px without the px. Zero-width assertions compose - you can stack a behind and an ahead on one pattern for 'sandwiched between' extraction without consuming either slice of bread.

How to use

  1. Number mentally left to right: the first ( is group 1, its inner ( is group 2 - count opening parens, including ones you meant to nest.
  2. Name anything a colleague will read: (?<year>\d{4}) costs six characters and buys self-documenting extraction via match.groups.
  3. Use lookarounds for context-without-capture: strip the $ from prices, keep the px out of measurements - the delimiter stays out of your result.

Frequently asked questions

How are capture groups numbered?

By counting OPENING parentheses left to right, ignoring whether they're capturing: in ((a)(b)), group 1 is the outer pair, group 2 is (a), group 3 is (b) - outer before inner, left before right. Non-capturing groups (?:...) don't take numbers, which is the escape hatch when you need grouping for precedence without disturbing the numbering your code depends on.

Why use named groups instead of numbered ones?

Self-documentation and refactoring safety. match.groups.year cannot break when someone inserts a new group early in the pattern - numbered captures shift and every match[3] silently becomes a different value. Named groups also survive into replace patterns as $<year> and into destructuring: const {groups: {year}} = str.match(re). The cost is a few characters; the benefit is a pattern whose extraction contract is written on its face.

What is the difference between the two lookarounds?

Direction. Lookahead (?=...) checks what FOLLOWS the match position; lookbehind (?<=...) checks what PRECEDES it. Neither consumes characters - the engine checks, then continues from the same spot. That's their superpower over plain groups: \d+(?=\s*px) matches the number in '24px' without the px joining the match, and (?<=\$)\d+ matches the price without the dollar sign - context asserted, result clean.

When do I need a backreference like \1?

When the SAME text must appear twice: (["'])\w*\1 matches a quoted string with the SAME quote character on both ends - group 1 captured whichever quote opened, and \1 demands the identical one closes. In replace patterns the syntax is $1. The modern niche: backreferences are the regex-only way to express 'exactly what I matched before', which no character class can express - though for validating structured input, a dedicated parser usually beats an unreadable backreference pattern.

Related tools