JavaScript Regex Groups Table
| Piece | What it captures | Field note |
|---|---|---|
(\d+) | Numbered group - $1 or match[1] | Numbering opens at the first opening paren |
(?<year>\d{4}) | Named group | match.groups.year - self-documenting extraction |
(?:...) | Non-capturing group | Groups without consuming a $N slot |
(?<=\$)\d+ | Lookbehind - preceded by | Matches the digits after a literal $ without capturing it |
\d+(?=px) | Lookahead - followed by | Matches numbers whose unit is px, unit not captured |
(?<i>word) | Case-insensitive group | i flag scoped to one group |
\1 / $1 | Backreference | Match the SAME text group 1 captured - doubles, quotes |
Groups turn a regex from a matcher into an extractor: parentheses capture the parts you actually want out of a match. The table below is the working seven - numbered and named captures, the non-capturing form, both lookarounds, and backreferences - with the numbering rule that trips everyone once.
Bottom line: numbering opens at the FIRST opening parenthesis and counts left to right - groups nest, and the outer group numbers before its inner ones. Named groups (?<year>\d{4}) are the upgrade that pays immediately: match.groups.year reads in code and in review, while match[3] needs a comment to explain what it holds.
The honest part: lookarounds look but never touch. (?<=\$)\d+ matches digits PRECEDED by a dollar sign without the dollar joining the match; \d+(?=px) matches numbers whose unit is px without the px. Zero-width assertions compose - you can stack a behind and an ahead on one pattern for 'sandwiched between' extraction without consuming either slice of bread.
How to use
- Number mentally left to right: the first ( is group 1, its inner ( is group 2 - count opening parens, including ones you meant to nest.
- Name anything a colleague will read: (?<year>\d{4}) costs six characters and buys self-documenting extraction via match.groups.
- Use lookarounds for context-without-capture: strip the $ from prices, keep the px out of measurements - the delimiter stays out of your result.
Frequently asked questions
How are capture groups numbered?
By counting OPENING parentheses left to right, ignoring whether they're capturing: in ((a)(b)), group 1 is the outer pair, group 2 is (a), group 3 is (b) - outer before inner, left before right. Non-capturing groups (?:...) don't take numbers, which is the escape hatch when you need grouping for precedence without disturbing the numbering your code depends on.
Why use named groups instead of numbered ones?
Self-documentation and refactoring safety. match.groups.year cannot break when someone inserts a new group early in the pattern - numbered captures shift and every match[3] silently becomes a different value. Named groups also survive into replace patterns as $<year> and into destructuring: const {groups: {year}} = str.match(re). The cost is a few characters; the benefit is a pattern whose extraction contract is written on its face.
What is the difference between the two lookarounds?
Direction. Lookahead (?=...) checks what FOLLOWS the match position; lookbehind (?<=...) checks what PRECEDES it. Neither consumes characters - the engine checks, then continues from the same spot. That's their superpower over plain groups: \d+(?=\s*px) matches the number in '24px' without the px joining the match, and (?<=\$)\d+ matches the price without the dollar sign - context asserted, result clean.
When do I need a backreference like \1?
When the SAME text must appear twice: (["'])\w*\1 matches a quoted string with the SAME quote character on both ends - group 1 captured whichever quote opened, and \1 demands the identical one closes. In replace patterns the syntax is $1. The modern niche: backreferences are the regex-only way to express 'exactly what I matched before', which no character class can express - though for validating structured input, a dedicated parser usually beats an unreadable backreference pattern.