JavaScript Regex Syntax Table
| Piece | Syntax | Field note |
|---|---|---|
anchors | ^ $ \b | Position, not characters - \b is the word edge, \B its opposite |
classes | [abc] [a-z] [^abc] | ^ INSIDE brackets negates - outside it anchors |
shorthands | \d \w \s | Digits, word chars, whitespace - uppercase twin negates (\D \W \S) |
quantifiers | * + ? {n,m} | Greedy by default - append ? for lazy (.*?) |
the dot | . | Any char EXCEPT newline - the s flag fixes dot to mean all |
alternation | (abc|xyz) | Groups capture - (?:...) matches without capturing |
flags | g i m s u | g makes the regex OBJECT stateful - lastIndex marches on (the test() gotcha) |
lookaround | (?=...) (?!...) | Zero-width peek - check what follows without consuming it |
Regex is a tiny language for describing SHAPES: anchors pin where a match sits, character classes define what a position may hold, and quantifiers count how often - composed left to right into one pattern. Eight building blocks cover essentially every pattern you will ever write.
Bottom line: two runtime surprises account for most regex bugs. The dot matches everything EXCEPT newlines until the s flag widens it, and the g flag makes the regex OBJECT stateful - lastIndex marches forward between calls, which is why .test() with a /g/ pattern alternates true and false inside a loop.
The honest part: greedy quantifiers are the default and the danger - .* will eat as much as it can and walk it back (backtracking), which is both the wrong match result and, on adversarial input, the catastrophic-blowup CPU pattern. Lazy (.*?) and character-class negation ([^>]*) are the two everyday fixes.
How to use
- Pin before you match: ^start and end$ turn a substring search into a full-string validation - the difference between contains and is.
- Capture only when you will use it: (?:...) for grouping, (...) for extraction - named captures (?<year>\d{4}) make long patterns self-documenting.
- Reset or clone stateful regexes: a /g/ pattern reused across strings keeps lastIndex - either drop the g for test(), or create the regex inside the loop.
Frequently asked questions
Why does my regex .test() return true, then false, then true?
The g flag made your regex stateful. With /g/, successful matches record their position in the regex object's lastIndex property, and the NEXT test() call resumes from there - so testing a string that matches once gives true (found), then false (resumed past the match), forever alternating. The fixes: drop g when you only test (test does not need it), create a fresh regex per call inside loops, or reset re.lastIndex = 0 between uses. This is the single most-reported regex surprise in JavaScript, and it is a property of the regex OBJECT, not of the string.
What is the difference between greedy and lazy quantifiers?
Direction of appetite. Greedy (.*, +, {1,}) takes as MUCH as possible, then backtracks if the rest of the pattern fails; lazy (.*?, +?) takes as LITTLE as possible and expands only on demand. On '<a><b>', /<.*>/ greedily matches the whole string; /<.*?>/ lazily stops at the first '>'. Greedy is also the performance risk: on non-matching input, the engine's backtracking can explode combinatorially (ReDoS), which is why security-sensitive patterns prefer explicit character classes like [^<]* over the open-ended dot-star.
Why doesn't the dot match my multiline text?
By definition the dot matches any character EXCEPT line terminators - a 1960s lineage decision that survived into JavaScript. The s flag (dotAll) redefines it to match everything: /a.b/s crosses newlines. Without it, the classic workaround was [\s\S] - a character class joining whitespace and non-whitespace, which by definition covers all characters. Prefer the s flag for readability in modern engines; recognize [\s\S] in older code as the same statement.
When do lookahead assertions earn their keep?
When a match depends on context you must not consume. A lookahead (?=...) checks what follows WITHOUT including it in the match - passwords like /^(?=.*\d)(?=.*[a-z]).{8,}$/ validate several conditions in one pass because each lookahead resets to the same starting point. Negative lookahead (?!...) excludes: /\d+(?!px)/ matches numbers not followed by px. The boundary: lookbehind (?<=...) exists in modern engines too, but lookarounds cannot contain unbounded quantifiers safely in every engine - keep the peeks simple and they stay both correct and fast.