XPath Cheatsheet Table

ExpressionSelects
//itemEvery element named item, at any depth
/catalog/itemitem children of catalog - a specific path from the root
//item[1]The first item of each parent (indexing starts at 1)
//item[last()]The last item under each parent
//item[@id]Items that have an id attribute at all
//item[@id='a1']Items whose id attribute equals a1
//item[price>10]Items whose price child compares greater than 10
//item[contains(name,'war')]Items whose name child contains the text war
//title/text()The text content of title elements - not the elements
//title | //authorUnion: both title and author elements in one result
../@idThe id attribute of the parent - XPath walks up, CSS cannot
//*Every element in the document
(//item)[3]The third item in document order - parentheses first, then index
XPath is the query language of XML - and the only selector language browsers expose that can walk up the tree, match on text, and count. Test expressions in DevTools with $x("//item[1]"). Two cost rules from the MDN XPath guide: a leading // scans the whole document, so anchor it to a path when the document is large; and XML is case-sensitive - //Item is not //item. Inside predicates, indexing is 1-based ([1] is the first), which trips everyone who arrives from JavaScript arrays. Bottom line: reach for XPath when CSS cannot say it - parent steps, text matching, positional unions - and stay with CSS selectors otherwise. Related tools: regex cheatsheet for the pattern-matching sibling, json formatter for the data side, extract urls and extract emails for scrape-style extraction without XPath at all.

XPath is the query language of XML - and the only selector language browsers expose that can walk up the tree, match on text content, and compare positions. CSS selectors cannot do any of those three. The table below is the working set: thirteen expressions that cover the overwhelming majority of scraping scripts and test selectors, from the plain descendant search to the union and the parent step.

Bottom line: reach for XPath when CSS cannot say it - //item[@id='a1'] for attributes, contains() for text, ../@id for the parent - and stay with CSS selectors otherwise, because they are faster to read and natively faster to evaluate. Test any expression instantly in browser DevTools with $x("//item[1]").

The honest part: two cost rules do most of the damage in slow scrapes. A leading // scans the entire document, so anchor it to a path when the document is large; and predicates are 1-based ([1] is the first element), which trips everyone arriving from JavaScript's 0-based arrays - the table's last row shows the parentheses trick for document-order indexing.

How to use

  1. Find the intent you need in the table - expression first, plain-English description second.
  2. Swap in your own element and attribute names; predicates combine, so //item[@id='a1'][price>10] is legal.
  3. Verify in DevTools: $x("expression") in the console returns the live node list, so you iterate real results before writing the script.

Frequently asked questions

XPath or CSS selectors - which should I use?

Default to CSS: it is shorter, every browser evaluates it natively, and libraries everywhere accept it. Switch to XPath for the three things CSS cannot express - walking upward (..), matching on text content (contains(text(), 'x')), and crossing the DOM in document order with unions. Scraping tools like Scrapy and test frameworks often accept both in the same API.

Why does //item[1] return many elements while (//item)[1] returns one?

Predicates apply at each step: //item[1] selects the first item under each parent - one result per parent, many results total. Wrapping the path in parentheses applies the index to the whole result set: (//item)[1] is the first item in document order, exactly one. This is the single most common XPath surprise.

Is XPath case-sensitive?

Yes - XML names are case-sensitive, so //Item and //item select different things. Text comparisons inside contains() are case-sensitive too; the standard workaround is translate(., 'ABC', 'abc') to normalize case before comparing. HTML parsed as HTML is forgiving in browsers' tag-name handling, but never rely on it in XPath you also run against real XML.

What is the difference between //title/text() and //title?

//title returns title elements - nodes you can keep navigating from. //title/text() returns the character data inside them - strings, end of the line. Scrapers that want values want text(); testers that want to assert on structure want the elements. Mixing them up is why a working expression sometimes returns objects instead of strings.

Related tools