Web Scraping

What Is Xpath? How Can We Use It?

Learn what XPath is, how XPath expressions work, and how they can be used to locate and extract data from web pages.

PParsedomSeptember 26, 20231,684 views
What Is Xpath? How Can We Use It?

1. What is XPath?

XPath stands for XML Path Language. It is a query language used to navigate and select elements in XML documents or HTML pages. XPath is essentially the address of an element. It provides a way to locate specific elements or nodes within a structured document using path expressions, and can be used in various programming languages including JavaScript, Java, XML Schema, Python with Scrapy, and others.

XPath expressions are written to specify a path to the desired elements or nodes in an XML document or HTML page. These expressions can include element names, attributes, and various operators to define the specific location.

2. What is the Importance of XPath in Automation?

XPath plays a crucial role in automation, particularly in web scraping and automated testing. Here is why it matters:

  • Precise Element Selection: XPath enables automation scripts to precisely locate elements within a webpage or XML document based on their structure, attributes, or content.
  • Targeted Data Extraction: XPath is used to navigate through a document's structure and extract desired information such as text, URLs, images, or structured data during web scraping.
  • Handling Dynamic Elements: XPath can adapt to changes by dynamically locating elements based on their properties or relationships with other elements, making it ideal for dynamic web pages.
  • Automated Testing: XPath is extensively used in frameworks like Selenium to locate and interact with elements during test execution, simplifying the identification of elements for testing purposes.
  • Cross-Platform Compatibility: XPath is a standardized language supported by various programming languages and tools, making it a versatile and reliable choice for automation tasks.

3. What Are the Types of XPath?

There are two main types of XPath expressions:

Absolute XPath

An absolute XPath expression specifies the complete path from the root element to the target element. It begins with a forward slash / to denote the root element and then traverses through the hierarchy of elements, specifying element names and their positions along the path. Absolute XPath expressions are less flexible and more prone to breaking if the structure of the document changes.

Example

/html/body/div[1]/div[2]/form/input[3]

Relative XPath

A relative XPath expression selects elements based on their relationship with other elements in the document. It starts with a double forward slash // to select elements from anywhere in the document, not just from the root. Relative XPath expressions provide more flexibility as they can adapt to changes in the document structure.

Example

//input[@id='username']

This selects the input element with the attribute id equal to username from anywhere in the document.

4. How to Write an XPath?

To write an XPath expression, identify the target element by specifying its name or attributes. Choose between absolute (root-based) or relative (anywhere in the document) XPath. Consider using axes to navigate relationships and predicates for conditions. Test and validate the expression using appropriate tools.

Single Forward Slash /

A single forward slash in an XPath expression indicates the root element of the document. It is typically used at the beginning of an absolute XPath expression to specify the path starting from the root. For example, /html/body/div selects the div element that is a direct child of body, which is in turn a direct child of the root html element.

Double Forward Slash //

A double forward slash selects elements from anywhere in the document, regardless of their position or level within the hierarchy. For example, //div selects all div elements in the document, regardless of their parent elements or depth. The double forward slash is particularly useful for writing relative XPath expressions, enabling selection of elements based on attributes or content without specifying their exact location.

5. XPath Functions

XPath provides a variety of functions that can be used within expressions to perform operations, manipulate data, or extract specific information from XML documents or HTML pages. Here are some commonly used XPath functions:

text()

Retrieves the text content of an element.

//p/text()

Selects the text content of all <p> elements.

@attributeName

Retrieves the value of a specific attribute.

//div/@class

Selects the value of the "class" attribute of all <div> elements.

contains(string1, string2)

Checks if string2 is a substring of string1.

//h2[contains(text(), 'example')]

Selects all <h2> elements that contain the word "example" in their text content.

starts-with(string1, string2)

Checks if string1 starts with string2.

//a[starts-with(@href, 'https://')]

Selects all <a> elements whose "href" attribute starts with "https://".

position()

Returns the position of the current element within the selection.

(//li)[position() = 1]

Selects the first <li> element among all <li> elements.

last()

Returns the position of the last element within the selection.

(//tr/td)[last()]

Selects the last <td> element among all <td> elements inside <tr> elements.

Relative XPath Using Axes

Relative XPath expressions can be further enhanced by utilizing axes to navigate through the document structure and specify the relationship between elements. Here are the most commonly used axes:

child::

Selects direct child elements of the current context node.

//div/child::p

Selects all <p> elements that are direct children of <div> elements.

parent::

Selects the parent element of the current context node.

//p/parent::div

Selects the <div> element that is the parent of <p> elements.

descendant::

Selects all descendant elements of the current context node, regardless of their level.

//div/descendant::span

Selects all <span> elements that are descendants of <div> elements.

ancestor::

Selects all ancestor elements of the current context node, up to the root element.

//span/ancestor::div

Selects all <div> elements that are ancestors of <span> elements.

following-sibling::

Selects sibling elements that appear after the current context node.

//div/following-sibling::p

Selects all <p> elements that are siblings of <div> elements and appear after them.

preceding-sibling::

Selects sibling elements that appear before the current context node.

//p/preceding-sibling::span

Selects all <span> elements that are siblings of <p> elements and appear before them.

Relative XPath Without Using Axes

If you want to select elements without using XPath axes, you can consider alternative methods provided by your programming language or automation tool:

  • CSS Selectors: Many programming languages and automation frameworks support CSS selectors, which are often simpler and more concise than XPath expressions. CSS selectors can target elements by tag name, class, ID, attributes, or other properties.
  • DOM Traversal Methods: Programming languages often provide methods to navigate the Document Object Model (DOM) of a webpage. You can traverse the DOM tree and select elements based on their parent-child or sibling relationships.
  • Element Properties and Attributes: Some automation tools offer direct access to element properties and attributes such as class names, IDs, or custom attributes to select elements without explicitly using XPath expressions.

6. What is the Right Platform to Write and Verify XPath?

Using SelectorsHub

SelectorsHub is a browser extension and web scraping tool that simplifies the process of selecting and extracting data from web pages. It allows users to visually select elements on a webpage and generates the corresponding CSS selectors or XPath expressions automatically.

  • Provides a user-friendly interface for building selectors visually.
  • Supports advanced features like attribute extraction, parent-child relationships, and handling dynamic web content.
  • Commonly used for web scraping, data extraction, and automating web interactions.
  • Available as a browser extension for both Chrome and Firefox.

7. Conclusion

XPath is a powerful and essential tool in automation for precise element selection in both web scraping and automated testing. It allows scripts to navigate XML documents or HTML pages and extract data based on element structure, attributes, or content. XPath offers flexibility, handles dynamic elements gracefully, and is supported across virtually all major automation platforms and programming languages.

SelectorsHub is a recommended tool for writing and verifying XPath expressions visually in your browser. Alternatives like CSS selectors and DOM traversal methods are also available depending on your specific use case.

Further Learning

You can check out this tutorial to learn more about XPath, or read our blog on How To Scrape IMDB Movies Using Python & Scrapy to see XPath selectors in action.

Need Custom Solutions?

We are a data scraping and mining company. If you need help with any scraping or automation work, feel free to reach out at info@parsedom.com.

Web ScrapingXPathAutomation
Share this article
P

Parsedom

We build the data pipelines, analytics, and software that power modern companies from proxy infrastructure to full-scale SaaS products.

Need help with your own data pipeline?

This is exactly the kind of problem our team solves every day let's talk about yours.

Get in touch