Skip to main content
Crawl is in Beta. Any active Zenrows API key can use it, with no extra setup.
Zenrows Crawl takes one start URL and returns the pages it finds from there. Point it at a homepage, a category page or a search results page, and Crawl follows the links for you and removes duplicates. You choose what you get back for each page: only its URL, or its HTML. You don’t write a link loop, a filter or a deduplication step. A crawl is a background job. You start it with one request, then read its status and its results while it runs.

How it works

  1. Crawl opens the start URL with Fetch (formerly, Universal Scraper API) and Adaptive Stealth Mode (mode=auto).
  2. It reads every link in the page’s HTML. It keeps the links on the same domain (subdomains included) that match your filters, and drops duplicates: http and https links to one page count as one. It skips a link longer than 2,048 characters.
  3. While depth allows, it opens the kept links and reads their links too. One step from a page to its links is one level of depth.
  4. It stops when nothing is left to open, or when it reaches max_items or max_pages.
include_patterns chooses which links Crawl opens, not only which URLs it returns. A homepage crawl with only /product/ opens the product links on that page and skips listing and pagination URLs that don’t contain that text, such as /page/, so it returns almost nothing from the rest of the site. To walk those pages, include their path as well and raise depth. Crawl returns the URLs that match.

Basic usage

Start a crawl with a POST to /v1/crawls:
cURL
With only /product/, the request above returns product URLs linked from the homepage. It doesn’t open /page/ or other listing URLs. The response is 202 Accepted with a crawl_id. Read the crawl with GET /v1/crawls/{crawl_id}, sending each response’s next_cursor as cursor on the next read, until next_cursor is null: then it has ended and you have every result. You can also try Crawl without code in the Crawl playground.

Next steps

Crawl Setup

Start a crawl, read its results, download them, and get each page’s HTML.

Crawl Endpoints

Every endpoint, parameter, limit and error code, plus pricing.