Crawl is in Beta. Any active Zenrows API key can use it, with no extra setup.
How it works
- Crawl opens the start URL with Fetch (formerly, Universal Scraper API) and Adaptive Stealth Mode (
mode=auto). - It reads every link in the page’s HTML. It keeps the links on the same domain (subdomains included) that match your filters, and drops duplicates:
httpandhttpslinks to one page count as one. It skips a link longer than 2,048 characters. - While
depthallows, it opens the kept links and reads their links too. One step from a page to its links is one level ofdepth. - It stops when nothing is left to open, or when it reaches
max_itemsormax_pages.
include_patterns chooses which links Crawl opens, not only which URLs it returns. A homepage crawl with only /product/ opens the product links on that page and skips listing and pagination URLs that don’t contain that text, such as /page/, so it returns almost nothing from the rest of the site. To walk those pages, include their path as well and raise depth. Crawl returns the URLs that match.
Basic usage
Start a crawl with aPOST to /v1/crawls:
cURL
/product/, the request above returns product URLs linked from the homepage. It doesn’t open /page/ or other listing URLs.
The response is 202 Accepted with a crawl_id. Read the crawl with GET /v1/crawls/{crawl_id}, sending each response’s next_cursor as cursor on the next read, until next_cursor is null: then it has ended and you have every result.
You can also try Crawl without code in the Crawl playground.
Next steps
Crawl Setup
Start a crawl, read its results, download them, and get each page’s HTML.
Crawl Endpoints
Every endpoint, parameter, limit and error code, plus pricing.