https://api.zenrows.com/v1/crawls. Send your API key in the X-API-Key header or the apikey query parameter.
Start a crawl
POST /v1/crawls
Starts a crawl and returns 202 Accepted at once, with the crawl in the body and its path in the Location header. The body is JSON.
Send an
Idempotency-Key header (up to 255 characters) to make a retry safe: a retry with the same key and body returns the same crawl and starts no second one.
List your crawls
GET /v1/crawls
Returns your crawls, newest first, without their results.
The response holds a
crawls array and a next_cursor. next_cursor is absent on the last page.
Read a crawl
GET /v1/crawls/{crawl_id}
Returns the crawl, its coverage so far, and one page of results.
While the crawl runs, and while a stopped crawl still reads the pages fetched before the stop,
next_cursor is never null: read again with the same cursor to get only the new URLs. next_cursor is null once the crawl has ended and you have read its last results. Each URL is returned once.
Response fields
Statuses
Download the results
GET /v1/crawls/{crawl_id}/download
Returns every result as one NDJSON file (one JSON object per line), named {crawl_id}.jsonl. Each line has the url. With output_format, each line also has content_status and, once fetched, the page HTML in content.
The X-Crawl-Status header holds the crawl’s status. While it is running, a later download may hold more lines. After a stop, it reads stopped, yet a later download may still hold more lines for up to 10 minutes: read the crawl until next_cursor is null to know the results are final.
Read a page
GET /v1/crawls/{crawl_id}/contents/{content_id}
Returns the HTML of one kept page, for a crawl started with output_format. Take the path from the result’s content_url.
Stop a crawl
POST /v1/crawls/{crawl_id}/stop
Stops a running crawl and returns 200 with its crawl_id, status, stop_reason and finished_at. Send no request body.
Response
coverage, and keeps the links on it within your patterns and max_items. With output_format, a URL kept after the stop reads content_status: pending, since no new page is fetched. status, stop_reason and finished_at stay as the stop set them. Read the crawl until next_cursor is null for its final results. A crawl that already ended returns its final state. If the stop returns 500 internal_error, send it again.
Limits
An account runs 3 active jobs by default, shared with Batch. One more returns
429 too_many_crawls, with a Retry-After header of 30 seconds. To run more than 3 at once, use the chat in your dashboard. Email success@zenrows.com if you can’t use the chat.
Pricing
Crawl has no fee of its own. Each page a crawl opens is one Fetch (formerly, Universal Scraper API) request with Adaptive Stealth Mode, billed at that mode’s price: from 1 to 25 credits per page, depending on the site. A page that fails to load (blocked, timed out or server error) is not billed. A page that answers 404 or 410 counts as failed but is billed, as in Fetch. A refused request starts no crawl and costs nothing.max_pages caps how many pages a crawl fetches, so it caps what the crawl costs. With output_format, each kept page counts as one fetch. A kept page that the crawl also opens is fetched once, unless the crawl fetches it for its HTML first and reaches it later from a page nearer the start: then it is fetched, and billed, twice.
Errors
Errors return asapplication/problem+json, with a code, a title, a detail that says what to do, the status and an instance to quote to support.
For authentication and plan errors, see the API error codes.