Skip to main content
This guide takes one start URL through a full crawl. Every request sends your API key in the X-API-Key header. Parameters, limits and errors are in the endpoint reference.

Start a crawl

Send the start URL and depth to POST /v1/crawls. Add max_items to keep more than the default 10 URLs. include_patterns keeps a URL only if it contains one of the strings you pass, and Crawl opens only those URLs. A homepage crawl with only /product/ doesn’t open listing paths that lack that text, such as /page/.
The response is 202 Accepted. The crawl runs in the background:
Response
Keep the crawl_id. Every other call takes it.

Read the status and the results

GET /v1/crawls/{crawl_id} returns the crawl’s status, its coverage so far and one page of results. Results grow while the crawl runs, so read them as you go:
  1. Read without cursor to get the first results.
  2. Send the next_cursor you received as cursor to get the URLs found since.
  3. Stop when next_cursor is null. This happens only once the crawl has ended and you have read its last results.
A crawl that has ended reads like this:
Response
stop_reason: max_items means the site may hold more URLs. To get them, start a new crawl with a higher max_items (and a higher max_pages if stop_reason is max_pages).

Download all results

GET /v1/crawls/{crawl_id}/download returns every URL in one NDJSON file: one JSON object per line.
cURL
Response
You can download while the crawl runs. The X-Crawl-Status header tells you the crawl’s status: running means a later download may hold more lines. So can a download in the 10 minutes after a stop, while the header reads stopped.

Get each page’s HTML

Send "output_format": "html" to get the HTML of each page the crawl keeps, not only its URL. Each kept page counts toward max_pages, so set max_pages to at least max_items plus the pages the crawl walks:
cURL
Each result then has a content_status, and a content_url once the page is fetched:
Response
content_status is pending until the page is fetched, then fetched or failed. After a stop, a page not fetched yet stays pending. Read a fetched page at its content_url:
cURL
The download also includes the HTML of each fetched page, in a content field.

Stop a crawl

POST /v1/crawls/{crawl_id}/stop stops a running crawl:
cURL
The crawl then reads status: stopped with stop_reason: user. The URLs found before the stop stay readable. Pages already being fetched finish, are billed, and stay in the crawl: for up to 10 minutes after the stop, the links on them join the results. Keep reading with next_cursor until it is null for the final results. A stopped crawl can’t resume.

Next steps

Crawl Endpoints

Every parameter, limit, status and error code, plus pricing.

Adaptive Stealth Mode

How Fetch picks the settings for each page a crawl opens.