X-API-Key header. Parameters, limits and errors are in the endpoint reference.
Start a crawl
Send the start URL anddepth to POST /v1/crawls. Add max_items to keep more than the default 10 URLs. include_patterns keeps a URL only if it contains one of the strings you pass, and Crawl opens only those URLs. A homepage crawl with only /product/ doesn’t open listing paths that lack that text, such as /page/.
202 Accepted. The crawl runs in the background:
Response
crawl_id. Every other call takes it.
Read the status and the results
GET /v1/crawls/{crawl_id} returns the crawl’s status, its coverage so far and one page of results. Results grow while the crawl runs, so read them as you go:
- Read without
cursorto get the first results. - Send the
next_cursoryou received ascursorto get the URLs found since. - Stop when
next_cursorisnull. This happens only once the crawl has ended and you have read its last results.
Response
stop_reason: max_items means the site may hold more URLs. To get them, start a new crawl with a higher max_items (and a higher max_pages if stop_reason is max_pages).
Download all results
GET /v1/crawls/{crawl_id}/download returns every URL in one NDJSON file: one JSON object per line.
cURL
Response
X-Crawl-Status header tells you the crawl’s status: running means a later download may hold more lines. So can a download in the 10 minutes after a stop, while the header reads stopped.
Get each page’s HTML
Send"output_format": "html" to get the HTML of each page the crawl keeps, not only its URL. Each kept page counts toward max_pages, so set max_pages to at least max_items plus the pages the crawl walks:
cURL
content_status, and a content_url once the page is fetched:
Response
content_status is pending until the page is fetched, then fetched or failed. After a stop, a page not fetched yet stays pending. Read a fetched page at its content_url:
cURL
content field.
Stop a crawl
POST /v1/crawls/{crawl_id}/stop stops a running crawl:
cURL
status: stopped with stop_reason: user. The URLs found before the stop stay readable. Pages already being fetched finish, are billed, and stay in the crawl: for up to 10 minutes after the stop, the links on them join the results. Keep reading with next_cursor until it is null for the final results. A stopped crawl can’t resume.
Next steps
Crawl Endpoints
Every parameter, limit, status and error code, plus pricing.
Adaptive Stealth Mode
How Fetch picks the settings for each page a crawl opens.