> ## Documentation Index
> Fetch the complete documentation index at: https://docs.zenrows.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> The Zenrows API (Fetch, Extract via extract=auto, and Batch) is described in OpenAPI 3.1 at https://www.zenrows.com/openapi.json and listed in the RFC 9727 API catalog at https://www.zenrows.com/.well-known/api-catalog. Use the spec for exact parameter names and types.

# Crawl Endpoints

> Reference for the Zenrows Crawl endpoints: parameters, response fields, statuses, limits, pricing and error codes.

Crawl's endpoints live under `https://api.zenrows.com/v1/crawls`. Send your API key in the `X-API-Key` header or the `apikey` query parameter.

| Method | Path | Does |
| - | - | - |
| `POST` | `/v1/crawls` | [Start a crawl](#start-a-crawl) |
| `GET` | `/v1/crawls` | [List your crawls](#list-your-crawls) |
| `GET` | `/v1/crawls/{crawl_id}` | [Read a crawl and its results](#read-a-crawl) |
| `GET` | `/v1/crawls/{crawl_id}/download` | [Download all results](#download-the-results) |
| `GET` | `/v1/crawls/{crawl_id}/contents/{content_id}` | [Read one page's HTML](#read-a-page) |
| `POST` | `/v1/crawls/{crawl_id}/stop` | [Stop a crawl](#stop-a-crawl) |

## Start a crawl

### `POST /v1/crawls`

Starts a crawl and returns `202 Accepted` at once, with the crawl in the body and its path in the `Location` header. The body is JSON.

| PARAMETER | TYPE | DEFAULT | DESCRIPTION |
| - | - | - | - |
| **url** `required` | `string` | | The start URL: a public `http` or `https` URL, up to 2,048 characters. The crawl starts on this exact URL, so one ending in `?page=3` starts on page 3. It stays on the URL's domain, subdomains included. A link longer than 2,048 characters on a crawled page is skipped. |
| **depth** `required` | `integer` | | How many link steps to follow from the start URL, from 1 to 100,000. At `1`, Crawl returns the links on the start page. At `2`, it also opens those pages and returns their links. |
| **max\_items** | `integer` | `10` | Stops the crawl after it keeps this many URLs, from 1 to 100,000. With `output_format`, `max_pages` can stop it first, because each kept page takes a fetch. |
| **max\_pages** | `integer` | `10` | Stops the crawl after it fetches this many pages, from 1 to 100,000. With `output_format`, each kept page's fetch counts too. If both limits are reached together, `stop_reason` is `max_pages`. |
| **include\_patterns** | `array` of `string` | | Keeps a URL only if it contains at least one of these texts, such as `/product/`, and opens only those URLs. A homepage crawl with only `/product/` skips listing URLs that lack that text, such as `/page/`. The start URL is never returned. |
| **exclude\_patterns** | `array` of `string` | | Drops a URL that contains any of these texts, even if it matches an include pattern. |
| **output\_format** | `string` | | `html` also returns the HTML of each kept page. Leave it out to get URLs only. |

Send an `Idempotency-Key` header (up to 255 characters) to make a retry safe: a retry with the same key and body returns the same crawl and starts no second one.

## List your crawls

### `GET /v1/crawls`

Returns your crawls, newest first, without their results.

| PARAMETER | TYPE | DEFAULT | DESCRIPTION |
| - | - | - | - |
| **cursor** | `string` | | The `next_cursor` from the previous page. |
| **limit** | `integer` | `20` | How many crawls to return, from 1 to 100. |

The response holds a `crawls` array and a `next_cursor`. `next_cursor` is absent on the last page.

## Read a crawl

### `GET /v1/crawls/{crawl_id}`

Returns the crawl, its coverage so far, and one page of results.

| PARAMETER | TYPE | DEFAULT | DESCRIPTION |
| - | - | - | - |
| **cursor** | `string` | | The `next_cursor` from the previous read. Leave it out to start from the first result. |
| **limit** | `integer` | `1000` | How many results to return, from 1 to 10,000. |

While the crawl runs, and while a stopped crawl still reads the pages fetched before the stop, `next_cursor` is never `null`: read again with the same cursor to get only the new URLs. `next_cursor` is `null` once the crawl has ended and you have read its last results. Each URL is returned once.

### Response fields

| Field | Description |
| - | - |
| `crawl_id` | The crawl's ID. |
| `status` | `running`, `completed`, `stopped` or `failed`. See [Statuses](#statuses). |
| `stop_reason` | Why the crawl ended early: `max_items`, `max_pages` or `user`. Absent when nothing was left to open, and when the crawl failed. |
| `error` | Why the crawl failed, with a `code` and a `detail`. Present only when `status` is `failed`. |
| `url`, `depth`, `max_items`, `max_pages`, `include_patterns`, `exclude_patterns`, `output_format` | The parameters the crawl runs with. |
| `discovery`, `extract_calls` | `["links"]` and `0`: the crawl follows the links in each page's HTML. |
| `coverage.pages_fetched` | Pages fetched so far, successful or not, at most `max_pages`. With `output_format`, this includes each kept page fetched for its HTML. |
| `coverage.pages_failed` | Pages that failed to load. |
| `coverage.items_found` | URLs the crawl returns so far, after duplicates and your patterns are applied. At most `max_items`. |
| `duplicates_removed` | Links the crawl met again and dropped. The results never hold a duplicate. |
| `created_at`, `finished_at` | When the crawl started and ended, in UTC. `finished_at` is absent while it runs. |
| `results[].url` | A URL the crawl found, normalized. Spellings of one URL count as one: `http` or `https`, with or without `www.` and a trailing slash. Tracking parameters, such as `utm_*`, `gclid` and `fbclid`, are removed. |
| `results[].content_status` | With `output_format` only: `pending`, `fetched` or `failed`. |
| `results[].content_url` | With `output_format` only, once `content_status` is `fetched`: the path to [read the page](#read-a-page). |
| `next_cursor` | Send it as `cursor` on the next read. |

### Statuses

| `status` | Meaning |
| - | - |
| `running` | The crawl is opening pages. Results grow as it goes. |
| `completed` | The crawl finished. Nothing was left to open, or it reached `max_items` or `max_pages` (see `stop_reason`). |
| `stopped` | You stopped it with [`POST /v1/crawls/{crawl_id}/stop`](#stop-a-crawl). `stop_reason` is `user`. The results hold the URLs kept before the stop and the links on the pages already fetched by then. |
| `failed` | The crawl could not go on. `error.code` says why. The URLs found so far stay readable. |

| `error.code` | Meaning | What to do |
| - | - | - |
| `seed_unreachable` | The start URL could not be fetched. | Check the URL. The site may be blocking it. |
| `no_items_found` | The crawl found no URL to return. | Check your `include_patterns` and `exclude_patterns`. Without patterns, the site may have served a block or consent page. |
| `insufficient_credits` | Your account ran out of credits. | Add credits or upgrade your plan, then start a new crawl. |
| `internal_error` | Something failed on our side. | Start a new crawl. If it fails again, contact support with the `crawl_id`. |

## Download the results

### `GET /v1/crawls/{crawl_id}/download`

Returns every result as one NDJSON file (one JSON object per line), named `{crawl_id}.jsonl`. Each line has the `url`. With `output_format`, each line also has `content_status` and, once fetched, the page HTML in `content`.

The `X-Crawl-Status` header holds the crawl's status. While it is `running`, a later download may hold more lines. After a stop, it reads `stopped`, yet a later download may still hold more lines for up to 10 minutes: read the crawl until `next_cursor` is `null` to know the results are final.

## Read a page

### `GET /v1/crawls/{crawl_id}/contents/{content_id}`

Returns the HTML of one kept page, for a crawl started with `output_format`. Take the path from the result's `content_url`.

| PARAMETER | TYPE | DEFAULT | DESCRIPTION |
| - | - | - | - |
| **download** | `boolean` | `false` | `true` returns the page as a file to save. |

## Stop a crawl

### `POST /v1/crawls/{crawl_id}/stop`

Stops a running crawl and returns `200` with its `crawl_id`, `status`, `stop_reason` and `finished_at`. Send no request body.

```json Response theme={"dark"}
{
  "crawl_id": "c_01M3XP7RZD4JT88F7PNPGBFRAF",
  "status": "stopped",
  "stop_reason": "user",
  "finished_at": "2026-10-09T09:01:30Z"
}
```

No new page is fetched after the stop. Pages already being fetched finish and are billed, and the crawl keeps them: for up to 10 minutes after the stop, it reads each one, counts it in `coverage`, and keeps the links on it within your patterns and `max_items`. With `output_format`, a URL kept after the stop reads `content_status: pending`, since no new page is fetched. `status`, `stop_reason` and `finished_at` stay as the stop set them. Read the crawl until `next_cursor` is `null` for its final results. A crawl that already ended returns its final state. If the stop returns `500 internal_error`, send it again.

## Limits

| Limit | Value |
| - | - |
| Active crawls per account | 3 by default, shared with your [Batch](/batch/introduction) jobs |
| `max_items` | 10 by default, 100,000 at most |
| `max_pages` | 10 by default, 100,000 at most |
| `depth` | Required, from 1 to 100,000 |
| Results per read | 1,000 by default, 10,000 at most |

An account runs 3 active jobs by default, shared with [Batch](/batch/introduction). One more returns `429 too_many_crawls`, with a `Retry-After` header of 30 seconds. To run more than 3 at once, use the chat in your [dashboard](https://app.zenrows.com). Email [success@zenrows.com](mailto:success@zenrows.com) if you can't use the chat.

## Pricing

Crawl has no fee of its own. Each page a crawl opens is one Fetch (formerly, Universal Scraper API) request with [Adaptive Stealth Mode](/fetch/features/adaptive-stealth-mode), billed at that mode's price: from 1 to 25 credits per page, depending on the site. A page that fails to load (blocked, timed out or server error) is not billed. A page that answers 404 or 410 counts as failed but is billed, as in Fetch. A refused request starts no crawl and costs nothing.

`max_pages` caps how many pages a crawl fetches, so it caps what the crawl costs. With `output_format`, each kept page counts as one fetch. A kept page that the crawl also opens is fetched once, unless the crawl fetches it for its HTML first and reaches it later from a page nearer the start: then it is fetched, and billed, twice.

## Errors

Errors return as `application/problem+json`, with a `code`, a `title`, a `detail` that says what to do, the `status` and an `instance` to quote to support.

| Code | Status | Meaning |
| - | - | - |
| `invalid_request` | `400` | The body is not valid JSON, the `Idempotency-Key` is over 255 characters, or the list's `limit` or `cursor` is not valid. |
| `unknown_parameter` | `400` | The request has a body field or query parameter this endpoint does not accept. `detail` names it. |
| `invalid_cursor` | `400` | The `cursor` does not belong to this crawl. |
| `invalid_parameter` | `400`, `422` | A value is missing or out of range. `detail` names the field. |
| `crawl_not_found` | `404` | No crawl with this ID exists on your account. |
| `content_not_found` | `404` | The page is not fetched yet, its fetch failed, or the crawl has no `output_format`. |
| `idempotency_request_in_flight` | `409` | A request with this `Idempotency-Key` is still in progress. Retry when it ends. |
| `invalid_start_url` | `422` | The start URL is not a public `http` or `https` URL: its host is an IP address or a private address, or it is over 2,048 characters. |
| `idempotency_key_reused` | `422` | This `Idempotency-Key` was used with a different body. |
| `too_many_crawls` | `429` | Your account already runs 3 active jobs by default, shared with [Batch](/batch/introduction). Wait for one to end. |
| `internal_error` | `500` | Something failed on our side. Retry later. |

For authentication and plan errors, see the [API error codes](/api-error-codes).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.