> ## Documentation Index
> Fetch the complete documentation index at: https://docs.zenrows.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> The Zenrows API (Fetch, Extract via extract=auto, and Batch) is described in OpenAPI 3.1 at https://www.zenrows.com/openapi.json and listed in the RFC 9727 API catalog at https://www.zenrows.com/.well-known/api-catalog. Use the spec for exact parameter names and types.

# Crawl Setup

> Start a Zenrows crawl, read its status and results with a cursor, download every result in one file, get each page's HTML, and stop a crawl.

This guide takes one start URL through a full crawl. Every request sends your API key in the `X-API-Key` header. Parameters, limits and errors are in the [endpoint reference](/crawl/endpoints).

## Start a crawl

Send the start URL and `depth` to `POST /v1/crawls`. Add `max_items` to keep more than the default 10 URLs. `include_patterns` keeps a URL only if it contains one of the strings you pass, and Crawl opens only those URLs. A homepage crawl with only `/product/` doesn't open listing paths that lack that text, such as `/page/`.

<CodeGroup>
  ```python Python theme={"dark"}
  # pip install requests
  import requests

  response = requests.post(
      'https://api.zenrows.com/v1/crawls',
      headers={'X-API-Key': 'YOUR_ZENROWS_API_KEY'},
      json={
          'url': 'https://www.scrapingcourse.com/ecommerce/',
          'depth': 2,
          'max_items': 100,
          'max_pages': 20,
          'include_patterns': ['/product/'],
      },
  )
  crawl = response.json()
  print(response.status_code, crawl['crawl_id'])
  ```

  ```javascript Node.js theme={"dark"}
  const response = await fetch('https://api.zenrows.com/v1/crawls', {
      method: 'POST',
      headers: {
          'X-API-Key': 'YOUR_ZENROWS_API_KEY',
          'Content-Type': 'application/json',
      },
      body: JSON.stringify({
          url: 'https://www.scrapingcourse.com/ecommerce/',
          depth: 2,
          max_items: 100,
          max_pages: 20,
          include_patterns: ['/product/'],
      }),
  });
  const crawl = await response.json();
  console.log(response.status, crawl.crawl_id);
  ```

  ```bash cURL theme={"dark"}
  curl -X POST "https://api.zenrows.com/v1/crawls" \
    -H "X-API-Key: YOUR_ZENROWS_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "url": "https://www.scrapingcourse.com/ecommerce/",
      "depth": 2,
      "max_items": 100,
      "max_pages": 20,
      "include_patterns": ["/product/"]
    }'
  ```
</CodeGroup>

The response is `202 Accepted`. The crawl runs in the background:

```json Response theme={"dark"}
{
  "crawl_id": "c_01M3XP7RZD4JT88F7PNPGBFRAF",
  "status": "running",
  "url": "https://www.scrapingcourse.com/ecommerce/",
  "depth": 2,
  "max_items": 100,
  "max_pages": 20,
  "include_patterns": ["/product/"],
  "exclude_patterns": [],
  "discovery": ["links"],
  "coverage": {
    "pages_fetched": 0,
    "pages_failed": 0,
    "items_found": 0
  },
  "duplicates_removed": 0,
  "extract_calls": 0,
  "created_at": "2026-10-09T09:00:00Z"
}
```

Keep the `crawl_id`. Every other call takes it.

## Read the status and the results

`GET /v1/crawls/{crawl_id}` returns the crawl's `status`, its `coverage` so far and one page of results. Results grow while the crawl runs, so read them as you go:

1. Read without `cursor` to get the first results.
2. Send the `next_cursor` you received as `cursor` to get the URLs found since.
3. Stop when `next_cursor` is `null`. This happens only once the crawl has ended and you have read its last results.

<CodeGroup>
  ```python Python theme={"dark"}
  # pip install requests
  import time
  import requests

  crawl_id = 'YOUR_CRAWL_ID'
  url = f'https://api.zenrows.com/v1/crawls/{crawl_id}'
  headers = {'X-API-Key': 'YOUR_ZENROWS_API_KEY'}

  urls, cursor = [], None
  while True:
      params = {'cursor': cursor} if cursor else {}
      body = requests.get(url, headers=headers, params=params).json()
      urls += [result['url'] for result in body['results']]
      if body['next_cursor'] is None:
          break
      cursor = body['next_cursor']
      if not body['results']:
          time.sleep(5)  # nothing new yet

  print(body['status'], len(urls))
  ```

  ```javascript Node.js theme={"dark"}
  const crawlId = 'YOUR_CRAWL_ID';
  const url = `https://api.zenrows.com/v1/crawls/${crawlId}`;
  const headers = { 'X-API-Key': 'YOUR_ZENROWS_API_KEY' };

  const urls = [];
  let cursor = null;
  let body;
  while (true) {
      const params = cursor ? `?cursor=${encodeURIComponent(cursor)}` : '';
      body = await (await fetch(url + params, { headers })).json();
      urls.push(...body.results.map((result) => result.url));
      if (body.next_cursor === null) break;
      cursor = body.next_cursor;
      if (body.results.length === 0) {
          await new Promise((resolve) => setTimeout(resolve, 5000)); // nothing new yet
      }
  }

  console.log(body.status, urls.length);
  ```

  ```bash cURL theme={"dark"}
  curl "https://api.zenrows.com/v1/crawls/YOUR_CRAWL_ID" \
    -H "X-API-Key: YOUR_ZENROWS_API_KEY"
  ```
</CodeGroup>

A crawl that has ended reads like this:

```json Response theme={"dark"}
{
  "crawl_id": "c_01M3XP7RZD4JT88F7PNPGBFRAF",
  "status": "completed",
  "stop_reason": "max_items",
  "url": "https://www.scrapingcourse.com/ecommerce/",
  "depth": 2,
  "max_items": 100,
  "max_pages": 20,
  "include_patterns": ["/product/"],
  "exclude_patterns": [],
  "discovery": ["links"],
  "coverage": {
    "pages_fetched": 12,
    "pages_failed": 0,
    "items_found": 100
  },
  "duplicates_removed": 214,
  "extract_calls": 0,
  "created_at": "2026-10-09T09:00:00Z",
  "finished_at": "2026-10-09T09:01:12Z",
  "results": [
    { "url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/" },
    { "url": "https://www.scrapingcourse.com/ecommerce/product/adrienne-trek-jacket/" }
  ],
  "next_cursor": null
}
```

`stop_reason: max_items` means the site may hold more URLs. To get them, start a new crawl with a higher `max_items` (and a higher `max_pages` if `stop_reason` is `max_pages`).

## Download all results

`GET /v1/crawls/{crawl_id}/download` returns every URL in one NDJSON file: one JSON object per line.

```bash cURL theme={"dark"}
curl "https://api.zenrows.com/v1/crawls/YOUR_CRAWL_ID/download" \
  -H "X-API-Key: YOUR_ZENROWS_API_KEY" \
  -o crawl.jsonl
```

```json Response theme={"dark"}
{"url":"https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/"}
{"url":"https://www.scrapingcourse.com/ecommerce/product/adrienne-trek-jacket/"}
```

You can download while the crawl runs. The `X-Crawl-Status` header tells you the crawl's status: `running` means a later download may hold more lines. So can a download in the 10 minutes after a [stop](#stop-a-crawl), while the header reads `stopped`.

## Get each page's HTML

Send `"output_format": "html"` to get the HTML of each page the crawl keeps, not only its URL. Each kept page counts toward `max_pages`, so set `max_pages` to at least `max_items` plus the pages the crawl walks:

```bash cURL theme={"dark"}
curl -X POST "https://api.zenrows.com/v1/crawls" \
  -H "X-API-Key: YOUR_ZENROWS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://www.scrapingcourse.com/ecommerce/",
    "depth": 1,
    "max_items": 10,
    "max_pages": 20,
    "include_patterns": ["/product/"],
    "output_format": "html"
  }'
```

Each result then has a `content_status`, and a `content_url` once the page is fetched:

```json Response theme={"dark"}
{
  "url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/",
  "content_status": "fetched",
  "content_url": "/v1/crawls/c_01M3XP24VJ55Q96QQYKT6FQW1J/contents/ct_4bf794e572e17fdd50710f61cc730066"
}
```

`content_status` is `pending` until the page is fetched, then `fetched` or `failed`. After a [stop](#stop-a-crawl), a page not fetched yet stays `pending`. Read a fetched page at its `content_url`:

```bash cURL theme={"dark"}
curl "https://api.zenrows.com/v1/crawls/c_01M3XP24VJ55Q96QQYKT6FQW1J/contents/ct_4bf794e572e17fdd50710f61cc730066" \
  -H "X-API-Key: YOUR_ZENROWS_API_KEY" \
  -o page.html
```

The [download](#download-all-results) also includes the HTML of each fetched page, in a `content` field.

## Stop a crawl

`POST /v1/crawls/{crawl_id}/stop` stops a running crawl:

```bash cURL theme={"dark"}
curl -X POST "https://api.zenrows.com/v1/crawls/YOUR_CRAWL_ID/stop" \
  -H "X-API-Key: YOUR_ZENROWS_API_KEY"
```

The crawl then reads `status: stopped` with `stop_reason: user`. The URLs found before the stop stay readable. Pages already being fetched finish, are billed, and stay in the crawl: for up to 10 minutes after the stop, the links on them join the results. Keep reading with `next_cursor` until it is `null` for the final results. A stopped crawl can't resume.

## Next steps

<CardGroup cols={2}>
  <Card title="Crawl Endpoints" icon="book" href="/crawl/endpoints">
    Every parameter, limit, status and error code, plus pricing.
  </Card>

  <Card title="Adaptive Stealth Mode" icon="shield-check" href="/fetch/features/adaptive-stealth-mode">
    How Fetch picks the settings for each page a crawl opens.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.