> ## Documentation Index
> Fetch the complete documentation index at: https://docs.convocore.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Create Crawler Job

> Creates a new crawler or scrape job using the V3 API.

## Overview

Creates a new crawler job for a workspace and starts processing it in the background.

<Note>
  Crawler jobs are immutable after creation. If you need different settings, delete the job and create a new one.
</Note>

## Supports

* Single-page scrape jobs
* Multi-URL scrape jobs
* Crawl jobs with `crawlOptions`
* Optional outbound webhooks for `page_scraped`, `job_completed`, and `job_failed`

## Job type

Created crawler jobs are tagged:

| `type`      | When                                          |
| ----------- | --------------------------------------------- |
| `kb_ingest` | `toAgentId` / `toAgentIds` is set (KB-linked) |
| `general`   | No agent destination                          |

For ad-hoc research scrapes that should not go through the workspace crawler, use [`POST /scrape`](/api-reference/v3/scrape/post).

## Billing

Credits are estimated at submission time and actually consumed per successful scraped page.

<Tip>
  Use `useProxy: true` only when needed. Proxy scraping has higher per-page credit cost.
</Tip>


## OpenAPI

````yaml POST /workspaces/{workspaceId}/crawler/jobs
openapi: 3.0.3
info:
  title: Convocore OpenAPI
  description: Full API reference for Convocore
  version: 1.0.4
servers:
  - url: https://eu-gcp-api.vg-stuff.com/v3
security: []
paths:
  /workspaces/{workspaceId}/crawler/jobs:
    post:
      tags:
        - Crawler
      summary: Create crawler job
      description: >-
        Creates a new scrape or crawl job for the authenticated workspace.
        Crawler jobs are immutable after submission. Updating a crawler job is
        not supported; create a new job or delete the existing job instead.
      operationId: createCrawlerJob
      parameters:
        - name: workspaceId
          in: path
          required: true
          schema:
            type: string
          description: The workspace that owns the crawler job.
          example: 360c48fb56eeaaaa322973c18
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                urls:
                  type: array
                  items:
                    type: string
                    format: uri
                  minItems: 1
                  description: >-
                    One or more source URLs to scrape or use as crawl entry
                    points.
                crawl:
                  type: boolean
                  description: >-
                    If true, discovered URLs can be followed and scraped as part
                    of the same job.
                crawlOptions:
                  type: object
                  properties:
                    maxPages:
                      $ref: '#/components/schemas/integer_61bce812'
                    urlMatchers:
                      $ref: '#/components/schemas/array_33ac78ec'
                    unMatchers:
                      $ref: '#/components/schemas/array_e9b67bd5'
                    stayOnDomain:
                      type: boolean
                      description: If true, only pages on the same domain are crawled.
                  additionalProperties: false
                deep:
                  type: boolean
                  description: If true, use deep scraping behavior.
                useProxy:
                  type: boolean
                  description: >-
                    If true, the crawler uses proxy scraping and paid proxy
                    pricing.
                aiPostProcess:
                  type: boolean
                  description: >-
                    Paid-only. After each page scrape, Gemini cleans markdown
                    (strip cookie/nav chrome, keep content/media/links).
                mode:
                  type: string
                  enum:
                    - markdown
                    - schema
                  description: >-
                    markdown = classic crawl; schema = AI listing→detail
                    structured crawl.
                schemaCrawl:
                  type: object
                  properties:
                    maxItems:
                      type: integer
                      minimum: 1
                      maximum: 500
                    freeForm:
                      type: boolean
                      description: >-
                        If true, AI collects page markdown without freezing a
                        Typesense schema (markdown-only destination).
                    collectMarkdown:
                      type: boolean
                      description: >-
                        If true with structured crawl (freeForm=false), also
                        store page markdown + aggregate for agent KB import
                        alongside Structured Search.
                    operatorNotes:
                      type: string
                      maxLength: 8000
                      description: >-
                        Notes/instructions for the AI crawler and Schema Crawl
                        Operator chat (injected into discovery/schema/extract).
                  additionalProperties: false
                  description: Options for schema listing crawl mode.
                refreshRate:
                  type: string
                  description: Optional refresh cadence for KB-linked scrapes.
                toAgentId:
                  type: string
                  description: Optional single agent destination for KB import.
                toAgentIds:
                  type: array
                  items:
                    type: string
                  description: Optional list of agent destinations for KB import.
                webhook:
                  type: object
                  properties:
                    url:
                      type: string
                      format: uri
                      description: >-
                        The HTTPS endpoint that should receive crawler
                        callbacks.
                    events:
                      type: array
                      items:
                        type: string
                        enum:
                          - page_scraped
                          - job_completed
                          - job_failed
                      minItems: 1
                      description: >-
                        Which crawler events should be delivered to your
                        webhook. Defaults to page and final events.
                    secret:
                      type: string
                      description: >-
                        Optional secret mirrored back as the
                        `x-convocore-crawler-secret` header.
                    bearerToken:
                      type: string
                      description: Optional bearer token sent as the Authorization header.
                    headers:
                      type: object
                      additionalProperties:
                        type: string
                      description: >-
                        Optional additional headers to send with each webhook
                        request.
                  required:
                    - url
                  additionalProperties: false
                  description: >-
                    Optional outbound webhook that receives `page_scraped`,
                    `job_completed`, and `job_failed` events.
              required:
                - urls
              additionalProperties: false
            example:
              urls:
                - https://example.com
              crawl: true
              crawlOptions:
                maxPages: 25
                urlMatchers:
                  - /
                stayOnDomain: true
              webhook:
                url: https://example.com/webhooks/crawler
                events:
                  - page_scraped
                  - job_completed
                  - job_failed
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                  message:
                    type: string
                  data:
                    $ref: '#/components/schemas/object_e29c2fd2'
                required:
                  - success
                  - message
                  - data
                additionalProperties: false
        default:
          $ref: '#/components/responses/error'
      security:
        - Authorization: []
components:
  schemas:
    integer_61bce812:
      type: integer
      minimum: 1
      maximum: 500
      description: Maximum number of pages the crawler is allowed to process.
    array_33ac78ec:
      type: array
      items:
        type: string
      description: URL path matchers that define which pages are in scope.
    array_e9b67bd5:
      type: array
      items:
        type: string
      description: URL path matchers that should be excluded.
    object_e29c2fd2:
      type: object
      properties:
        id:
          type: string
        workspaceId:
          type: string
        status:
          $ref: '#/components/schemas/string_2c1ccf80'
        primaryUrl:
          type: string
        urls:
          type: array
          items:
            type: string
        crawl:
          type: boolean
        crawlOptions:
          $ref: >-
            #/components/schemas/maxPages_urlMatchers_unMatchers_stayOnDomain_b95006
        useProxy:
          type: boolean
        aiPostProcess:
          type: boolean
        deep:
          type: boolean
        mode:
          type: string
          enum:
            - markdown
            - schema
        schemaCrawl:
          $ref: '#/components/schemas/object_fc2128c8'
        refreshRate:
          type: string
          nullable: true
        toAgentId:
          type: string
          nullable: true
        toAgentIds:
          type: array
          items:
            type: string
        done:
          type: boolean
        failed:
          type: boolean
        isCancelled:
          type: boolean
        message:
          type: string
          nullable: true
        resultError:
          type: string
          nullable: true
        createdAt:
          type: string
          nullable: true
        ts:
          type: number
        currentPageIndex:
          type: number
        scrapedPagesNum:
          type: number
        failedPagesNum:
          type: number
        pageLimit:
          type: number
        creditsPerPage:
          type: number
        estimatedCredits:
          type: number
        activeScrapeUrl:
          type: string
          nullable: true
        crawlerJobId:
          type: string
          nullable: true
        webhook:
          $ref: '#/components/schemas/url_events_hasSecret_hasBearerToken_e213dd'
      required:
        - id
        - workspaceId
        - status
        - primaryUrl
        - urls
        - crawl
        - crawlOptions
        - useProxy
        - deep
        - refreshRate
        - toAgentId
        - toAgentIds
        - done
        - failed
        - isCancelled
        - message
        - resultError
        - createdAt
        - ts
        - currentPageIndex
        - scrapedPagesNum
        - failedPagesNum
        - pageLimit
        - creditsPerPage
        - estimatedCredits
        - activeScrapeUrl
        - crawlerJobId
      additionalProperties: false
    string_2c1ccf80:
      type: string
      enum:
        - queued
        - processing
        - completed
        - failed
        - cancelled
    maxPages_urlMatchers_unMatchers_stayOnDomain_b95006:
      type: object
      properties:
        maxPages:
          $ref: '#/components/schemas/integer_61bce812'
        urlMatchers:
          $ref: '#/components/schemas/array_33ac78ec'
        unMatchers:
          $ref: '#/components/schemas/array_e9b67bd5'
        stayOnDomain:
          type: boolean
          description: If true, only pages on the same domain are crawled.
      additionalProperties: false
      nullable: true
    object_fc2128c8:
      type: object
      properties:
        phase:
          type: string
        maxItems:
          type: number
        discoveredCount:
          type: number
        extractedCount:
          type: number
        entityName:
          type: string
          nullable: true
        fieldCount:
          type: number
        pageCreditsUsed:
          type: number
        llmCreditsUsed:
          type: number
        totalCreditsUsed:
          type: number
        resultPath:
          type: string
          nullable: true
        lastDiscoveryReason:
          type: string
          nullable: true
      additionalProperties: false
      nullable: true
    url_events_hasSecret_hasBearerToken_e213dd:
      type: object
      properties:
        url:
          type: string
          format: uri
        events:
          $ref: '#/components/schemas/array_8ed97274'
        hasSecret:
          type: boolean
        hasBearerToken:
          type: boolean
        headerKeys:
          type: array
          items:
            type: string
      required:
        - url
        - events
        - hasSecret
        - hasBearerToken
        - headerKeys
      additionalProperties: false
    message_11569e:
      type: object
      properties:
        message:
          type: string
      required:
        - message
      additionalProperties: false
    array_8ed97274:
      type: array
      items:
        type: string
        enum:
          - page_scraped
          - job_completed
          - job_failed
  responses:
    error:
      description: Error response
      content:
        application/json:
          schema:
            type: object
            properties:
              message:
                type: string
              code:
                type: string
              issues:
                type: array
                items:
                  $ref: '#/components/schemas/message_11569e'
            required:
              - message
              - code
            additionalProperties: false
  securitySchemes:
    Authorization:
      type: http
      scheme: bearer

````