> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Configure ingestion

Ingestion configuration is optional. If omitted, [Jockey](/v1.3/agents/concepts/jockey) uses the default extraction. Configure ingestion when you know what to extract. The platform then prioritizes that information during indexing, instead of the general-purpose default.

Choose the approach that fits your use case:

* **Default (no configuration)**: No setup required. Use it for getting started or quick prototyping.
* **Natural language description**: Describe what to extract in plain language. Use it when you know the domain but not the exact fields.
* **JSON Schema**: Define the exact fields and types to extract. Use it when downstream code expects specific typed fields.

# Natural language description

Use this approach for exploratory work when you don't yet know the exact fields you need.

## Example

The following example focuses extraction on brand mentions, product appearances, audience reactions, and visual tone. Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=15
from twelvelabs import TwelveLabs, IngestionConfig, EnrichmentConfig_Description

client = TwelveLabs(api_key="<YOUR_API_KEY>")

store = client.knowledge_stores.create(
    name="<YOUR_KNOWLEDGE_STORE_NAME>",
    ingestion_config=IngestionConfig(
        enrichment_config=EnrichmentConfig_Description(
            description="Focus on brand mentions, product appearances, audience reactions, and visual tone"
        )
    ),
)
print(f"Knowledge store created: {store.id}")
```

**`Node.js`**

```javascript Node.js maxlines=15
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });

const store = await client.knowledgeStores.create({
  name: "<YOUR_KNOWLEDGE_STORE_NAME>",
  ingestionConfig: {
    enrichmentConfig: {
      type: "description",
      description: "Focus on brand mentions, product appearances, audience reactions, and visual tone",
    },
  },
});
console.log(`Knowledge store created: ${store.id}`);
```

## Code explanation

#### Python

To configure ingestion with a natural-language description, pass an `ingestion_config` to the [`knowledge_stores.create`](/v1.3/sdk-reference/python/knowledge-stores#create-a-knowledge-store) method.\

**Parameters**:

* `ingestion_config`: An object that configures what the platform extracts during indexing. It contains an `enrichment_config` object with the following properties:
  * `type`: The enrichment type. Set to `"description"`.
  * `description`: Natural-language instructions. Jockey interprets your description to guide extraction.\


**Return value**: An object of type `KnowledgeStore` with a field named `id` representing the unique identifier of the newly created knowledge store.

#### Node.js

To configure ingestion with a natural-language description, pass an `ingestionConfig` to the [`knowledgeStores.create`](/v1.3/sdk-reference/node-js/knowledge-stores#create-a-knowledge-store) method. You pass all parameters as properties of a single object.\

**Parameters**:

* `ingestionConfig`: An object that configures what the platform extracts during indexing. It contains an `enrichmentConfig` object with the following properties:
  * `type`: The enrichment type. Set to `"description"`.
  * `description`: Natural-language instructions. Jockey interprets your description to guide extraction.\


**Return value**: An `HttpResponsePromise` that resolves to an object of type `KnowledgeStore` with a field named `id` representing the unique identifier of the newly created knowledge store.

# JSON Schema

Use this approach when downstream code expects specific typed fields.

## Example

The following example extracts metadata about people, locations, and activities from each video shot. Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=30
from twelvelabs import (
    TwelveLabs,
    IngestionConfig,
    EnrichmentConfig_JsonSchema,
    EnrichmentConfigJsonSchemaJsonSchema,
)

client = TwelveLabs(api_key="<YOUR_API_KEY>")

store = client.knowledge_stores.create(
    name="<YOUR_KNOWLEDGE_STORE_NAME>",
    ingestion_config=IngestionConfig(
        enrichment_config=EnrichmentConfig_JsonSchema(
            json_schema=EnrichmentConfigJsonSchemaJsonSchema(
                type="object",
                properties={
                    "people_count": {"type": "integer", "description": "Number of people visible in the frame"},
                    "location": {"type": "string", "description": "Name or type of the location"},
                    "suspicious_activity": {"type": "boolean", "description": "Whether suspicious activity is detected"},
                    "scene_description": {"type": "string", "description": "Brief description of what is happening"},
                },
                required=["people_count", "scene_description"],
            )
        )
    ),
)
print(f"Knowledge store created: {store.id}")
```

**`Node.js`**

```javascript Node.js maxlines=30
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });

const store = await client.knowledgeStores.create({
  name: "<YOUR_KNOWLEDGE_STORE_NAME>",
  ingestionConfig: {
    enrichmentConfig: {
      type: "json_schema",
      jsonSchema: {
        type: "object",
        properties: {
          people_count: { type: "integer", description: "Number of people visible in the frame" },
          location: { type: "string", description: "Name or type of the location" },
          suspicious_activity: { type: "boolean", description: "Whether suspicious activity is detected" },
          scene_description: { type: "string", description: "Brief description of what is happening" },
        },
        required: ["people_count", "scene_description"],
      },
    },
  },
});
console.log(`Knowledge store created: ${store.id}`);
```

## Code explanation

#### Python

To configure ingestion with a JSON Schema, pass an `ingestion_config` to the [`knowledge_stores.create`](/v1.3/sdk-reference/python/knowledge-stores#create-a-knowledge-store) method.\

**Parameters**:

* `ingestion_config`: An object that configures what the platform extracts during indexing. It contains an `enrichment_config` object with the following properties:
  * `type`: The enrichment type. Set to `"json_schema"`.
  * `json_schema`: A JSON Schema (draft 2020-12). The root `type` keyword must be `"object"`. See the [JSON Schema keywords](#json-schema-keywords) section for the accepted keywords.\


**Return value**: An object of type `KnowledgeStore` with a field named `id` representing the unique identifier of the newly created knowledge store.

#### Node.js

To configure ingestion with a JSON Schema, pass an `ingestionConfig` to the [`knowledgeStores.create`](/v1.3/sdk-reference/node-js/knowledge-stores#create-a-knowledge-store) method. You pass all parameters as properties of a single object.\

**Parameters**:

* `ingestionConfig`: An object that configures what the platform extracts during indexing. It contains an `enrichmentConfig` object with the following properties:
  * `type`: The enrichment type. Set to `"json_schema"`.
  * `jsonSchema`: A JSON Schema (draft 2020-12). The root `type` keyword must be `"object"`. See the [JSON Schema keywords](#json-schema-keywords) section for the accepted keywords.\


**Return value**: An `HttpResponsePromise` that resolves to an object of type `KnowledgeStore` with a field named `id` representing the unique identifier of the newly created knowledge store.

# JSON Schema keywords

> **Note**
>
> The constraints in this section apply only to ingestion. The schema for the [structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) feature of the Responses API supports a different set of keywords, and unsupported keywords do not return an error. They produce incomplete or malformed output.

The platform accepts the following JSON Schema keywords:

| Category | Keywords                                         |
| -------- | ------------------------------------------------ |
| Core     | `type`, `title`, `description`, `enum`           |
| Object   | `properties`, `required`, `additionalProperties` |
| Array    | `items`, `prefixItems`, `minItems`, `maxItems`   |
| Number   | `minimum`, `maximum`                             |
| String   | `format`                                         |

Notes:

* Include a `description` field in every property. The platform uses this text to guide extraction, and omitting it returns a `422` error.
* Use the `required` keyword to specify which fields must appear in every result.
* Set the `additionalProperties` keyword to `true` or `false` to control strict shapes.
* Do not use nullable fields. To make a field optional, omit it from the `required` array.
* Do not include unknown keywords. The platform rejects them with a `422` error.

## Not supported

The platform does not support the following JSON Schema keywords:

* Schema composition: `anyOf`, `allOf`, `oneOf`, `not`
* Conditional schemas: `if` / `then` / `else`
* Property dependencies: `dependentSchemas`, `dependentRequired`
* Object size constraints: `minProperties`, `maxProperties`
* String constraints: `pattern`, `minLength`, `maxLength`
* Number constraints: `exclusiveMinimum`, `exclusiveMaximum`, `multipleOf`
* Value constraints: `const`
* Null handling: `nullable`, `type` arrays (e.g., `["string", "null"]`)
* Annotative keywords: `default`, `examples`, `readOnly`, `writeOnly`
* References and reuse: `$ref`, `$defs`, `definitions`

# Common pitfalls

* **Overly specific schemas can limit extraction.** If the schema is too specific, Jockey may miss relevant content. Start broad, then narrow the schema as you learn what you need.
* **General descriptions can lead to broad extraction.** If the description is too general, Jockey may return results that are broader than you need. Name the fields, entities, or patterns you want Jockey to emphasize.
* **Programmatically generated schemas may include unsupported keywords.** If you use tools like `pydantic.model_json_schema()`, the output may contain annotative keywords such as `default`, `examples`, and `readOnly`. Strip these before submitting. The platform rejects unknown keywords with a `422` error.

# Jupyter notebook

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/twelvelabs-io/twelvelabs-developer-experience/blob/main/quickstarts/jockey/guides/ingestion_config.ipynb)