> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Generate a corpus overview

> Summarize all your videos and images (themes, subjects, content types, patterns, and key statistics) in a single request.

This recipe shows you how to summarize all your videos and images (themes, subjects, content types, patterns, and key statistics) in a single request. You can receive the overview as prose or as structured JSON.

**Use cases**:

* **Collection onboarding**: Summarize a new collection before starting development
* **Dashboard generation**: Create collection summaries for a UI or reporting tool
* **Pipeline input**: Feed collection context to a downstream agent or workflow

# Key concepts

* **Knowledge store**: A persistent store of your videos and images plus the understanding the platform derives from them - spatiotemporal context, a typed ontology, and embeddings - that together enable corpus-level reasoning.
* **Structured output**: A JSON schema you provide with your request. Jockey constrains its output to match your schema, so you retrieve machine-readable results instead of plain text. For a guide on designing schemas and parsing responses, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page.

# Workflow

This recipe shows two approaches: a plain-text overview that needs only a prompt, and a structured overview that adds a JSON schema to define the output format. [Jockey](/v1.3/agents/concepts/jockey) reasons across your knowledge store, identifies themes and patterns, and returns results in the format you chose. You can adapt this recipe to target a specific domain or ask a different question about your collection.

# Prerequisites

* You've already created a knowledge store with at least one item in `ready` status. See the [Create a knowledge store](/v1.3/agents/guides/create-a-knowledge-store) and [Add assets to a knowledge store](/v1.3/agents/guides/add-assets) guides for details.
* You're familiar with the request and response format. See the [Create a response](/v1.3/agents/guides/create-a-response) page for details.

# Plain-text overview

Return the overview as prose. This approach needs only a prompt.

## Example

Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=15
from twelvelabs import TwelveLabs

client = TwelveLabs(api_key="<YOUR_API_KEY>")
STORE_ID = "<YOUR_KNOWLEDGE_STORE_ID>"

# Generate the overview
response = client.responses.create(
    knowledge_store_id=STORE_ID,
    input=[{"type": "message", "role": "user", "content": "Give me a comprehensive overview of these videos and images. Include main themes, recurring subjects, content types, and any notable patterns."}],
)

# Read the overview text
for output in response.output:
    if output.type == "message":
        for content in output.content:
            print(content.text)
```

**`Node.js`**

```javascript Node.js maxlines=15
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });
const storeId = "<YOUR_KNOWLEDGE_STORE_ID>";

// Generate the overview
const response = await client.responses.create({
  knowledgeStoreId: storeId,
  input: [{ type: "message", role: "user", content: "Give me a comprehensive overview of these videos and images. Include main themes, recurring subjects, content types, and any notable patterns." }],
});

// Read the overview text
for (const output of response.output ?? []) {
  if (output.type === "message") {
    for (const content of output.content ?? []) {
      console.log(content.text);
    }
  }
}
```

## Code explanation

#### Python

To generate a plain-text overview, call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with a prompt that requests a collection summary.\

**Parameters**:

* `knowledge_store_id`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey. This example requests a comprehensive overview covering themes, subjects, content types, and patterns.\


**Return value**: An object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledge_store_id`: The knowledge store this response was generated against.
* `session_id`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. Read the `text` field of each content part in a `message` item for the overview.
* `usage`: Token usage statistics, including the `input_tokens` and `output_tokens` fields.

#### Node.js

To generate a plain-text overview, call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with a prompt that requests a collection summary. You pass all parameters as properties of a single object.\

**Parameters**:

* `knowledgeStoreId`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey. This example requests a comprehensive overview covering themes, subjects, content types, and patterns.\


**Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledgeStoreId`: The knowledge store this response was generated against.
* `sessionId`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. Read the `text` field of each content part in a `message` item for the overview.
* `usage`: Token usage statistics, including the `inputTokens` and `outputTokens` fields.

## Example response

The `text` field inside each content part contains Jockey's overview formatted as Markdown. A typical response looks like this:

```text
### Overall character
- **Content type:** product tutorials, customer interviews, and feature demos
- **Primary audience:** developers and product managers
- **Structure:** a growing library of short-form instructional and promotional content

### Main themes
- **Product onboarding**: step-by-step walkthroughs of initial setup and core workflows
- **Feature announcements**: quarterly release highlights and capability demos
- **Customer success**: interviews and testimonials from active users

### Recurring subjects
- Dashboard and analytics tooling
- API authentication and integration workflows
- Mobile SDK setup and troubleshooting

### Notable patterns
- Tutorial output clusters around product launch dates
- Most webinars include a Q&A segment in the final third
- Interview content appears on a quarterly cadence
```

> **Note**
>
> Jockey returns Markdown-formatted text. If your downstream pipeline expects plain prose, strip the Markdown formatting before processing.

# Structured overview

Return the overview as JSON that matches a schema you define. Use this approach when downstream code expects typed fields.

## Example

Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=35
import json
from twelvelabs import TwelveLabs, TextParam
from twelvelabs.types.text_param_format import TextParamFormat_JsonSchema

client = TwelveLabs(api_key="<YOUR_API_KEY>")
STORE_ID = "<YOUR_KNOWLEDGE_STORE_ID>"

# Define a JSON schema for the overview
overview_schema = {
    "type": "object",
    "properties": {
        "total_items": {"type": "integer"},             # Number of items in the knowledge store
        "themes": {                                      # Recurring themes across items
            "type": "array",
            "items": {"type": "string"},
        },
        "content_types": {                               # Production formats (tutorial, demo, etc.)
            "type": "array",
            "items": {"type": "string"},
        },
        "key_subjects": {                                # Main subjects or topics covered
            "type": "array",
            "items": {"type": "string"},
        },
        "patterns": {                                    # Notable patterns or trends
            "type": "array",
            "items": {"type": "string"},
        },
        "summary": {"type": "string"},                  # Overall narrative summary
    },
}

# Generate the overview
response = client.responses.create(
    knowledge_store_id=STORE_ID,
    input=[{"type": "message", "role": "user", "content": "Give me a structured overview of these videos and images."}],
    text=TextParam(format=TextParamFormat_JsonSchema(name="corpus_overview", schema_=overview_schema)),
)

# Parse the JSON overview
for output in response.output:
    if output.type == "message":
        for content in output.content:
            overview = json.loads(content.text)
            print(f"Items: {overview['total_items']}")
            print(f"Themes: {', '.join(overview['themes'])}")
            print(f"Summary: {overview['summary']}")
```

**`Node.js`**

```javascript Node.js maxlines=35
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });
const storeId = "<YOUR_KNOWLEDGE_STORE_ID>";

// Define a JSON schema for the overview
const overviewSchema = {
  type: "object",
  properties: {
    total_items: { type: "integer" }, // Number of items in the knowledge store
    themes: { // Recurring themes across items
      type: "array",
      items: { type: "string" },
    },
    content_types: { // Production formats (tutorial, demo, etc.)
      type: "array",
      items: { type: "string" },
    },
    key_subjects: { // Main subjects or topics covered
      type: "array",
      items: { type: "string" },
    },
    patterns: { // Notable patterns or trends
      type: "array",
      items: { type: "string" },
    },
    summary: { type: "string" }, // Overall narrative summary
  },
};

// Generate the overview
const response = await client.responses.create({
  knowledgeStoreId: storeId,
  input: [{ type: "message", role: "user", content: "Give me a structured overview of these videos and images." }],
  text: { format: { type: "json_schema", name: "corpus_overview", schema: overviewSchema } },
});

// Parse the JSON overview
for (const output of response.output ?? []) {
  if (output.type === "message") {
    for (const content of output.content ?? []) {
      const overview = JSON.parse(content.text);
      console.log(`Items: ${overview.total_items}`);
      console.log(`Themes: ${overview.themes.join(", ")}`);
      console.log(`Summary: ${overview.summary}`);
    }
  }
}
```

## Code explanation

#### Python

To receive a structured overview, call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with a `text` parameter that describes your JSON Schema.\

**Parameters**:

* `knowledge_store_id`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey.
* `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `schema_` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\


**Return value**: An object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledge_store_id`: The knowledge store this response was generated against.
* `session_id`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `json.loads()` method into a native object.
* `usage`: Token usage statistics, including the `input_tokens` and `output_tokens` fields.

#### Node.js

To receive a structured overview, call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with a `text` parameter that describes your JSON Schema. You pass all parameters as properties of a single object.\

**Parameters**:

* `knowledgeStoreId`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey.
* `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `type` field must be `"json_schema"`, the `schema` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\


**Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledgeStoreId`: The knowledge store this response was generated against.
* `sessionId`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `JSON.parse()` method into a native object.
* `usage`: Token usage statistics, including the `inputTokens` and `outputTokens` fields.

## Example response

The `text` field inside each content part is a JSON string that matches your schema. After parsing, a typical result looks like this:

```json
{
  "total_items": 12,
  "themes": ["product onboarding", "feature announcements", "customer success", "API integration"],
  "content_types": ["tutorials", "demos", "webinars", "interviews"],
  "key_subjects": ["analytics dashboard", "API authentication", "mobile SDK", "billing workflows"],
  "patterns": ["tutorial output clusters around launch dates", "webinars include a Q&A segment in the final third"],
  "summary": "A product-focused collection of 12 videos spanning tutorials, demos, webinars, and interviews. Content centers on onboarding and feature adoption, with tutorial output concentrated around launch cycles."
}
```

> **Note**
>
> Jockey evaluates the actual content in your knowledge store. Results change depending on the items you have indexed.

# Variations

Adapt the analysis by changing the prompt or adding the `instructions` parameter.

* **Target a domain**: Add the [`instructions`](/v1.3/agents/guides/create-a-response#customize-behavior-with-instructions) parameter, such as "You are a media analyst", to shape the overview for a specific audience.
* **Compare collections**: Change the prompt to "How do these videos and images compare to a typical corporate training library?"
* **Find gaps**: Change the prompt to "What topics are underrepresented in these videos and images?"

# Jupyter notebook

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/twelvelabs-io/twelvelabs-developer-experience/blob/main/quickstarts/jockey/recipes/get_corpus_overview.ipynb)

# See also

* [Organize a video library](/v1.3/agents/recipes/organize-a-video-library) - uses a corpus overview as the first step in a full workflow
* [Structured output](/v1.3/agents/guides/create-a-response/structured-output) - more on JSON Schema responses
* [Create a response](/v1.3/api-reference/responses/create) - API reference for the Responses endpoint