> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.twelvelabs.io/v1.3/agents/recipes/get-a-corpus-overview/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server. # Generate a corpus overview > Summarize all your videos and images (themes, subjects, content types, patterns, and key statistics) in a single request. This recipe shows you how to summarize all your videos and images (themes, subjects, content types, patterns, and key statistics) in a single request. You can receive the overview as prose or as structured JSON. **Use cases**: * **Collection onboarding**: Summarize a new collection before starting development * **Dashboard generation**: Create collection summaries for a UI or reporting tool * **Pipeline input**: Feed collection context to a downstream agent or workflow # Key concepts * **Knowledge store**: A persistent store of your videos and images plus the understanding the platform derives from them - spatiotemporal context, a typed ontology, and embeddings - that together enable corpus-level reasoning. * **Structured output**: A JSON schema you provide with your request. Jockey constrains its output to match your schema, so you retrieve machine-readable results instead of plain text. For a guide on designing schemas and parsing responses, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page. # Workflow This recipe shows two approaches: a plain-text overview that needs only a prompt, and a structured overview that adds a JSON schema to define the output format. [Jockey](/v1.3/agents/concepts/jockey) reasons across your knowledge store, identifies themes and patterns, and returns results in the format you chose. You can adapt this recipe to target a specific domain or ask a different question about your collection. # Prerequisites * You've already created a knowledge store with at least one item in `ready` status. See the [Create a knowledge store](/v1.3/agents/guides/create-a-knowledge-store) and [Add assets to a knowledge store](/v1.3/agents/guides/add-assets) guides for details. * You're familiar with the request and response format. See the [Create a response](/v1.3/agents/guides/create-a-response) page for details. # Plain-text overview Return the overview as prose. This approach needs only a prompt. ## Example Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values. **`Python`** ```python Python maxlines=15 from twelvelabs import TwelveLabs client = TwelveLabs(api_key="") STORE_ID = "" # Generate the overview response = client.responses.create( knowledge_store_id=STORE_ID, input=[{"type": "message", "role": "user", "content": "Give me a comprehensive overview of these videos and images. Include main themes, recurring subjects, content types, and any notable patterns."}], ) # Read the overview text for output in response.output: if output.type == "message": for content in output.content: print(content.text) ``` **`Node.js`** ```javascript Node.js maxlines=15 import { TwelveLabs } from "twelvelabs-js"; const client = new TwelveLabs({ apiKey: "" }); const storeId = ""; // Generate the overview const response = await client.responses.create({ knowledgeStoreId: storeId, input: [{ type: "message", role: "user", content: "Give me a comprehensive overview of these videos and images. Include main themes, recurring subjects, content types, and any notable patterns." }], }); // Read the overview text for (const output of response.output ?? []) { if (output.type === "message") { for (const content of output.content ?? []) { console.log(content.text); } } } ``` ## Code explanation #### Python To generate a plain-text overview, call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with a prompt that requests a collection summary.\ **Parameters**: * `knowledge_store_id`: The unique identifier of the knowledge store to reason over. * `input`: An array of input items. Each item is a message you send to Jockey. This example requests a comprehensive overview covering themes, subjects, content types, and patterns.\ **Return value**: An object of type `ResponseObject` containing, among other information, the following fields: * `id`: The unique identifier of the response. * `knowledge_store_id`: The knowledge store this response was generated against. * `session_id`: The session identifier. Pass it in a follow-up request to continue the conversation. * `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`. * `output`: The response output items. Read the `text` field of each content part in a `message` item for the overview. * `usage`: Token usage statistics, including the `input_tokens` and `output_tokens` fields. #### Node.js To generate a plain-text overview, call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with a prompt that requests a collection summary. You pass all parameters as properties of a single object.\ **Parameters**: * `knowledgeStoreId`: The unique identifier of the knowledge store to reason over. * `input`: An array of input items. Each item is a message you send to Jockey. This example requests a comprehensive overview covering themes, subjects, content types, and patterns.\ **Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject` containing, among other information, the following fields: * `id`: The unique identifier of the response. * `knowledgeStoreId`: The knowledge store this response was generated against. * `sessionId`: The session identifier. Pass it in a follow-up request to continue the conversation. * `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`. * `output`: The response output items. Read the `text` field of each content part in a `message` item for the overview. * `usage`: Token usage statistics, including the `inputTokens` and `outputTokens` fields. ## Example response The `text` field inside each content part contains Jockey's overview formatted as Markdown. A typical response looks like this: ```text ### Overall character - **Content type:** product tutorials, customer interviews, and feature demos - **Primary audience:** developers and product managers - **Structure:** a growing library of short-form instructional and promotional content ### Main themes - **Product onboarding**: step-by-step walkthroughs of initial setup and core workflows - **Feature announcements**: quarterly release highlights and capability demos - **Customer success**: interviews and testimonials from active users ### Recurring subjects - Dashboard and analytics tooling - API authentication and integration workflows - Mobile SDK setup and troubleshooting ### Notable patterns - Tutorial output clusters around product launch dates - Most webinars include a Q&A segment in the final third - Interview content appears on a quarterly cadence ``` > **Note** > > Jockey returns Markdown-formatted text. If your downstream pipeline expects plain prose, strip the Markdown formatting before processing. # Structured overview Return the overview as JSON that matches a schema you define. Use this approach when downstream code expects typed fields. ## Example Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values. **`Python`** ```python Python maxlines=35 import json from twelvelabs import TwelveLabs, TextParam from twelvelabs.types.text_param_format import TextParamFormat_JsonSchema client = TwelveLabs(api_key="") STORE_ID = "" # Define a JSON schema for the overview overview_schema = { "type": "object", "properties": { "total_items": {"type": "integer"}, # Number of items in the knowledge store "themes": { # Recurring themes across items "type": "array", "items": {"type": "string"}, }, "content_types": { # Production formats (tutorial, demo, etc.) "type": "array", "items": {"type": "string"}, }, "key_subjects": { # Main subjects or topics covered "type": "array", "items": {"type": "string"}, }, "patterns": { # Notable patterns or trends "type": "array", "items": {"type": "string"}, }, "summary": {"type": "string"}, # Overall narrative summary }, } # Generate the overview response = client.responses.create( knowledge_store_id=STORE_ID, input=[{"type": "message", "role": "user", "content": "Give me a structured overview of these videos and images."}], text=TextParam(format=TextParamFormat_JsonSchema(name="corpus_overview", schema_=overview_schema)), ) # Parse the JSON overview for output in response.output: if output.type == "message": for content in output.content: overview = json.loads(content.text) print(f"Items: {overview['total_items']}") print(f"Themes: {', '.join(overview['themes'])}") print(f"Summary: {overview['summary']}") ``` **`Node.js`** ```javascript Node.js maxlines=35 import { TwelveLabs } from "twelvelabs-js"; const client = new TwelveLabs({ apiKey: "" }); const storeId = ""; // Define a JSON schema for the overview const overviewSchema = { type: "object", properties: { total_items: { type: "integer" }, // Number of items in the knowledge store themes: { // Recurring themes across items type: "array", items: { type: "string" }, }, content_types: { // Production formats (tutorial, demo, etc.) type: "array", items: { type: "string" }, }, key_subjects: { // Main subjects or topics covered type: "array", items: { type: "string" }, }, patterns: { // Notable patterns or trends type: "array", items: { type: "string" }, }, summary: { type: "string" }, // Overall narrative summary }, }; // Generate the overview const response = await client.responses.create({ knowledgeStoreId: storeId, input: [{ type: "message", role: "user", content: "Give me a structured overview of these videos and images." }], text: { format: { type: "json_schema", name: "corpus_overview", schema: overviewSchema } }, }); // Parse the JSON overview for (const output of response.output ?? []) { if (output.type === "message") { for (const content of output.content ?? []) { const overview = JSON.parse(content.text); console.log(`Items: ${overview.total_items}`); console.log(`Themes: ${overview.themes.join(", ")}`); console.log(`Summary: ${overview.summary}`); } } } ``` ## Code explanation #### Python To receive a structured overview, call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with a `text` parameter that describes your JSON Schema.\ **Parameters**: * `knowledge_store_id`: The unique identifier of the knowledge store to reason over. * `input`: An array of input items. Each item is a message you send to Jockey. * `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `schema_` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\ **Return value**: An object of type `ResponseObject` containing, among other information, the following fields: * `id`: The unique identifier of the response. * `knowledge_store_id`: The knowledge store this response was generated against. * `session_id`: The session identifier. Pass it in a follow-up request to continue the conversation. * `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`. * `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `json.loads()` method into a native object. * `usage`: Token usage statistics, including the `input_tokens` and `output_tokens` fields. #### Node.js To receive a structured overview, call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with a `text` parameter that describes your JSON Schema. You pass all parameters as properties of a single object.\ **Parameters**: * `knowledgeStoreId`: The unique identifier of the knowledge store to reason over. * `input`: An array of input items. Each item is a message you send to Jockey. * `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `type` field must be `"json_schema"`, the `schema` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\ **Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject` containing, among other information, the following fields: * `id`: The unique identifier of the response. * `knowledgeStoreId`: The knowledge store this response was generated against. * `sessionId`: The session identifier. Pass it in a follow-up request to continue the conversation. * `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`. * `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `JSON.parse()` method into a native object. * `usage`: Token usage statistics, including the `inputTokens` and `outputTokens` fields. ## Example response The `text` field inside each content part is a JSON string that matches your schema. After parsing, a typical result looks like this: ```json { "total_items": 12, "themes": ["product onboarding", "feature announcements", "customer success", "API integration"], "content_types": ["tutorials", "demos", "webinars", "interviews"], "key_subjects": ["analytics dashboard", "API authentication", "mobile SDK", "billing workflows"], "patterns": ["tutorial output clusters around launch dates", "webinars include a Q&A segment in the final third"], "summary": "A product-focused collection of 12 videos spanning tutorials, demos, webinars, and interviews. Content centers on onboarding and feature adoption, with tutorial output concentrated around launch cycles." } ``` > **Note** > > Jockey evaluates the actual content in your knowledge store. Results change depending on the items you have indexed. # Variations Adapt the analysis by changing the prompt or adding the `instructions` parameter. * **Target a domain**: Add the [`instructions`](/v1.3/agents/guides/create-a-response#customize-behavior-with-instructions) parameter, such as "You are a media analyst", to shape the overview for a specific audience. * **Compare collections**: Change the prompt to "How do these videos and images compare to a typical corporate training library?" * **Find gaps**: Change the prompt to "What topics are underrepresented in these videos and images?" # Jupyter notebook [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/twelvelabs-io/twelvelabs-developer-experience/blob/main/quickstarts/jockey/recipes/get_corpus_overview.ipynb) # See also * [Organize a video library](/v1.3/agents/recipes/organize-a-video-library) - uses a corpus overview as the first step in a full workflow * [Structured output](/v1.3/agents/guides/create-a-response/structured-output) - more on JSON Schema responses * [Create a response](/v1.3/api-reference/responses/create) - API reference for the Responses endpoint > Summarize all your videos and images (themes, subjects, content types, patterns, and key statistics) in a single request.