> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Extract entities

> List all entities found across your videos and images (people, places, objects, brands, and concepts) with their type, frequency, and the items they appear in.

This recipe shows you how to extract every distinct entity from your videos and images. The response includes each entity's name, type, frequency, and the items it appears in, ready for a content index, tagging system, or downstream pipeline.

**Use cases**:

* **Content indexing**: Build a searchable catalog of who and what appears across your videos and images
* **Classification pipelines**: Feed entity lists to downstream tagging or categorization systems
* **Cross-video analysis**: Find which entities appear most frequently or span multiple videos

# Key concepts

* **Knowledge store**: A persistent store of your videos and images plus the understanding the platform derives from them - spatiotemporal context, a typed ontology, and embeddings - that together enable corpus-level reasoning.
* **Structured output**: A JSON schema you provide with your request. Jockey constrains its output to match your schema, so you retrieve machine-readable results instead of plain text. For a guide on designing schemas and parsing responses, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page.

# Workflow

This recipe combines two elements: a JSON schema that defines the entity data structure, and a prompt that describes what to extract. [Jockey](/v1.3/agents/concepts/jockey) reasons across your knowledge store, identifies every distinct entity, and returns structured results that match your schema. You can adapt this recipe to extract different entity types or apply different extraction criteria.

# Prerequisites

* You've already created a knowledge store with at least one item in `ready` status. See the [Create a knowledge store](/v1.3/agents/guides/create-a-knowledge-store) and [Add assets to a knowledge store](/v1.3/agents/guides/add-assets) guides for details.
* You're familiar with the request and response format. See the [Create a response](/v1.3/agents/guides/create-a-response) page for details.

# Complete example

Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=38
import json
from twelvelabs import TwelveLabs, TextParam
from twelvelabs.types.text_param_format import TextParamFormat_JsonSchema

client = TwelveLabs(api_key="<YOUR_API_KEY>")
STORE_ID = "<YOUR_KNOWLEDGE_STORE_ID>"

# Define a JSON schema for the entity list
entity_schema = {
    "type": "object",
    "properties": {
        "entities": {
            "type": "array",
            "items": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},            # Entity name (person, place, brand, etc.)
                    "type": {"type": "string"},            # Category: person, place, object, brand, concept
                    "frequency": {"type": "string"},       # How often the entity appears
                    "appears_in": {                        # Items containing this entity
                        "type": "array",
                        "items": {"type": "string"},
                    },
                },
            },
        },
        "entity_count": {"type": "integer"},                # Total number of distinct entities
    },
}

# Extract the entities
response = client.responses.create(
    knowledge_store_id=STORE_ID,
    input=[{"type": "message", "role": "user", "content": "List every distinct entity across all videos and images - people, places, objects, brands, and concepts. Include how frequently each appears."}],
    text=TextParam(format=TextParamFormat_JsonSchema(name="entity_list", schema_=entity_schema)),
)

# Parse and read the entities
for output in response.output:
    if output.type == "message":
        for content in output.content:
            data = json.loads(content.text)
            print(f"Found {data['entity_count']} entities:\n")
            for entity in data["entities"]:
                print(f"  [{entity['type']}] {entity['name']} - {entity['frequency']}")
                if entity.get("appears_in"):
                    print(f"    Items: {', '.join(entity['appears_in'])}")
```

**`Node.js`**

```javascript Node.js maxlines=38
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });
const storeId = "<YOUR_KNOWLEDGE_STORE_ID>";

// Define a JSON schema for the entity list
const entitySchema = {
  type: "object",
  properties: {
    entities: {
      type: "array",
      items: {
        type: "object",
        properties: {
          name: { type: "string" }, // Entity name (person, place, brand, etc.)
          type: { type: "string" }, // Category: person, place, object, brand, concept
          frequency: { type: "string" }, // How often the entity appears
          appears_in: { // Items containing this entity
            type: "array",
            items: { type: "string" },
          },
        },
      },
    },
    entity_count: { type: "integer" }, // Total number of distinct entities
  },
};

// Extract the entities
const response = await client.responses.create({
  knowledgeStoreId: storeId,
  input: [{ type: "message", role: "user", content: "List every distinct entity across all videos and images - people, places, objects, brands, and concepts. Include how frequently each appears." }],
  text: { format: { type: "json_schema", name: "entity_list", schema: entitySchema } },
});

// Parse and read the entities
for (const output of response.output ?? []) {
  if (output.type === "message") {
    for (const content of output.content ?? []) {
      const data = JSON.parse(content.text);
      console.log(`Found ${data.entity_count} entities:\n`);
      for (const entity of data.entities) {
        console.log(`  [${entity.type}] ${entity.name} - ${entity.frequency}`);
        if (entity.appears_in) {
          console.log(`    Items: ${entity.appears_in.join(", ")}`);
        }
      }
    }
  }
}
```

# Code explanation

#### Python

To extract entities, call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with a prompt and a `text` parameter that describes your result schema.\

**Parameters**:

* `knowledge_store_id`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey. This example asks Jockey to list every distinct entity across all videos and images, including frequency.
* `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `schema_` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\


**Return value**: An object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledge_store_id`: The knowledge store this response was generated against.
* `session_id`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `json.loads()` method, then iterate the `entities` array to read each entity.
* `usage`: Token usage statistics, including the `input_tokens` and `output_tokens` fields.

#### Node.js

To extract entities, call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with a prompt and a `text` parameter that describes your result schema. You pass all parameters as properties of a single object.\

**Parameters**:

* `knowledgeStoreId`: The unique identifier of the knowledge store to reason over.
* `input`: An array of input items. Each item is a message you send to Jockey. This example asks Jockey to list every distinct entity across all videos and images, including frequency.
* `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `type` field must be `"json_schema"`, the `schema` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\


**Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject` containing, among other information, the following fields:

* `id`: The unique identifier of the response.
* `knowledgeStoreId`: The knowledge store this response was generated against.
* `sessionId`: The session identifier. Pass it in a follow-up request to continue the conversation.
* `status`: The status of the response. The possible values are `completed`, `failed`, `in_progress`, and `incomplete`.
* `output`: The response output items. The `text` field of each content part in a `message` item is a JSON string that matches your schema. Parse it with the `JSON.parse()` method, then iterate the `entities` array to read each entity.
* `usage`: Token usage statistics, including the `inputTokens` and `outputTokens` fields.

# Example response

The `text` field inside each content part is a JSON string that matches your schema. After parsing, a typical result looks like this:

```json
{
  "entities": [
    {
      "name": "Sarah Chen",
      "type": "person",
      "frequency": "42 spans",
      "appears_in": ["069eb4e8-aeb0-7e83-8000-86413fcc296a", "069e1e97-27f8-7a8e-8000-bf32ecd7fc8c"]
    },
    {
      "name": "Dashboard Analytics",
      "type": "concept",
      "frequency": "18 spans",
      "appears_in": ["069eb4e8-aeb0-7e83-8000-86413fcc296a"]
    },
    {
      "name": "San Francisco",
      "type": "place",
      "frequency": "5 spans",
      "appears_in": ["069eb4e8-aeb0-7e83-8000-86413fcc296a", "069e1e97-27f8-7a8e-8000-bf32ecd7fc8c"]
    }
  ],
  "entity_count": 3
}
```

> **Notes**
>
> * The `frequency` field reports the number of indexed spans where the entity appears, not the number of videos.
> * The `appears_in` field lists the asset identifiers of the knowledge store items that contain the entity.

# Variations

Change the prompt to adapt this recipe for different extraction needs.

* **Filter by entity type**: Change the prompt to "List only the people who appear in these videos and images."
* **Cross-video presence**: Change the prompt to "Which entities appear in more than one video?"
* **Include relationships**: Change the prompt to "List entities and describe how they relate to each other."

# Jupyter notebook

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/twelvelabs-io/twelvelabs-developer-experience/blob/main/quickstarts/jockey/recipes/extract_entities.ipynb)

# See also

* [Track entities across videos](/v1.3/agents/recipes/track-entities-across-videos) - follow a specific entity across multiple videos
* [Organize a video library](/v1.3/agents/recipes/organize-a-video-library) - use entity extraction for categorization
* [Structured output](/v1.3/agents/guides/create-a-response/structured-output) - more on JSON Schema responses
* [Create a response](/v1.3/api-reference/responses/create) - API reference for the Responses endpoint