> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.twelvelabs.io/v1.3/agents/recipes/track-entities-across-videos/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server. # Track entities across videos > Follow a subject across multiple videos and build a chronological timeline of appearances with timestamps, locations, and context. This recipe shows you how to follow a subject across multiple videos and build a chronological timeline of appearances. The response includes timestamps, video references, locations, and context for each sighting, ready for your tracking dashboard or analysis pipeline. > **Note** > > Cross-video entity tracking is in research preview. Behavior and accuracy may change as the feature develops. **Use cases**: * **Security investigations**: Track a person across multiple camera angles or footage sources * **Timeline construction**: Build a chronological record of a subject's appearances * **Content auditing**: Identify every moment where a specific entity appears in a collection # Key concepts * **Knowledge store**: A persistent store of your videos and images plus the understanding the platform derives from them - spatiotemporal context, a typed ontology, and embeddings - that together enable corpus-level reasoning. * **Instructions**: Additional guidance that shapes Jockey's behavior for a specific domain or task. * **Structured output**: A JSON schema you provide with your request. Jockey constrains its output to match your schema, so you retrieve machine-readable results instead of plain text. For a guide on designing schemas and parsing responses, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page. # Workflow This recipe combines three elements: a JSON schema that defines the appearance data structure, instructions that give [Jockey](/v1.3/agents/concepts/jockey) a specific role and domain guidance, and a prompt that describes the subject to track. Jockey reasons across your knowledge store and returns structured tracking results. You can then continue the conversation using the session identifier from the response. For example, drill into a specific appearance or ask about the surrounding context. You can adapt this recipe to track different subjects or focus on different appearance details. # Prerequisites * You've already created a knowledge store with at least one item in `ready` status. See the [Create a knowledge store](/v1.3/agents/guides/create-a-knowledge-store) and [Add assets to a knowledge store](/v1.3/agents/guides/add-assets) guides for details. * You're familiar with the request and response format. See the [Create a response](/v1.3/agents/guides/create-a-response) page for details. # Complete example Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values. **`Python`** ```python Python maxlines=45 import json from twelvelabs import TwelveLabs, TextParam from twelvelabs.types.text_param_format import TextParamFormat_JsonSchema client = TwelveLabs(api_key="") STORE_ID = "" # Define a JSON schema for the tracking timeline tracking_schema = { "type": "object", "properties": { "subject": {"type": "string"}, # Description of the tracked subject "timeline": { "type": "array", "items": { "type": "object", "properties": { "timestamp": {"type": "string"}, # When the appearance occurs "video_reference": {"type": "string"}, # Source video identifier "location": {"type": "string"}, # Where in the scene "action": {"type": "string"}, # What the subject is doing }, }, }, "summary": {"type": "string"}, # Overall tracking summary }, } # Step 1: Track the subject response = client.responses.create( knowledge_store_id=STORE_ID, instructions="You are a security analyst. Prioritize temporal accuracy. Flag any low-confidence identifications.", input=[{"type": "message", "role": "user", "content": "Track the person in the blue hoodie across all cameras. Give me a chronological timeline."}], text=TextParam(format=TextParamFormat_JsonSchema(name="entity_tracking", schema_=tracking_schema)), ) # Step 2: Parse and read the timeline for output in response.output: if output.type == "message": for content in output.content: data = json.loads(content.text) print(f"Subject: {data['subject']}") print(f"Summary: {data['summary']}\n") for entry in data["timeline"]: print(f" [{entry['timestamp']}] {entry['video_reference']}") print(f" Location: {entry['location']}") print(f" Action: {entry['action']}") # Step 3: Refine with a follow-up turn follow_up = client.responses.create( knowledge_store_id=STORE_ID, session_id=response.session_id, input=[{"type": "message", "role": "user", "content": "Tell me more about what happened at the third timestamp. What was happening around the subject?"}], ) # Read the follow-up response for output in follow_up.output: if output.type == "message": for content in output.content: print(content.text) ``` **`Node.js`** ```javascript Node.js maxlines=45 import { TwelveLabs } from "twelvelabs-js"; const client = new TwelveLabs({ apiKey: "" }); const storeId = ""; // Define a JSON schema for the tracking timeline const trackingSchema = { type: "object", properties: { subject: { type: "string" }, // Description of the tracked subject timeline: { type: "array", items: { type: "object", properties: { timestamp: { type: "string" }, // When the appearance occurs video_reference: { type: "string" }, // Source video identifier location: { type: "string" }, // Where in the scene action: { type: "string" }, // What the subject is doing }, }, }, summary: { type: "string" }, // Overall tracking summary }, }; // Step 1: Track the subject const response = await client.responses.create({ knowledgeStoreId: storeId, instructions: "You are a security analyst. Prioritize temporal accuracy. Flag any low-confidence identifications.", input: [{ type: "message", role: "user", content: "Track the person in the blue hoodie across all cameras. Give me a chronological timeline." }], text: { format: { type: "json_schema", name: "entity_tracking", schema: trackingSchema } }, }); // Step 2: Parse and read the timeline for (const output of response.output ?? []) { if (output.type === "message") { for (const content of output.content ?? []) { const data = JSON.parse(content.text); console.log(`Subject: ${data.subject}`); console.log(`Summary: ${data.summary}\n`); for (const entry of data.timeline) { console.log(` [${entry.timestamp}] ${entry.video_reference}`); console.log(` Location: ${entry.location}`); console.log(` Action: ${entry.action}`); } } } } // Step 3: Refine with a follow-up turn const followUp = await client.responses.create({ knowledgeStoreId: storeId, sessionId: response.sessionId, input: [{ type: "message", role: "user", content: "Tell me more about what happened at the third timestamp. What was happening around the subject?" }], }); // Read the follow-up response for (const output of followUp.output ?? []) { if (output.type === "message") { for (const content of output.content ?? []) { console.log(content.text); } } } ``` # Code explanation #### Python #### Track the subject Search the collection for the subject and return a structured timeline of appearances.\ **Function call**: You call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method with tracking `instructions`, a prompt, and a `text` parameter that describes your timeline schema.\ **Parameters**: * `knowledge_store_id`: The unique identifier of the knowledge store to search. * `input`: An array of input items. Each item is a message you send to Jockey. This example describes the subject to track and requests a chronological timeline. * `instructions`: A per-request system prompt that shapes Jockey's behavior for a domain or task. This example uses a security analyst role that prioritizes temporal accuracy. * `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `schema_` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\ > **Tip** > > Describe your subject with specific physical details. Clothing, features, and distinguishing characteristics improve tracking accuracy. **Return value**: An object of type `ResponseObject`. Read the `session_id` field to continue the conversation, and the `output` items for the timeline. The response also includes `id`, `status`, and `usage`. #### Process the results The `text` field of each content part in a `message` item of the `output` array is a JSON string that matches your schema. Parse it with the `json.loads()` method, then iterate the `timeline` array to read each appearance. #### Refine with a follow-up turn Continue the conversation with the `session_id` from the first response. Jockey retains the full context from the previous turn, so you can drill into a specific appearance without repeating the original prompt.\ **Function call**: You call the [`responses.create`](/v1.3/sdk-reference/python/responses#create-a-response) method again with the `session_id`.\ **Parameters**: * `knowledge_store_id`: The unique identifier of the knowledge store. Must match the store from the first request. * `session_id`: The session identifier from the first response. * `input`: An array of input items. Each item is a message you send to Jockey. This example asks for more detail about a specific timestamp.\ **Return value**: An object of type `ResponseObject`. This turn omits the `text` parameter, so the `text` field of each content part in a `message` item is plain text. Read it directly. #### Node.js #### Track the subject Search the collection for the subject and return a structured timeline of appearances.\ **Function call**: You call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method with tracking `instructions`, a prompt, and a `text` parameter that describes your timeline schema. You pass all parameters as properties of a single object.\ **Parameters**: * `knowledgeStoreId`: The unique identifier of the knowledge store to search. * `input`: An array of input items. Each item is a message you send to Jockey. This example describes the subject to track and requests a chronological timeline. * `instructions`: A per-request system prompt that shapes Jockey's behavior for a domain or task. This example uses a security analyst role that prioritizes temporal accuracy. * `text`: An object that specifies the response format. To return structured JSON, set the `format` field. Within it, the `type` field must be `"json_schema"`, the `schema` field defines the structure the output must match, and the `name` field identifies the schema. Design the schema to fit your use case; the inline comments in the example describe each field. For the schema rules, see the [Structured output](/v1.3/agents/guides/create-a-response/structured-output#json-schema-requirements) page.\ > **Tip** > > Describe your subject with specific physical details. Clothing, features, and distinguishing characteristics improve tracking accuracy. **Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject`. Read the `sessionId` field to continue the conversation, and the `output` items for the timeline. The response also includes `id`, `status`, and `usage`. #### Process the results The `text` field of each content part in a `message` item of the `output` array is a JSON string that matches your schema. Parse it with the `JSON.parse()` method, then iterate the `timeline` array to read each appearance. #### Refine with a follow-up turn Continue the conversation with the `sessionId` from the first response. Jockey retains the full context from the previous turn, so you can drill into a specific appearance without repeating the original prompt.\ **Function call**: You call the [`responses.create`](/v1.3/sdk-reference/node-js/responses#create-a-response) method again with the `sessionId`.\ **Parameters**: * `knowledgeStoreId`: The unique identifier of the knowledge store. Must match the store from the first request. * `sessionId`: The session identifier from the first response. * `input`: An array of input items. Each item is a message you send to Jockey. This example asks for more detail about a specific timestamp.\ **Return value**: An `HttpResponsePromise` that resolves to an object of type `ResponseObject`. This turn omits the `text` parameter, so the `text` field of each content part in a `message` item is plain text. Read it directly. # Example response The `text` field inside each content part is a JSON string that matches your schema. After parsing, a typical result looks like this: ```json { "subject": "Person in blue hoodie", "timeline": [ { "timestamp": "02:15-03:40", "video_reference": "069eb4e8-aeb0-7e83-8000-86413fcc296a", "location": "Building lobby, near the front entrance", "action": "Enters through the main door and walks toward the elevator" }, { "timestamp": "05:42-06:10", "video_reference": "069e1e97-27f8-7a8e-8000-bf32ecd7fc8c", "location": "Third floor hallway", "action": "Exits the elevator and walks down the corridor" } ], "summary": "Subject appears in 2 videos over a 3-minute span, moving from the lobby to the third floor." } ``` > **Notes** > > * The `video_reference` field contains the asset identifier of the knowledge store item. > * The `timestamp` field is a range in `MM:SS-MM:SS` format indicating when the appearance starts and ends. > * If Jockey cannot find the subject in the collection (for example, because the footage does not contain a match or the subject description is ambiguous), it returns an empty array named `timeline` and explains the result in the `summary` field. > * Jockey identifies subjects visually and does not match by voice. > * Each request tracks one subject. To track multiple subjects, make separate requests. > * Tracking accuracy depends on video quality, camera angles, lighting, and how distinctive the subject appears. # Variations Change the instructions and prompt to adapt this recipe for different tracking scenarios. * **Change the analytical perspective**: Set the `instructions` parameter to "documentary researcher" or "video editor" for different emphasis. * **Track objects instead of people**: Change the prompt to track a vehicle, branded object, or recurring visual element. * **Multi-episode tracking**: Track a recurring character or subject across an episode series. # See also * [Extract entities](/v1.3/agents/recipes/extract-entities) - list all entities before deciding which to track * [Structured output](/v1.3/agents/guides/create-a-response/structured-output) - more on JSON Schema responses * [Multi-turn sessions](/v1.3/agents/guides/create-a-response/multi-turn-sessions) - more on continuing conversations * [Create a response](/v1.3/api-reference/responses/create) - API reference for the Responses endpoint > Follow a subject across multiple videos and build a chronological timeline of appearances with timestamps, locations, and context.