> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Streaming

Receive response text in real time as it is generated, instead of waiting for the complete response. Use streaming to display text as it arrives, show progress for long responses, or reduce perceived latency in a chat interface. This guide shows you how to stream a response and read the events it returns.

# Key concepts

* **Events**: The stream arrives as a series of events. Most contain a small fragment of the generated text, which you concatenate to build the full response; a final event marks the end of the stream. The stream may also emit a `keepalive` heartbeat about every 10 seconds while no other events are being emitted (for example, during a long tool call). It contains no response data and can be ignored. `sequence_number` is a single monotonic counter shared by every event type, so skipping `keepalive` frames creates gaps in it; those gaps are not dropped events. The SDK manages the underlying connection, so you handle each event as it arrives.

# Prerequisites

* You've already uploaded your videos and images, and the assets have reached the `ready` status. See the [Upload content](/v1.3/agents/guides/upload-content) page for details.
* You've already created a knowledge store. See the [Create a knowledge store](/v1.3/agents/guides/create-a-knowledge-store) page for details.
* You've already added at least one asset to the knowledge store, and the item has reached the `ready` status. See the [Add assets to a knowledge store](/v1.3/agents/guides/add-assets) page for details.
* You've already read the [Create a response](/v1.3/agents/guides/create-a-response) page and understand the basic request and response format.

# Example

Copy and paste the code below, replacing the placeholders surrounded by `<>` with your values.

**`Python`**

```python Python maxlines=20
from twelvelabs import TwelveLabs

client = TwelveLabs(api_key="<YOUR_API_KEY>")

stream = client.responses.create_stream(
    knowledge_store_id="<YOUR_KNOWLEDGE_STORE_ID>",
    input=[{"type": "message", "role": "user", "content": "Describe what these videos and images show"}],
    # session_id=...,  # Optional. Continue a multi-turn conversation. See Multi-turn sessions.
    # instructions=...,  # Optional. Add a per-request system prompt. See Create a response.
    # include=...,  # Optional. Return intermediate reasoning steps. See Create a response.
    # selections=...,  # Optional. Restrict the request to specific items or item collections.
    # text=...,  # Optional. Return typed JSON output. See Structured output.
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
```

**`Node.js`**

```javascript Node.js maxlines=20
import { TwelveLabs } from "twelvelabs-js";

const client = new TwelveLabs({ apiKey: "<YOUR_API_KEY>" });

const stream = await client.responses.createStream({
  knowledgeStoreId: "<YOUR_KNOWLEDGE_STORE_ID>",
  input: [{ type: "message", role: "user", content: "Describe what these videos and images show" }],
  // sessionId: ...,  // Optional. Continue a multi-turn conversation. See Multi-turn sessions.
  // instructions: ...,  // Optional. Add a per-request system prompt. See Create a response.
  // include: ...,  // Optional. Return intermediate reasoning steps. See Create a response.
  // selections: ...,  // Optional. Restrict the request to specific items or item collections.
  // text: ...,  // Optional. Return typed JSON output. See Structured output.
});

for await (const event of stream) {
  if (event.type === "response.output_text.delta") {
    process.stdout.write(event.delta ?? "");
  }
}
```

# Code explanation

#### Python

To stream a response as it is generated, call the [`responses.create_stream`](/v1.3/sdk-reference/python/responses#stream-a-response) method and iterate over the events it returns. It accepts the same parameters as the `responses.create` method. This example writes each incremental text fragment to the standard output as it arrives.\

**Parameters**:

* `knowledge_store_id`: The unique identifier of the knowledge store to reason over. Every request requires exactly one.
* `input`: An array of input items. Each item is a message you send to Jockey.
* *(Optional)* `session_id`: Continues a multi-turn conversation. Pass the identifier from a previous response. See the [Multi-turn sessions](/v1.3/agents/guides/create-a-response/multi-turn-sessions) page.
* *(Optional)* `instructions`: Adds a per-request system prompt that shapes Jockey's behavior for your domain or task. See the [Create a response](/v1.3/agents/guides/create-a-response#customize-behavior-with-instructions) section for details.
* *(Optional)* `include`: An array of extra items to include in the response. Pass `intermediate_outputs` to return Jockey's reasoning steps in the `output` array. See the [Create a response](/v1.3/agents/guides/create-a-response#inspect-intermediate-outputs) section for details.
* *(Optional)* `selections`: An array that restricts the request to specific knowledge store items or item collections. Omit to run against every item.
* *(Optional)* `text`: An object that specifies the response format. Provide a JSON Schema to return typed JSON output. See the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page for details.\


**Return value**: A stream of typed events. Iterate the stream: each `response.output_text.delta` event contains an incremental text fragment in its `delta` field, a `keepalive` event is a heartbeat with no response data and can be ignored, and a `response.completed` event marks the end of the stream. Concatenate the deltas to build the full response. Skipping `keepalive` frames creates gaps in `sequence_number`; those gaps are not dropped events. Citations come later than the text: the `annotations` array is empty while a content part is being generated, and the `response.content_part.done` event contains the completed part with all of its citations.

#### Node.js

To stream a response as it is generated, call the [`responses.createStream`](/v1.3/sdk-reference/node-js/responses#stream-a-response) method and iterate over the events it returns. It accepts the same parameters as the `responses.create` method. You pass all parameters as properties of a single object. This example writes each incremental text fragment to the standard output as it arrives.\

**Parameters**:

* `knowledgeStoreId`: The unique identifier of the knowledge store to reason over. Every request requires exactly one.
* `input`: An array of input items. Each item is a message you send to Jockey.
* *(Optional)* `sessionId`: Continues a multi-turn conversation. Pass the identifier from a previous response. See the [Multi-turn sessions](/v1.3/agents/guides/create-a-response/multi-turn-sessions) page.
* *(Optional)* `instructions`: Adds a per-request system prompt that shapes Jockey's behavior for your domain or task. See the [Create a response](/v1.3/agents/guides/create-a-response#customize-behavior-with-instructions) section for details.
* *(Optional)* `include`: An array of extra items to include in the response. Pass `intermediate_outputs` to return Jockey's reasoning steps in the `output` array. See the [Create a response](/v1.3/agents/guides/create-a-response#inspect-intermediate-outputs) section for details.
* *(Optional)* `selections`: An array that restricts the request to specific knowledge store items or item collections. Omit to run against every item.
* *(Optional)* `text`: An object that specifies the response format. Provide a JSON Schema to return typed JSON output. See the [Structured output](/v1.3/agents/guides/create-a-response/structured-output) page for details.\


**Return value**: A stream of typed events. Iterate the stream: each `response.output_text.delta` event contains an incremental text fragment in its `delta` field, a `keepalive` event is a heartbeat with no response data and can be ignored, and a `response.completed` event marks the end of the stream. Concatenate the deltas to build the full response. Skipping `keepalive` frames creates gaps in `sequence_number`; those gaps are not dropped events. Citations come later than the text: the `annotations` array is empty while a content part is being generated, and the `response.content_part.done` event contains the completed part with all of its citations.

# Example response

Each `response.output_text.delta` event contains a fragment of the generated text in its `delta` field:

```json
{"type": "response.output_text.delta", "delta": "The main"}
```

Events arrive incrementally. Concatenate the deltas to build the full response.

The stream may also emit a `keepalive` heartbeat about every 10 seconds when no other events are being emitted, for example while a tool call is still running. It carries no response data:

```json
{"type": "keepalive", "sequence_number": 42}
```

`sequence_number` uses the same counter as all other event types. If you skip `keepalive` frames, the numbers you see will have gaps; these gaps do not mean events were dropped.

The `response.content_part.done` event marks the point where a content part is complete. It contains the full text and all citations:

```json
{"type": "response.content_part.done", "item_id": "msg_abc123", "output_index": 0, "content_index": 0, "part": {"type": "output_text", "text": "...", "annotations": []}}
```

# Common pitfalls

* **Iterate to completion.** Consume the full stream so the connection closes cleanly and you receive every text fragment.
* **Concatenate the deltas.** Each `response.output_text.delta` event contains a fragment of the generated text; concatenate them in arrival order to build the full response.
* **Reconnect manually if the stream drops.** If the stream ends before completion, start a new request. Reconnection is not built in for this endpoint.

# Next steps

* [Structured output](/v1.3/agents/guides/create-a-response/structured-output) - receive typed JSON by providing a schema
* [Multi-turn sessions](/v1.3/agents/guides/create-a-response/multi-turn-sessions) - maintain conversation context across requests
* [Create a response](/v1.3/agents/guides/create-a-response) - review the basic request and response format

# Jupyter notebook

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/twelvelabs-io/twelvelabs-developer-experience/blob/main/quickstarts/jockey/guides/streaming.ipynb)