> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Rate limits

> Rate limits by plan and tier. Monitor usage with response headers and handle 429 errors with best practices.

Rate limits control how many requests you can make within a specific time window. These limits ensure fair usage and maintain optimal performance for all users.

# Understand rate limits

The platform uses multiple dimensions to measure your usage. Each dimension tracks different aspects of your requests. Understanding how rate limits work helps you optimize your usage and avoid errors.

## Rate limit dimensions

The platform measures your usage across these dimensions:

* **DPD**: Duration per day (in minutes)
* **DPH**: Duration per hour (in minutes)
* **RPD**: Requests per day
* **RPM**: Requests per minute
* **TPD**: Tokens per day
* **TPM**: Tokens per minute

Rate limits vary by endpoint based on computational requirements:

* **Duration-based limits** (DPH, DPD) apply to endpoints that process video or audio content. They do not apply to embedding with Marengo 3.5.
* **Token limits** (TPM, TPD) count the output tokens on endpoints that generate text. For embedding, they count the input tokens and apply only with Marengo 3.5.
* **Request limits** (RPM, RPD) apply to all endpoints.

## Exceeding a limit

The platform checks your usage against each applicable dimension. You receive an error when you exceed any limit, even if other dimensions have remaining capacity.

For example, you might have remaining request capacity (RPM) but exceed your duration limit (DPH) when processing a long video. The platform returns an error based on the duration limit.

> **Note**
>
> When using an organization account, rate limits apply in aggregate to all the API keys in the organization.

# Your rate limits

Your plan and monthly spending determine your rate limits.

## Plans overview

TwelveLabs offers three plans:

* **Free plan**: Basic limits at no cost
* **Developer plan**: Three tiers that increase with monthly spending
* **Enterprise plan**: Custom limits based on your requirements

## Rate limit categories

Rate limits are grouped into categories. Endpoints in the same category share a rate limit. All requests to these endpoints count toward the shared limit. An endpoint can have different rate limits based on the type of content you process.

* **Index**: [`POST /tasks`](/v1.3/api-reference/upload-content/tasks/create), [`POST /indexes/{index-id}/indexed-assets`](/v1.3/api-reference/index-content/create)
* **Upload**: [`POST /assets`](/v1.3/api-reference/upload-content/direct-uploads/create)
* **Embed - Video**: [`POST /embed`](/v1.3/api-reference/create-embeddings-v1/text-image-audio-embeddings/create-text-image-audio-embeddings), [`POST /embed/tasks`](/v1.3/api-reference/create-embeddings-v1/video-embeddings/create-video-embedding-task), [`POST /embed-v2`](/v1.3/api-reference/create-embeddings-v2/create-embeddings), [`POST /embed-v2/tasks`](/v1.3/api-reference/create-embeddings-v2/create-async-embedding-task)
* **Embed - Audio**: [`POST /embed`](/v1.3/api-reference/create-embeddings-v1/text-image-audio-embeddings/create-text-image-audio-embeddings), [`POST /embed-v2`](/v1.3/api-reference/create-embeddings-v2/create-embeddings), [`POST /embed-v2/tasks`](/v1.3/api-reference/create-embeddings-v2/create-async-embedding-task)
* **Embed - Image**: [`POST /embed`](/v1.3/api-reference/create-embeddings-v1/text-image-audio-embeddings/create-text-image-audio-embeddings), [`POST /embed-v2`](/v1.3/api-reference/create-embeddings-v2/create-embeddings)
* **Embed - Document**: [`POST /embed-v2/tasks`](/v1.3/api-reference/create-embeddings-v2/create-async-embedding-task)
* **Embed - Text**: [`POST /embed`](/v1.3/api-reference/create-embeddings-v1/text-image-audio-embeddings/create-text-image-audio-embeddings), [`POST /embed-v2`](/v1.3/api-reference/create-embeddings-v2/create-embeddings)
* **Embed - Text\_Image**: [`POST /embed-v2`](/v1.3/api-reference/create-embeddings-v2/create-embeddings)
* **Search**: [`POST /search`](/v1.3/api-reference/any-to-video-search/make-search-request), [`POST /knowledge-stores/{knowledge_store_id}/search`](/v1.3/api-reference/knowledge-store-search/search)
* **Analyze**: [`POST /analyze`](/v1.3/api-reference/analyze-videos/analyze)

### Free plan

New accounts start with the Free plan at no cost.

| Category | DPD (minutes) | DPH (minutes) | RPD   | RPM | Output TPD (thousands) | Output TPM (thousands) |
| :------- | :------------ | :------------ | :---- | :-- | :--------------------- | :--------------------- |
| Index    | 3,000         | 600           | 3,000 | 60  | —                      | —                      |
| Upload   | —             | —             | 3,000 | 60  | —                      | —                      |
| Search   | —             | —             | 3,000 | 600 | —                      | —                      |
| Analyze  | 3,000         | 600           | 1,000 | 60  | 500                    | 30                     |

The embedding limits depend on the model you use:

#### Marengo 3.5

| Category         | RPD   | RPM | Input TPD (millions) | Input TPM (millions) |
| :--------------- | :---- | :-- | :------------------- | :------------------- |
| Embed - Video    | 3,000 | 25  | 500                  | 50                   |
| Embed - Audio    | 3,000 | 25  | 225                  | 20                   |
| Embed - Image    | 3,000 | 600 | 4                    | 1                    |
| Embed - Document |       |     | 10                   | 2                    |
| Embed - Text     | 3,000 | 600 | 6                    | 1                    |

The request limits for `Embed - Document` are not published.

#### Marengo 3.0

| Category            | DPD (minutes) | DPH (minutes) | RPD   | RPM |
| :------------------ | :------------ | :------------ | :---- | :-- |
| Embed - Video       | 3,000         | 600           | 3,000 | 25  |
| Embed - Audio       | 3,000         | 600           | 3,000 | 25  |
| Embed - Image       | —             | —             | 3,000 | 600 |
| Embed - Text        | —             | —             | 3,000 | 600 |
| Embed - Text\_Image | —             | —             | 3,000 | 600 |

### Developer plan

The Developer plan has three tiers. You start at Tier 1 when you add a payment method. Your tier increases automatically based on your monthly spending.

#### Tier qualifications

| Tier   | Qualification     |
| :----- | :---------------- |
| Tier 1 | Default tier      |
| Tier 2 | Spend \$200/month |
| Tier 3 | Spend \$400/month |

See the [Pricing](https://www.twelvelabs.io/pricing) page to calculate your spending.

#### Rate limits by tier

#### Tier 1

| Category | DPD (minutes) | DPH (minutes) | RPD   | RPM | Output TPD (thousands) | Output TPM (thousands) |
| :------- | :------------ | :------------ | :---- | :-- | :--------------------- | :--------------------- |
| Index    | 3,000         | 600           | 3,000 | 60  | —                      | —                      |
| Upload   | —             | —             | 3,000 | 60  | —                      | —                      |
| Search   | —             | —             | 3,000 | 600 | —                      | —                      |
| Analyze  | 3,000         | 600           | 1,000 | 60  | 500                    | 30                     |

The embedding limits depend on the model you use:

#### Marengo 3.5

| Category         | RPD   | RPM | Input TPD (millions) | Input TPM (millions) |
| :--------------- | :---- | :-- | :------------------- | :------------------- |
| Embed - Video    | 3,000 | 25  | 500                  | 50                   |
| Embed - Audio    | 3,000 | 25  | 225                  | 20                   |
| Embed - Image    | 3,000 | 600 | 4                    | 1                    |
| Embed - Document |       |     | 10                   | 2                    |
| Embed - Text     | 3,000 | 600 | 6                    | 1                    |

The request limits for `Embed - Document` are not published.

#### Marengo 3.0

| Category            | DPD (minutes) | DPH (minutes) | RPD   | RPM |
| :------------------ | :------------ | :------------ | :---- | :-- |
| Embed - Video       | 3,000         | 600           | 3,000 | 25  |
| Embed - Audio       | 3,000         | 600           | 3,000 | 25  |
| Embed - Image       | —             | —             | 3,000 | 600 |
| Embed - Text        | —             | —             | 3,000 | 600 |
| Embed - Text\_Image | —             | —             | 3,000 | 600 |

#### Tier 2

| Category | DPD (minutes) | DPH (minutes) | RPD   | RPM   | Output TPD (thousands) | Output TPM (thousands) |
| :------- | :------------ | :------------ | :---- | :---- | :--------------------- | :--------------------- |
| Index    | 6,000         | 1,200         | 6,000 | 120   | —                      | —                      |
| Upload   | —             | —             | 6,000 | 120   | —                      | —                      |
| Search   | —             | —             | 6,000 | 1,200 | —                      | —                      |
| Analyze  | 6,000         | 1,200         | 2,000 | 120   | 1,000                  | 60                     |

The embedding limits depend on the model you use:

#### Marengo 3.5

| Category         | RPD    | RPM   | Input TPD (millions) | Input TPM (millions) |
| :--------------- | :----- | :---- | :------------------- | :------------------- |
| Embed - Video    | 6,000  | 50    | 1,000                | 100                  |
| Embed - Audio    | 6,000  | 50    | 450                  | 40                   |
| Embed - Image    | 12,000 | 1,200 | 8                    | 2                    |
| Embed - Document |        |       | 20                   | 4                    |
| Embed - Text     | 12,000 | 1,200 | 12                   | 2                    |

The request limits for `Embed - Document` are not published.

#### Marengo 3.0

| Category            | DPD (minutes) | DPH (minutes) | RPD    | RPM   |
| :------------------ | :------------ | :------------ | :----- | :---- |
| Embed - Video       | 6,000         | 1,200         | 6,000  | 50    |
| Embed - Audio       | 6,000         | 1,200         | 6,000  | 50    |
| Embed - Image       | —             | —             | 12,000 | 1,200 |
| Embed - Text        | —             | —             | 12,000 | 1,200 |
| Embed - Text\_Image | —             | —             | 12,000 | 1,200 |

#### Tier 3

| Category | DPD (minutes) | DPH (minutes) | RPD    | RPM   | Output TPD (thousands) | Output TPM (thousands) |
| :------- | :------------ | :------------ | :----- | :---- | :--------------------- | :--------------------- |
| Index    | 12,000        | 2,400         | 12,000 | 240   | —                      | —                      |
| Upload   | —             | —             | 12,000 | 240   | —                      | —                      |
| Search   | —             | —             | 12,000 | 2,400 | —                      | —                      |
| Analyze  | 12,000        | 2,400         | 3,000  | 240   | 1,500                  | 120                    |

The embedding limits depend on the model you use:

#### Marengo 3.5

| Category         | RPD    | RPM   | Input TPD (millions) | Input TPM (millions) |
| :--------------- | :----- | :---- | :------------------- | :------------------- |
| Embed - Video    | 9,000  | 75    | 2,000                | 200                  |
| Embed - Audio    | 9,000  | 75    | 900                  | 80                   |
| Embed - Image    | 30,000 | 2,400 | 16                   | 4                    |
| Embed - Document |        |       | 40                   | 8                    |
| Embed - Text     | 30,000 | 2,400 | 24                   | 4                    |

The request limits for `Embed - Document` are not published.

#### Marengo 3.0

| Category            | DPD (minutes) | DPH (minutes) | RPD    | RPM   |
| :------------------ | :------------ | :------------ | :----- | :---- |
| Embed - Video       | 12,000        | 2,400         | 9,000  | 75    |
| Embed - Audio       | 12,000        | 2,400         | 9,000  | 75    |
| Embed - Image       | —             | —             | 30,000 | 2,400 |
| Embed - Text        | —             | —             | 30,000 | 2,400 |
| Embed - Text\_Image | —             | —             | 30,000 | 2,400 |

### Enterprise plan

The Enterprise plan provides custom rate limits. [Contact us](https://www.twelvelabs.io/contact) to discuss your requirements.

### Tier upgrades and downgrades

**Upgrades**: The platform upgrades your account when you reach the spending requirement for a higher tier. The upgrade takes effect immediately. You receive an email notification when your tier changes.

**Downgrades within the Developer plan**: When your monthly spending falls below your current tier's threshold, a one-month grace period applies before a downgrade.

Example:

* First month: You spend less than \$200 (below the Tier 2 threshold)
* Beginning of the second month: The platform sends you an email notification
* Second month: You spend less than \$200 again
* Beginning of the third month: The platform downgrades you to Tier 1

**Plan downgrades**: When you downgrade from the Developer plan to the Free plan, the tier change takes effect immediately, with no grace period.

## Input token limits for embedding

When you embed content with Marengo 3.5, the platform counts input tokens. For Marengo 3.0, duration limits apply instead.

Each type of content has its own limit: video, audio, image, document, and text. The platform never adds the counts of different types into a shared total. One request can count against several limits at once. A request that combines a video and an image counts against the video limit and the image limit. The response contains a `usage.input_tokens` object with the token count for each type of content. Each key identifies the limit those tokens counted against.

The platform counts the tokens of a request only after the model processes it. You can exceed a limit before you see an error:

* You still receive a success response for the request that exceeds a limit. An asynchronous task that exceeds a limit still completes.
* You receive the `HTTP 429 - Too Many Requests` error on your next request. The platform checks your limits before it processes anything, and it does not count any tokens for that request.
* A success response is not proof that you were within your limits. To check your remaining capacity, read the `Remaining` header for the type of content you sent.

The platform checks only the content you include in your request. A file can contain more than one type of content, such as a video with an audio track. The platform counts the tokens of the extra type only after it returns the response. If you have exceeded only that limit, one more request can succeed before you see the error.

# Work with rate limits

Use response headers to track your usage, implement best practices to handle errors, and upgrade your plan when you need higher limits.

## Monitor your usage

Check your usage using HTTP response headers.

### Response headers

Each response includes headers for the active dimensions:

| Header                                    | Description                                                              |
| :---------------------------------------- | :----------------------------------------------------------------------- |
| `X-Ratelimit-Dimensions`                  | The limits the request was measured against, comma-separated             |
| `X-Ratelimit-Request-Limit`               | Requests allowed per time window                                         |
| `X-Ratelimit-Request-Remaining`           | Requests remaining in current window                                     |
| `X-Ratelimit-Request-Reset`               | Unix timestamp when request limit resets                                 |
| `X-Ratelimit-Duration-Limit`              | Duration (in minutes) allowed per time window                            |
| `X-Ratelimit-Duration-Remaining`          | Duration (in minutes) remaining in current window                        |
| `X-Ratelimit-Duration-Reset`              | Unix timestamp when duration limit resets                                |
| `X-Ratelimit-Outputtoken-Limit`           | Output tokens allowed per time window                                    |
| `X-Ratelimit-Outputtoken-Remaining`       | Output tokens remaining in current window                                |
| `X-Ratelimit-Outputtoken-Reset`           | Unix timestamp when output token limit resets                            |
| `X-Ratelimit-Inputtoken-<type>-Limit`     | Input tokens allowed per time window for one type of content             |
| `X-Ratelimit-Inputtoken-<type>-Remaining` | Input tokens remaining in current window for one type of content         |
| `X-Ratelimit-Inputtoken-<type>-Reset`     | Unix timestamp when the input token limit for one type of content resets |

For the input token headers, `<type>` is `Video`, `Audio`, `Image`, `Document`, or `Text`.

> **Note**
>
> Responses include only the headers that apply to the endpoint you call.

#### Build a header name from `X-Ratelimit-Dimensions`

The `X-Ratelimit-Dimensions` header lists the label of every limit the request was measured against. The labels are comma-separated and sorted alphabetically. Replace `<label>` with an entry from the list to build the header names for that limit: `X-Ratelimit-<label>-Limit`, `X-Ratelimit-<label>-Remaining`, and `X-Ratelimit-<label>-Reset`.

The list keeps the mixed-case spelling of each label, such as `InputToken-Video`. The platform sends the headers in the canonical form, such as `X-Ratelimit-Inputtoken-Video-Limit`. HTTP header names are case-insensitive, so the two spellings match. If your code copies raw header names into a case-sensitive map, lowercase them first.

#### Legacy headers

These headers contain aggregate rate limit information. The platform maintains them for backward compatibility. Use the `X-Ratelimit-<label>-*` headers instead. The `/embed-v2` and `/embed-v2/tasks` endpoints do not send these headers. On those endpoints, only the `X-Ratelimit-Request-*` headers contain the request limit.

| Header                  | Description                               |
| :---------------------- | :---------------------------------------- |
| `X-Ratelimit-Limit`     | Maximum requests per time window          |
| `X-Ratelimit-Remaining` | Requests remaining in current time window |
| `X-Ratelimit-Used`      | Requests made in the current time window  |
| `X-Ratelimit-Reset`     | Unix timestamp when time window resets    |

### Example response headers

A request measured on requests and duration:

```
X-Ratelimit-Dimensions: Duration,Request
X-Ratelimit-Request-Limit: 600
X-Ratelimit-Request-Remaining: 542
X-Ratelimit-Request-Reset: 1735689600
X-Ratelimit-Duration-Limit: 3000
X-Ratelimit-Duration-Remaining: 2847
X-Ratelimit-Duration-Reset: 1735689600
```

A Marengo 3.5 request on the Developer plan at Tier 1, combining a video, an image, and text:

```
X-Ratelimit-Dimensions: InputToken-Image,InputToken-Text,InputToken-Video,Request
X-Ratelimit-Request-Limit: 25
X-Ratelimit-Request-Remaining: 24
X-Ratelimit-Request-Reset: 1735689600
X-Ratelimit-Inputtoken-Video-Limit: 50000000
X-Ratelimit-Inputtoken-Video-Remaining: 49994000
X-Ratelimit-Inputtoken-Video-Reset: 1735689600
X-Ratelimit-Inputtoken-Image-Limit: 1000000
X-Ratelimit-Inputtoken-Image-Remaining: 999000
X-Ratelimit-Inputtoken-Image-Reset: 1735689600
X-Ratelimit-Inputtoken-Text-Limit: 1000000
X-Ratelimit-Inputtoken-Text-Remaining: 999975
X-Ratelimit-Inputtoken-Text-Reset: 1735689600
```

## Handle rate limit errors

The platform returns an `HTTP 429 - Too Many Requests` error when you exceed a rate limit. The `message` field contains the limit you exceeded and the time it resets. The response includes a `Retry-After` header with the number of seconds to wait.

For an input token limit, you receive the error on the request after the one that exceeded it. See [Input token limits for embedding](#input-token-limits-for-embedding).

### Error response format

When you exceed a duration limit:

```json
{
  "code": "too_many_requests",
  "message": "You have exceeded the rate limit (3000sec/1hour). Please try again later after 2025-12-03T14:30:00Z."
}
```

When you exceed an input token limit:

```json
{
  "code": "too_many_requests",
  "message": "You have exceeded the rate limit (1000000token/1minute). Please try again later after 2025-12-03T14:30:00Z."
}
```

### Best practices

Follow these practices to handle rate limit errors:

* **Check response headers**: Monitor your remaining capacity before making requests
* **Respect the `Retry-After` header**: Wait the number of seconds it specifies before retrying
* **Implement exponential backoff**: Increase the time between retries
* **Distribute requests**: Spread the calls evenly instead of sending bursts
* **Cache responses**: Reuse responses when possible to reduce requests

## Increase your limits

Upgrade your plan or tier to increase rate limits:

1. **Add a payment method**: Access Tier 1
2. **Increase spending**: Reach Tier 2 or Tier 3
3. **Contact sales**: [Request custom Enterprise limits](https://www.twelvelabs.io/contact)