> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Modalities

> Understand how the platform analyzes your video content.

Modalities represent the sources of information that the platform processes and analyzes in a video.

**Visual includes**:

* Actions, objects, and events in the video.
* Text that appears on screen (through OCR).
* Brand logos and visual elements.

**Audio includes**:

* Ambient sounds, music, and sound effects.
* Human speech (Agents).
* Non-speech audio only (Models). For speech content, use the transcription modality.

**Transcription includes**:

* Spoken words extracted from the audio track. In Agents, speech is part of the audio modality instead.

You specify modalities through different parameters depending on your task:

* **Model options**: when you create an index.
* **Search options**: when you search videos.
* **Embedding option**: when you retrieve embeddings.

# Model options

When you create an index, specify which modalities the platform must process. You can include the following values in the [`model_options`](/v1.3/api-reference/indexes/create#request.body.models.model_options) array:

* `visual`: To process visual content
* `audio`: To process audio content

You can enable one or both model options. The platform processes only the modalities you specify.

## Related topics

* [Python SDK Reference > Create an index](/v1.3/sdk-reference/python/manage-indexes#create-an-index)
* [Node.js SDK Reference > Create an index](/v1.3/sdk-reference/node-js/manage-indexes#create-an-index)
* [API Reference > Create an index](/v1.3/api-reference/indexes/create)

# Search options

The modalities you can search, and the parameter you pass them in, depend on which product you use.

## Models

When you search videos, use the [`search_options`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.search_options.search_options) parameter to specify which modalities the platform uses to find relevant matches.

Marengo separates audio into speech and non-speech content.

**To find visual content**:

Set `search_options` to `visual` to search for:

* Actions, objects, and events in the video
* Text that appears on screen (through OCR)
* Brand logos and visual elements

**Example use cases:**

* Finding scenes with specific objects: "red car in parking lot"
* Locating on-screen text: "company logo on building"
* Identifying actions: "person running"

**To find non-speech audio**:

Set `search_options` to `audio` to search for sounds other than human speech:

* Musical tones and melodies
* Beeping, alarms, and mechanical sounds
* Environmental sounds (rain, traffic, nature)

**Example use cases:**

* Finding background music: "upbeat electronic music"
* Locating sound effects: "door slamming"
* Identifying ambient sounds: "rainfall"

**Find spoken words**

Set `search_options` to `transcription` to search the spoken content in your videos.

**Example use cases:**

* Finding mentions of topics: "climate change discussion"
* Locating product names: "iPhone 15 Pro Max"
* Identifying speakers discussing concepts: "quarterly revenue growth"

### Transcription options

Use the [`transcription_options`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.transcription_options.transcription_options) parameter to specify how the platform matches your query against spoken words:

* `lexical`: Matches the exact words or phrases in your query, allowing for minor spelling variations.
* `semantic`: Matches the meaning of your query, even when the spoken words differ.

**Exact word matching (`lexical`)**

* Matches the specific words or phrases in your query
* Allows for minor spelling variations

**Best for**: Product names, technical terminology, proper nouns.

**Meaning-based matching (`semantic`)**

* Matches the meaning of your query, even with different wording
* Finds conceptually similar content

**Best for**: General concepts, topics that can be expressed in multiple ways.

**Using both methods (default)**

* Specify both `lexical` and `semantic`, or omit `transcription_options` entirely
* Returns the broadest set of results

**Best for**: Comprehensive searches where you want both exact matches and related content.

### Combine multiple modalities

You can search across multiple modalities simultaneously by specifying multiple values for the `search_options` parameter. Control how results are combined using the [`operator`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.operator.operator) parameter.

| `search_options`                   | `operator` | `transcription_options` | Result                                   |
| ---------------------------------- | ---------- | ----------------------- | ---------------------------------------- |
| `["visual", "transcription"]`      | `or`       | `lexical`               | Product shown OR exact name spoken       |
| `visual`, `transcription`          | `and`      | `lexical`               | Product shown WHILE exact name spoken    |
| `visual`, `transcription`          | `or`       | `semantic`              | Product shown OR discussed (any wording) |
| `visual`, `transcription`          | `and`      | `semantic`              | Product shown WHILE discussed            |
| `visual`, `audio`                  | `or`       | N/A                     | Visuals OR sounds (non-speech)           |
| `visual`, `audio`                  | `and`      | N/A                     | Visuals WITH sounds together             |
| `visual`, `audio`, `transcription` | `or`       | Both                    | Any modality matches                     |
| `visual`, `audio`, `transcription` | `and`      | Both                    | All modalities match simultaneously      |

### Related topics

* [Python SDK Reference > Make a search request](/v1.3/sdk-reference/python/search#make-a-search-request)
* [Node.js SDK Reference > Make a search request](/v1.3/sdk-reference/node-js/search#make-a-search-request)
* [API Reference > Make a search request](/v1.3/api-reference/any-to-video-search/make-search-request)

## Agents

When you search a knowledge store, use the `search_options.video.modalities` array in the [`POST /knowledge-stores/{knowledge_store_id}/search`](/v1.3/api-reference/knowledge-store-search/search) endpoint to specify which modalities the platform uses to find relevant matches:

* `visual`: Searches visual content.
* `audio`: Searches audio content, including speech and non-speech sounds.

Transcription is not a separate search modality in Agents; the `audio` modality already includes speech.

### Related topics

* [Guides > Search a knowledge store](/v1.3/agents/guides/search-a-knowledge-store)
* [Python SDK Reference > Search a knowledge store](/v1.3/sdk-reference/python/search-knowledge-store#search-a-knowledge-store)
* [Node.js SDK Reference > Search a knowledge store](/v1.3/sdk-reference/node-js/search-knowledge-store#search-a-knowledge-store)
* [API Reference > Search a knowledge store](/v1.3/api-reference/knowledge-store-search/search)

# Embedding options

When you create video embeddings, specify the types of embeddings the platform must return. You can include the following values in the [`embedding_option`](/v1.3/api-reference/create-embeddings-v1/video-embeddings/retrieve-video-embeddings#request.query.embedding_option.embedding_option) array:

* `visual`: To retrieve visual embeddings.
* `audio`: To retrieve embeddings for non-verbal audio (musical tones, beeping, environmental sounds).
* `transcription`: To retrieve embeddings for transcribed speech (the actual words spoken in the video).

## Related topics

* [Python SDK Reference > Create sync embeddings](/v1.3/sdk-reference/python/create-embeddings-v-2/create-sync-embeddings#create-sync-embeddings)
* [Python SDK Reference > Create an async embedding task](/v1.3/sdk-reference/python/create-embeddings-v-2/create-async-embeddings#create-an-async-embedding-task)
* [Node.js SDK Reference > Create sync embeddings](/v1.3/sdk-reference/node-js/create-embeddings-v-2/create-sync-embeddings)
* [Node.js SDK Reference > Create an async embedding task](/v1.3/sdk-reference/node-js/create-embeddings-v-2/create-async-embeddings#create-an-async-embedding-task)
* [API Reference > Create sync embeddings](/v1.3/api-reference/create-embeddings-v2/create-embeddings)
* [API Reference > Create an async embedding task](/v1.3/api-reference/create-embeddings-v2/create-async-embedding-task)