> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.twelvelabs.io/v1.3/docs/concepts/modalities/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server. # Modalities > Understand how the platform analyzes your video content. Modalities represent the sources of information that the platform processes and analyzes in a video. **Visual includes**: * Actions, objects, and events in the video. * Text that appears on screen (through OCR). * Brand logos and visual elements. **Audio includes**: * Ambient sounds, music, and sound effects. * Human speech (Agents). * Non-speech audio only (Models). For speech content, use the transcription modality. **Transcription includes**: * Spoken words extracted from the audio track. In Agents, speech is part of the audio modality instead. You specify modalities through different parameters depending on your task: * **Model options**: when you create an index. * **Search options**: when you search videos. * **Embedding option**: when you retrieve embeddings. # Model options When you create an index, specify which modalities the platform must process. You can include the following values in the [`model_options`](/v1.3/api-reference/indexes/create#request.body.models.model_options) array: * `visual`: To process visual content * `audio`: To process audio content You can enable one or both model options. The platform processes only the modalities you specify. ## Related topics * [Python SDK Reference > Create an index](/v1.3/sdk-reference/python/manage-indexes#create-an-index) * [Node.js SDK Reference > Create an index](/v1.3/sdk-reference/node-js/manage-indexes#create-an-index) * [API Reference > Create an index](/v1.3/api-reference/indexes/create) # Search options The modalities you can search, and the parameter you pass them in, depend on which product you use. ## Models When you search videos, use the [`search_options`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.search_options.search_options) parameter to specify which modalities the platform uses to find relevant matches. Marengo separates audio into speech and non-speech content. **To find visual content**: Set `search_options` to `visual` to search for: * Actions, objects, and events in the video * Text that appears on screen (through OCR) * Brand logos and visual elements **Example use cases:** * Finding scenes with specific objects: "red car in parking lot" * Locating on-screen text: "company logo on building" * Identifying actions: "person running" **To find non-speech audio**: Set `search_options` to `audio` to search for sounds other than human speech: * Musical tones and melodies * Beeping, alarms, and mechanical sounds * Environmental sounds (rain, traffic, nature) **Example use cases:** * Finding background music: "upbeat electronic music" * Locating sound effects: "door slamming" * Identifying ambient sounds: "rainfall" **Find spoken words** Set `search_options` to `transcription` to search the spoken content in your videos. **Example use cases:** * Finding mentions of topics: "climate change discussion" * Locating product names: "iPhone 15 Pro Max" * Identifying speakers discussing concepts: "quarterly revenue growth" ### Transcription options Use the [`transcription_options`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.transcription_options.transcription_options) parameter to specify how the platform matches your query against spoken words: * `lexical`: Matches the exact words or phrases in your query, allowing for minor spelling variations. * `semantic`: Matches the meaning of your query, even when the spoken words differ. **Exact word matching (`lexical`)** * Matches the specific words or phrases in your query * Allows for minor spelling variations **Best for**: Product names, technical terminology, proper nouns. **Meaning-based matching (`semantic`)** * Matches the meaning of your query, even with different wording * Finds conceptually similar content **Best for**: General concepts, topics that can be expressed in multiple ways. **Using both methods (default)** * Specify both `lexical` and `semantic`, or omit `transcription_options` entirely * Returns the broadest set of results **Best for**: Comprehensive searches where you want both exact matches and related content. ### Combine multiple modalities You can search across multiple modalities simultaneously by specifying multiple values for the `search_options` parameter. Control how results are combined using the [`operator`](/v1.3/api-reference/any-to-video-search/make-search-request#request.body.operator.operator) parameter. | `search_options` | `operator` | `transcription_options` | Result | | ---------------------------------- | ---------- | ----------------------- | ---------------------------------------- | | `["visual", "transcription"]` | `or` | `lexical` | Product shown OR exact name spoken | | `visual`, `transcription` | `and` | `lexical` | Product shown WHILE exact name spoken | | `visual`, `transcription` | `or` | `semantic` | Product shown OR discussed (any wording) | | `visual`, `transcription` | `and` | `semantic` | Product shown WHILE discussed | | `visual`, `audio` | `or` | N/A | Visuals OR sounds (non-speech) | | `visual`, `audio` | `and` | N/A | Visuals WITH sounds together | | `visual`, `audio`, `transcription` | `or` | Both | Any modality matches | | `visual`, `audio`, `transcription` | `and` | Both | All modalities match simultaneously | ### Related topics * [Python SDK Reference > Make a search request](/v1.3/sdk-reference/python/search#make-a-search-request) * [Node.js SDK Reference > Make a search request](/v1.3/sdk-reference/node-js/search#make-a-search-request) * [API Reference > Make a search request](/v1.3/api-reference/any-to-video-search/make-search-request) ## Agents When you search a knowledge store, use the `search_options.video.modalities` array in the [`POST /knowledge-stores/{knowledge_store_id}/search`](/v1.3/api-reference/knowledge-store-search/search) endpoint to specify which modalities the platform uses to find relevant matches: * `visual`: Searches visual content. * `audio`: Searches audio content, including speech and non-speech sounds. Transcription is not a separate search modality in Agents; the `audio` modality already includes speech. ### Related topics * [Guides > Search a knowledge store](/v1.3/agents/guides/search-a-knowledge-store) * [Python SDK Reference > Search a knowledge store](/v1.3/sdk-reference/python/search-knowledge-store#search-a-knowledge-store) * [Node.js SDK Reference > Search a knowledge store](/v1.3/sdk-reference/node-js/search-knowledge-store#search-a-knowledge-store) * [API Reference > Search a knowledge store](/v1.3/api-reference/knowledge-store-search/search) # Embedding options When you create video embeddings, specify the types of embeddings the platform must return. You can include the following values in the [`embedding_option`](/v1.3/api-reference/create-embeddings-v1/video-embeddings/retrieve-video-embeddings#request.query.embedding_option.embedding_option) array: * `visual`: To retrieve visual embeddings. * `audio`: To retrieve embeddings for non-verbal audio (musical tones, beeping, environmental sounds). * `transcription`: To retrieve embeddings for transcribed speech (the actual words spoken in the video). ## Related topics * [Python SDK Reference > Create sync embeddings](/v1.3/sdk-reference/python/create-embeddings-v-2/create-sync-embeddings#create-sync-embeddings) * [Python SDK Reference > Create an async embedding task](/v1.3/sdk-reference/python/create-embeddings-v-2/create-async-embeddings#create-an-async-embedding-task) * [Node.js SDK Reference > Create sync embeddings](/v1.3/sdk-reference/node-js/create-embeddings-v-2/create-sync-embeddings) * [Node.js SDK Reference > Create an async embedding task](/v1.3/sdk-reference/node-js/create-embeddings-v-2/create-async-embeddings#create-an-async-embedding-task) * [API Reference > Create sync embeddings](/v1.3/api-reference/create-embeddings-v2/create-embeddings) * [API Reference > Create an async embedding task](/v1.3/api-reference/create-embeddings-v2/create-async-embedding-task) > Understand how the platform analyzes your video content.