> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Marengo

> Marengo is an embedding model. It analyzes visual, audio, and text information.

Marengo is an embedding model for comprehensive video understanding. It analyzes visual, audio, and text information to provide a holistic understanding similar to human comprehension.

# Available versions

#### [Marengo 3.5](/v1.3/docs/concepts/models/marengo/marengo-3-5)

Create embeddings from video, audio, images, text, and documents.

#### [Marengo 3.0](/v1.3/docs/concepts/models/marengo/marengo-3-0)

Search your content with the platform, and create embeddings from video, audio, images, and text.

# Examples

The examples in this section are from the [Playground](https://playground.twelvelabs.io). However, the principles demonstrated are similar when invoking the API programmatically.

## Steve Jobs introducing the iPhone

In the example screenshot below, the query was "How did Steve Jobs introduce the iPhone?".  The Marengo video understanding model used information found in the visual and conversation audio to perform the following tasks:

* Visual recognition of a famous person (Steve Jobs)
* Joint speech and visual recognition to semantically search for the moment when Steve Jobs introduced the iPhone. Note that semantic search finds information based on the intended meaning of the query rather than the literal words you used, meaning that the platform identified the matching video fragments even if Steve Jobs didn't explicitly say the words in the query.

![](/_fern-img/c18654c85c03cbe61b796eca685042cdb565b5c96cd1ab72348de69e293fda14.webp)

## Polar bear holding a Coca-Cola bottle

In the example screenshot below, the query was "Polar bear holding a Coca-Cola bottle." The Marengo video understanding model used information found in the visual and logo modalities to perform the following tasks:

* Recognition of a cartoon character (polar bear)
* Identification of an object (bottle)
* Detection of a specific brand logo (Coca-Cola)
* Identification of an action (polar bear holding a bottle)

![](/_fern-img/24b43c456ac39ff582e3aae21167cfdac9bd1e751a9df681812821994a1b3605.webp)

To see this example in the Playground, ensure you're logged in, and then open [this URL](https://playground.twelvelabs.io/indexes/691bc2e7d4fd5772d3a41e02/search?qt=Polar+bear+holding+a+Coca-Cola+bottle.) in your browser.

## Using different languages

This section provides examples of using different languages to perform search requests.

### Spanish

In the example screenshot below, the query was "¿Cómo presentó Steve Jobs el iPhone?" ("How did Steve Jobs introduce the iPhone?"). The Marengo video understanding model used information from the visual and audio modalities.

![](/_fern-img/d8df0adc221ba986d05536059887bc7c1d72ef9a168be754d0f08216a927ff13.webp)

### Chinese

In the example screenshot below, the query was "猫做有趣的事情" ("Cats doing funny things."). The Marengo video understanding model used information from the visual modality.

![](/_fern-img/be100105b4db77b186cd9cdcc9a328ebd6b87c77b8d0e15182d19a1600df4856.webp)

To see this example in the Playground, ensure you're logged in, and then open [this URL](https://playground.twelvelabs.io/indexes/691bc2e7d4fd5772d3a41e02/search?qt=猫做有趣的事情) in your browser.

# Support

For support or feedback regarding Marengo, contact [support@twelvelabs.io](mailto:support@twelvelabs.io).