> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Models

> Use Marengo for embeddings, and Pegasus for text generation.

TwelveLabs' video understanding models consist of a family of deep neural networks built on our multimodal foundation model for video understanding that you can use for the following downstream tasks:

* Search using natural language queries
* Analyze videos to generate text

Videos contain multiple types of information, including visuals, sounds, spoken words, and text. The human brain combines all types of information and their relations with each other to comprehend the overall meaning of a scene. For example, you're watching a video of a person jumping and clapping, both visual cues, but the sound is muted. You might realize they're happy, but you can't understand why they're happy without the sound. However, if the sound is unmuted, you could realize they're cheering for a soccer team that scored a goal.

Thus, an application that analyzes a single type of information can't provide a comprehensive understanding of a video. TwelveLabs' video understanding models, however, analyze and combine information from all the modalities to accurately interpret the meaning of a video holistically, similar to how humans watch, listen, and read simultaneously to understand videos.

These models can identify, analyze, and interpret a variety of elements, including but not limited to the following:

| Element                              | Modality | Example                                                            |
| :----------------------------------- | :------- | :----------------------------------------------------------------- |
| People, including famous individuals | Visual   | Michael Jordan, Steve Jobs                                         |
| Actions                              | Visual   | Running, dancing, kickboxing                                       |
| Objects                              | Visual   | Cars, computers, stadiums                                          |
| Animals or pets                      | Visual   | Monkeys, cats, horses                                              |
| Nature                               | Visual   | Mountains, lakes, forests                                          |
| Text displayed on the screen (OCR)   | Visual   | License plates, handwritten words, number on a player's jersey     |
| Brand logos                          | Visual   | Nike, Starbucks, Mercedes                                          |
| Shot techniques and effects          | Visual   | Aerial shots, slow motion, time-lapse                              |
| Counting objects                     | Visual   | Number of people in a crowd, items on a shelf, vehicles in traffic |
| Sounds                               | Audio    | Chirping (birds), applause, fireworks popping or exploding         |
| Human speech                         | Audio    | "Good morning. How may I help you?"                                |
| Music                                | Audio    | Ominous music, whistling, lyrics                                   |

# Available models

TwelveLabs provides different models for video understanding tasks. This section describes each model and its capabilities, helping you understand which one fits your needs.

## Marengo

**Task type**

* Search for specific content in your videos using natural language queries.
* Create video embeddings for downstream tasks.

**Use cases**

* Find scenes where a person appears, locate brand logos, search for spoken phrases, identify specific actions or objects.
* Build recommendation systems, perform similarity searches, integrate with custom ML pipelines.

#### [Marengo](/v1.3/docs/concepts/models/marengo)

## Pegasus

**Task type**

* Analyze videos and generate text based on their content.

**Use cases**

* Create video summaries, generate social media captions, extract key information, identify when events occur in videos.

#### [Pegasus](/v1.3/docs/concepts/models/pegasus)