> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Pegasus

> Pegasus is a generative model. It analyzes multiple modalities to generate contextually relevant text.

Pegasus is a generative model for video-to-text generation. Pegasus analyzes multiple modalities to generate contextually relevant text based on the content of your videos.

The current version is **Pegasus 1.5**.

# New in Pegasus 1.5

Pegasus 1.5 analyzes videos directly from a URL, asset, or base64 string, with no pre-indexing required. Compared to Pegasus 1.2, it introduces the following capabilities:

* **Video segmentation**: Transform raw videos into structured, timestamped data. Define the types of segments you want to detect, such as editorial narratives, sports plays, speaker changes, or brand appearances, specify custom fields for each segment, and receive structured results in JSON format. For details, see the [Segment videos](/v1.3/docs/guides/segment-videos) page.
* **Multimodal prompting**: Include reference images to provide visual context for your prompts or to help identify specific segments during video segmentation.
* **Video clipping**: Analyze a specific portion of the video.
* **Per-definition time ranges**: Restrict segment extraction to specific time windows within the video for finer control over video segmentation.
* **Context window**: Pegasus 1.5 uses a shared [context window](#context-window) of 261,120 tokens for input and output per request.
* **Longer responses**: Pegasus 1.5 supports responses up to 98,304 tokens.

# Key features

* **Video-to-text generation**: Creates detailed textual descriptions based on video content
* **Extended processing capacity**: Analyzes up to 2 hours of video, or a portion of a video up to 4 hours long, with asynchronous analysis
* **Granular visual comprehension**: Analyzes objects, on-screen text, and numerical content
* **Temporal grounding**: Accurately identifies timestamps of specific events
* **Multimodal understanding**: Combines visual, audio, and textual information for comprehensive analysis

# Context window

Pegasus 1.5 uses a context window of 261,120 tokens. The context window is the maximum number of tokens a single request can use. This limit covers both the input and the response.

The following inputs and outputs count toward the context window:

* Video content
* Audio transcription
* Prompt text
* Reference images (in prompts or segment definitions)
* JSON schema (if you request structured responses)
* Segment definitions (for video segmentation)
* Generated output (text or JSON)

# Use cases

* **Content summarization**: Generate concise summaries of video content
* **Detailed descriptions**: Create comprehensive textual descriptions of visual scenes
* **Timestamp identification**: Answer questions about when specific events occur in videos
* **Content analysis**: Extract key information from video content for further processing
* **Image-guided analysis**: Reference images in your prompt to ask about specific objects, people, or scenes in the video
* **Video segmentation**: Detect and extract structured metadata for editorial segments, scene changes, sports plays, or brand appearances

# Input requirements

The specifications on this page reflect the maximum capabilities of the model. Your actual requirements depend on the upload method and operation you choose. For details about the available upload methods and the corresponding limits, see the [Upload and processing methods](/v1.3/docs/concepts/upload-methods) page.

## Video file requirements

* **Duration**: The video can be up to 2 hours long, or up to 4 hours when you analyze only a portion of it. You can analyze between 1 second and 2 hours of it.
* **File size**: ≤ 10 GB
* **Resolution**: 360x360 to 5184x2160
* **Aspect ratio**: Between 1:1 and 1:2.4, or between 2.4:1 and 1:1. For example, you can use 1:1, 4:3, 4:5, 5:4, 16:9, 9:16, or 17:9.
* **Formats**: [FFmpeg supported](https://ffmpeg.org/ffmpeg-formats.html)

> **Notes**
>
> * If you upload files using publicly accessible URLs, use direct links to raw videos that play without user interaction or custom video players (example: `https://example.com/videos/sample-video.mp4`). Video hosting platforms and cloud storage sharing links are not supported.
>
> * For videos in other formats or if you require different options, contact us at [support@twelvelabs.io](mailto:support@twelvelabs.io).

# Supported languages

Pegasus supports the following languages for processing visual and audio content, understanding prompts, and generating outputs:

* **Full support**: English
* **Partial support**: Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, Vietnamese

# Examples

## Summarizing educational videos

This example prompt summarizes an educational video without any customization.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Summarize this video
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
This video presents an educational overview of advanced techniques for adapting pre-trained large language models (LLMs) to specialized tasks, with a focus on prompt tuning as a more efficient alternative to traditional fine-tuning and manual prompt engineering. The speaker, Martin Kevner, Master Inventor at IBM, begins by explaining that foundation models like ChatGPT are highly flexible but may need customization for specific applications. He contrasts three approaches: fine-tuning, which requires large amounts of labeled data and retraining; prompt engineering, where humans craft explicit instructions (hard prompts) to guide model behavior; and prompt tuning, which uses AI-generated, tunable soft prompts—numerical embeddings that are invisible to humans but highly effective. The video highlights that soft prompts outperform human-designed ones, enable faster adaptation, and are particularly valuable in multi-task and continual learning scenarios, where models must switch tasks or learn new concepts without forgetting prior knowledge. Despite the efficiency and power of prompt tuning, the speaker notes its lack of interpretability, as the AI cannot explain why certain embeddings were chosen. The video concludes with a reflective tone, acknowledging that while prompt engineering may become obsolete, the future lies in leveraging the underlying embedding layers, where "a string of numbers is worth a thousand words." The presentation is visually enhanced with dynamic digital annotations, including animated text, virtual writing, and diagrams illustrating model architectures and data flow, reinforcing the technical concepts through interactive visualization.
```

## Generating captions for social media

This example prompt generates a caption for a social media post.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Generate an attention-grabbing caption for a social media post. Keep it shorter than 200 characters.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
Master the basics of golf in just 4 weeks! From alignment to grip, swing mechanics to mindset—learn how to play like a pro. Hit 30 balls this week, stay consistent, and have fun. #HowToPlayGolf #GolfTips
```

## Table of contents

This example prompt creates a table of contents detailing the main sections.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Provide a table of contents detailing the main sections of this video.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
- Planning  
- Execution  
- Review & Improvement  
- Checklist: Enthusiast, Brand Message, Cheese  
- Pre-Armed Model (Pre-trained Model)  
- Model Architecture: Input, Player, Target, Output, Softmax  
- Feedback Loop and Data Flow  
- Prompt Tuning vs. Fine Tuning vs. Prompt Engineering  
- Applications in Multi-task and Continual Learning  
- Conclusion: The Future of Prompt Engineering
```

## Company-wide memo

This example prompt generates a company-wide memo.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Generate a company-wide memo based on the content of this video.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
Subject: Enhancing Our Understanding of Machine Learning and Its Applications

Dear Team,

I hope this message finds you well. I’m sharing a concise yet comprehensive overview of a recent video that explores the fundamentals and real-world applications of machine learning—a field increasingly central to innovation across industries.

Machine learning enables computers to learn from data and improve over time without being explicitly programmed for each task. First introduced in 1959 by Arthur Samuel at IBM, who developed a checkers-playing program that learned from experience, the concept has evolved dramatically. Today, predictive models powered by machine learning are embedded in everyday technologies—from self-driving cars and smart thermostats to personalized content recommendations and medical diagnostics.

The core process begins with acquiring and cleaning large volumes of high-quality data, as the principle “garbage in, garbage out” underscores the importance of data integrity. Data scientists then perform feature engineering to transform raw data into meaningful inputs that better represent the underlying problem. The dataset is split into training and testing sets: the model learns from the training data, and its performance is validated using the test data.

A wide range of algorithms can be used, from simple models like linear regression to complex ones such as convolutional neural networks (CNNs), which automatically generate features—especially valuable for image and natural language processing. These models improve by minimizing error through feedback mechanisms, using metrics like accuracy for classification or mean absolute error for regression.

Python remains the dominant language in the field, supported by powerful frameworks like TensorFlow, PyTorch, and Scikit-learn, which streamline development and deployment. The final output is a trained model—a compact file that takes input data and produces predictions, ready to be deployed on devices or in the cloud to power real-world applications.

This video not only demystifies the technical process but also highlights the transformative impact of machine learning across sectors. I encourage everyone to watch it to deepen their understanding of how data-driven intelligence is shaping the future of technology.

Best regards,  
[Your Name]  
[Your Position]
```

## Video annotations

This example prompt identifies and lists key visual elements, scene changes, and notable events, briefly describing each.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Identify and list key visual elements, scene changes, and notable events in the video, briefly describing each.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
The video opens with the Universal Sports logo (00:00–00:03), setting the stage for a high-profile athletics broadcast. It quickly transitions to a scoreboard displaying the results of the Men’s 100-meter final at the IAAF World Championships, highlighting Usain Bolt in first place (00:03–00:11). The scene shifts to the starting line, where the sprinters are lined up, and the camera pans across the athletes, emphasizing the elite field of runners who have all broken the 10-second barrier (00:11–00:16).  

At the start of the race (00:16), the runners explode from the blocks, with Daniel Bailey making a strong initial move, but Usain Bolt quickly takes control of the race, accelerating into his stride (00:16–00:26). The camera follows the race from a side-on perspective, capturing Bolt’s dominant lead as he pulls away from the field.  

As Bolt crosses the finish line (00:26–00:32), the camera shifts to a higher vantage point, emphasizing his commanding margin of victory—two full meters ahead of Tyson Gay, who finishes in second (00:32–00:45). The video then replays the race from a different angle, focusing on Bolt’s consistent speed and the growing gap between him and the rest of the field (00:45–01:17).  

A close-up shot captures Bolt’s reaction upon realizing his performance, showing a mix of disbelief and elation (01:17–01:28). The camera then switches to a top-down view, visually reinforcing Bolt’s lead throughout the race (01:28–01:42).  

The climax comes at 01:42, when the electronic timer displays the new world record: 9.58 seconds. The screen flashes the iconic image of Bolt’s victory, accompanied by the announcement of the record-breaking time (01:49–02:01). The moment is underscored by dramatic commentary, noting the significance of the performance and the slight tailwind of +0.9 m/s that aided the time but remained within the allowable limit (01:38–01:44).  

The video concludes with a final shot of Bolt celebrating, followed by the Universal Sports logo reappearing (02:03–02:09), marking the end of the broadcast. The entire sequence captures not only the athletic brilliance of Usain Bolt but also the global significance of his record-breaking run.
```

## Video question answering

This example prompt identifies the key takeaways.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
What are the key takeaways of this video?
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
The video explains the transformative power of transformer models in artificial intelligence, emphasizing their ability to process and generate human-like text and perform diverse tasks beyond language translation. Key takeaways include: transformers are sequence-to-sequence models composed of an encoder and decoder, which work together to understand context and generate outputs; they leverage an attention mechanism that allows parallel processing of input sequences, significantly speeding up training compared to older models like RNNs; they are pre-trained on large unlabeled datasets and fine-tuned with supervised learning, enabling versatility across applications such as joke generation, document summarization, blog writing, and even image processing and game playing; the video concludes by highlighting that while current outputs—like a banana joke—may not always be funny, the underlying technology is rapidly advancing and holds immense potential for future applications.
```

## Timestamp breakdown

This example prompt lists all timestamps in an advertisement where a close-up of the product appears.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Tell me all the timestamps in the advertisement where a closeup of the product appears.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
The advertisement features a close-up of the product, Gatorade, at the following timestamps:

- [00:55] – A young man is shown drinking directly from a Gatorade bottle, captured in a clear close-up shot that highlights the product.
```

## Police report

This example prompt creates a police report using a specific template for a video showing a robbery.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Create a police report based on what happened in the video. Provide the exact time range where the suspect appears in the video
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
Police Report

Incident Description:  
A suspicious individual was observed approaching a package located on a wooden deck, picking it up, and departing through a white gate. The incident was captured on video from a fixed camera angle, providing a clear visual of the suspect’s actions and appearance.

Suspect Description:  
- Gender: Male  
- Clothing: Blue hoodie, grey pants  
- Additional Features: None observed  

Time Range of Suspect’s Appearance:  
00:02 to 00:10  

Video Context:  
The video footage begins at 00:00 with the package resting on the deck. At 00:02, the suspect enters the frame, walks toward the package, and takes it. The individual is seen moving steadily and deliberately, exiting the scene through the white gate by 00:10. The audio clip from 00:07 to 00:10 includes the phrase “so you,” which may be part of a conversation unrelated to the incident or possibly a fragment of the suspect’s speech, though no further context is available.

Additional Information:  
The video provides a continuous and unobstructed view of the suspect’s movements. No other individuals are visible during the incident. The package was unattended and left on the deck prior to the suspect’s arrival.

Please contact the police department if you have any further information or if you are the owner of the stolen package.
```

## Using different languages

### Spanish

This example prompt summarizes a video and indicates that the response should be in Spanish. Note that the prompt is in English, and the output is in Spanish.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Write a summary in Spanish.
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
Este video presenta a Tim Shields, fotógrafo profesional, quien comparte una guía práctica para tomar retratos de alta calidad. Comienza con una introducción sobre la importancia de los retratos, el escenario más común en fotografía, y explica cómo aplicar la regla de los tercios para una composición equilibrada: colocar al sujeto en uno de los tercios horizontales o verticales, evitando que mire hacia fuera del cuadro para no desviar la atención del espectador. Destaca que el rostro debe situarse en la línea superior de los tercios, preferiblemente en una intersección, y que elementos interesantes del fondo deben alinearse con los otros tercios. También advierte contra distracciones detrás de la cabeza del sujeto, como postes o ramas.

A continuación, revela el "ingrediente secreto" para fotos profesionales: desenfocar el fondo para resaltar al sujeto. Explica cómo lograrlo con cámaras inteligentes (usando el modo retrato) o con cámaras DSLR ajustando el diafragma al valor más bajo (f-stop más pequeño) y usando el zoom al máximo. También aborda la iluminación, recomendando evitar que el sujeto esté de espaldas al sol y proponiendo soluciones como acercarse, usar flash o posicionar al sujeto con luz natural o artificial (clave y relleno) para lograr un efecto suave y profesional.

Tim introduce su libro, *The Photo Cookbook*, que contiene 30 recetas fotográficas para dominar técnicas como la profundidad de campo, la composición, el equilibrio, las líneas guía, el enfoque y el uso de la luz. Destaca que no se trata de tener una cámara cara, sino de seguir pasos claros y estructurados. El video concluye con una invitación a adquirir el libro para mejorar las fotos, ganar reconocimiento en redes sociales, participar en concursos y tomar imágenes de las que sentirse orgulloso.
```

### French

This example prompt summarizes the three main takeaways of a video. Note that the prompt and the output are in French.

**`Prompt`**

```markdown title="Prompt" showLineNumbers={false} wordWrap maxLines=5
Résumez les trois principaux points à retenir de cette vidéo
```

**`Output`**

```markdown title="Output" showLineNumbers={false} wordWrap maxLines=5
Cette vidéo explique les stratégies de réplication de bases de données essentielles pour les entretiens de conception système. Les trois principaux points à retenir sont :  
1. La réplication est cruciale pour la résilience, la scalabilité et la réduction de la latence, notamment dans les systèmes distribués où les réseaux sont instables.  
2. La réplication leader-follower, la plus courante, repose sur un nœud principal (leader) qui gère les écritures et les réplique vers des nœuds secondaires (followers), avec deux variantes : synchrones (haute cohérence, mais plus lente) et asynchrones (plus rapide, mais risque d’incohérence en cas de panne du leader).  
3. Des alternatives comme la réplication multi-leader (améliore la disponibilité) et la réplication sans leader (comme dans DynamoDB, avec gestion des conflits par des mécanismes comme read-repair) offrent des compromis entre performance, cohérence et résilience, et doivent être choisies selon les besoins spécifiques du système.
```

# Support

For support or feedback regarding Pegasus, contact [support@twelvelabs.io](mailto:support@twelvelabs.io).