Pegasus 1.6
Pegasus 1.6 is a generative model. It analyzes videos or images to generate contextually relevant text based on their content.
New in Pegasus 1.6
- Egocentric video understanding: Analyzes first-person footage from sources such as wearable cameras and teleoperated systems
- Image analysis: Analyzes one to twenty images per request
- Improved entity recognition: Identifies and names people, objects, and characters in a scene more consistently than Pegasus 1.5
- Improved metadata extraction: Extracts metadata that is less dependent on the segment definitions, enabling richer fields in a single analysis
- In-segment events: Extracts a list of timestamped events inside each segment when you use video segmentation
- Segmentation without a token limit: Segments long videos and returns the segments it produced, even when the output is incomplete
Context window
Pegasus 1.6 uses a context window of 261,120 tokens. The context window is the maximum number of tokens a single request can use. This limit covers both the input and the response.
The following inputs and outputs count toward the context window:
- Video content
- Audio transcription
- Prompt text
- Reference images (in prompts or segment definitions)
- JSON schema (if you request structured responses)
- Segment definitions (for video segmentation)
- Generated output (text or JSON)
Use cases
Pegasus 1.6 adds the following use cases on top of those available with Pegasus 1.5:
- Action labeling: Label actions in first-person footage from wearable and teleoperated cameras
- Quality control: Review robotics and physical-AI footage against expected behavior
- Trajectory and next-action prediction: Predict the operator’s next action in teleoperated video
- Image analysis: Analyze and compare product shots, scenes, and diagrams
- Video segmentation: Extract timestamped segment metadata, including in-segment events
Input requirements
The specifications on this page reflect the maximum capabilities of the model. Your actual requirements depend on the upload method and operation you choose. For details about the available upload methods and the corresponding limits, see the Upload and processing methods page.
Video file requirements
- Duration: The video can be up to 2 hours long, or up to 4 hours when you analyze only a portion of it. You can analyze between 1 second and 2 hours of it.
- File size: ≤ 10 GB
- Resolution: 360x360 to 5184x2160
- Aspect ratio: Between 1:1 and 1:2.4, or between 2.4:1 and 1:1. For example, you can use 1:1, 4:3, 4:5, 5:4, 16:9, 9:16, or 17:9.
- Formats: FFmpeg supported
-
If you upload files using publicly accessible URLs, use direct links to raw videos that play without user interaction or custom video players (example:
https://example.com/videos/sample-video.mp4). Video hosting platforms and cloud storage sharing links are not supported. -
For videos in other formats or if you require different options, contact us at support@twelvelabs.io.
Image file requirements
- Number of images: one to twenty per request
- Formats: JPEG, PNG, WebP, GIF, and BMP
- File size: ≤ 20 MB per image
- Pixel count: ≤ 16,777,216 pixels per image (width × height)
Supported languages
Pegasus 1.6 supports the following languages for processing visual and audio content, understanding prompts, and generating outputs:
- Full support: English
- Partial support: Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, Vietnamese
Support
For support or feedback regarding Pegasus, contact support@twelvelabs.io.