Pegasus 1.5
Pegasus 1.5 is a generative model for video-to-text generation. It analyzes multiple modalities to generate contextually relevant text based on the content of your videos.
Key features
- Video-to-text generation: Creates detailed textual descriptions directly from a URL, asset, or base64 string, with no indexing required
- Extended processing capacity: Analyzes up to 2 hours of video, or a portion of a video up to 4 hours long, with asynchronous analysis
- Granular visual comprehension: Analyzes objects, on-screen text, and numerical content
- Temporal grounding: Accurately identifies timestamps of specific events
- Multimodal understanding: Combines visual, audio, and textual information for comprehensive analysis
- Video segmentation: Transforms raw videos into structured, timestamped data with custom segment definitions and JSON results
- Multimodal prompting: Includes reference images in prompts to provide visual context or help identify specific segments
- Video clipping: Analyzes a specific portion of the video
- Per-definition time ranges: Restricts segment extraction to specific time windows within the video
- Context window: Covers both the input and the response within 261,120 tokens per request
- Longer responses: Generates responses up to 98,304 tokens
Context window
Pegasus 1.5 uses a context window of 261,120 tokens. The context window is the maximum number of tokens a single request can use. This limit covers both the input and the response.
The following inputs and outputs count toward the context window:
- Video content
- Audio transcription
- Prompt text
- Reference images (in prompts or segment definitions)
- JSON schema (if you request structured responses)
- Segment definitions (for video segmentation)
- Generated output (text or JSON)
Use cases
- Content summarization: Generate concise summaries of video content
- Detailed descriptions: Create comprehensive textual descriptions of visual scenes
- Timestamp identification: Answer questions about when specific events occur in videos
- Content analysis: Extract key information from video content for further processing
- Image-guided analysis: Reference images in your prompt to ask about specific objects, people, or scenes in the video
- Video segmentation: Detect and extract structured metadata for editorial segments, scene changes, sports plays, or brand appearances. For details, see the Segment videos page.
Input requirements
The specifications on this page reflect the maximum capabilities of the model. Your actual requirements depend on the upload method and operation you choose. For details about the available upload methods and the corresponding limits, see the Upload and processing methods page.
Video file requirements
- Duration: The video can be up to 2 hours long, or up to 4 hours when you analyze only a portion of it. You can analyze between 1 second and 2 hours of it.
- File size: ≤ 10 GB
- Resolution: 360x360 to 5184x2160
- Aspect ratio: Between 1:1 and 1:2.4, or between 2.4:1 and 1:1. For example, you can use 1:1, 4:3, 4:5, 5:4, 16:9, 9:16, or 17:9.
- Formats: FFmpeg supported
-
If you upload files using publicly accessible URLs, use direct links to raw videos that play without user interaction or custom video players (example:
https://example.com/videos/sample-video.mp4). Video hosting platforms and cloud storage sharing links are not supported. -
For videos in other formats or if you require different options, contact us at support@twelvelabs.io.
Supported languages
Pegasus 1.5 supports the following languages for processing visual and audio content, understanding prompts, and generating outputs:
- Full support: English
- Partial support: Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, Vietnamese
Support
For support or feedback regarding Pegasus, contact support@twelvelabs.io.