Marengo 3.0
Marengo 3.0 is an embedding model for comprehensive video understanding. It analyzes visual, audio, and text information to provide a holistic understanding similar to human comprehension.
Marengo 3.0 is available for search, indexing, and creating embeddings.
Key features
- Multimodal processing: Combines visual, audio, and text elements for comprehensive understanding.
- Fine-grained search: Detects brand logos, text, and small objects (as small as 10% of the video frame).
- Motion search: Identifies and analyzes movement within videos.
- Counting capabilities: Accurately counts objects in video frames.
- Audio comprehension: Analyzes music, lyrics, sound, and silence.
- Expanded language support: Query videos in 36 languages.
- Composed text and image search: Combine text descriptions with images in a single search query for more precise results.
- Improved cinematography understanding: Enhanced search performance for cinematography terms like zoom, pan, and tracking shot.
- Sports intelligence: Improved recognition of soccer and basketball actions. Support for baseball, ice hockey, and American football.
- Faster indexing: Significant performance improvement with the new indexing technology.
- Extended text processing: Maximum text length increased from 77 to 500 tokens for both search queries and text embeddings.
- Optimized embeddings: 512-dimensional embeddings for faster processing and reduced storage.
- Long content support: Process up to four hours of video and audio content while maintaining context.
Context window
Marengo 3.0 uses a context window of 500 tokens. The context window is the maximum number of tokens a single request can use. This limit applies to search queries and text embeddings.
Use cases
- Search: Use text, images, video clips, or audio to find specific content. The model supports any-to-any search across multiple modalities.
- Embeddings: Create video embeddings for various downstream applications.
Input requirements
The specifications on this page reflect the maximum capabilities of the model. Your actual requirements depend on the upload method and operation you choose. For details about the available upload methods and the corresponding limits, see the Upload and processing methods page.
Video file requirements
- Duration: 4 sec to 4 hours
- File size: ≤ 4 GB
- Resolution: 360x360 to 5184x2160
- Aspect ratio: Between 1:1 and 1:2.4, or between 2.4:1 and 1:1. For example, you can use 1:1, 4:3, 4:5, 5:4, 16:9, 9:16, or 17:9.
- Formats: FFmpeg supported
Notes
-
If you upload files using publicly accessible URLs, use direct links to raw video files that play without user interaction or custom video players (example:
https://example.com/videos/sample-video.mp4). Video hosting platforms and cloud storage sharing links are not supported. -
For videos in other formats or if you require different options, contact us at support@twelvelabs.io.
Image file requirements
- Formats: JPEG, PNG
- Minimum size: 128x128 pixels
- File size: ≤ 32 MB
Audio file requirements
- Formats: WAV (uncompressed), MP3 (lossy), and FLAC (lossless)
- Duration: Up to 4 hours
- File size: ≤ 4 GB
Supported languages
Arabic, Bengali, Chinese (Simplified), Croatian, Cusco, Czech, Danish, Dutch, English, Farsi, Filipino, Finnish, French, German, Greek, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Maori, Norwegian, Polish, Portuguese, Romanian, Russian, Spanish, Swahili, Swedish, Telugu, Thai, Turkish, Ukrainian, and Vietnamese.
Support
For support or feedback regarding Marengo, contact support@twelvelabs.io.