> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Analyze videos and images

> Analyze videos and images to generate text based on their content, including quality control of AI-generated video, such as checking a clip against its prompt and detecting visual defects.

The platform uses a multimodal approach to analyze videos and images and generate text, processing visuals, sounds, spoken words, and on-screen text to provide a comprehensive understanding. This method captures nuances that unimodal interpretations might miss, enabling accurate, context-rich text generation based on your content.

**Key features**:

* **Multimodal analysis**: Processes visuals, sounds, spoken words, and text for a holistic understanding of your content.
* **Egocentric video understanding**: Analyzes first-person footage from sources such as wearable cameras and teleoperated systems.
* **Image analysis**: Analyzes images with the same prompt-based workflow as videos.
* **Entity recognition**: Identifies and names people, objects, and characters in a scene.
* **Customizable prompts**: Allows tailored outputs through instructive, descriptive, or question-based prompts.
* **Flexible text generation**: Supports various tasks, including summarization, chaptering, and open-ended text generation.
* **Segment videos**: Extract structured, timestamped segments from your videos by defining custom segment types and fields, including lists of timestamped events inside each segment.
* **Segmentation without a token limit**: Segments long videos and returns the segments it produced, even when the output is incomplete.

**Use cases**:

* **Robotics and physical AI**: Label actions, check quality, and predict teleoperated trajectories and next actions in egocentric footage.
* **Content structuring**: Organize and structure content for e-learning platforms to improve usability.
* **Product analysis**: Compare product images or extract product details and visible features.
* **SEO optimization**: Optimize content to rank higher in search engine results.
* **Highlight creation**: Create short, engaging video clips for media and broadcasting.
* **Incident reporting**: Record and report incidents for security and law enforcement purposes.
* **Generative video QA**: Check AI-generated clips against their prompt or flag visual defects before you use them.

For details on how your usage is measured and billed, see the [Pricing](https://www.twelvelabs.io/pricing) page.

# Workflow

The platform provides two methods to analyze content. Choose the method that fits your use case:

#### [Analyze videos](/v1.3/docs/guides/analyze-videos-and-images/videos)

Upload a video as an asset or pass it inline, then analyze it synchronously or asynchronously.

#### [Analyze images](/v1.3/docs/guides/analyze-videos-and-images/images)

Upload your images as assets or pass them inline, and analyze them in a single request.

**Customize text generation**

You can configure the temperature to control output randomness, set the maximum response length, and request [structured JSON responses](/v1.3/docs/guides/analyze-videos-and-images/structured-responses) for programmatic processing. To extract timestamped segments with custom fields, see the [Segment videos](/v1.3/docs/guides/segment-videos) page.