> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.twelvelabs.io/v1.3/docs/guides/analyze-videos-and-images/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server. # Analyze videos and images > Analyze videos and images to generate text based on their content, including quality control of AI-generated video, such as checking a clip against its prompt and detecting visual defects. The platform uses a multimodal approach to analyze videos and images and generate text, processing visuals, sounds, spoken words, and on-screen text to provide a comprehensive understanding. This method captures nuances that unimodal interpretations might miss, enabling accurate, context-rich text generation based on your content. **Key features**: * **Multimodal analysis**: Processes visuals, sounds, spoken words, and text for a holistic understanding of your content. * **Egocentric video understanding**: Analyzes first-person footage from sources such as wearable cameras and teleoperated systems. * **Image analysis**: Analyzes images with the same prompt-based workflow as videos. * **Entity recognition**: Identifies and names people, objects, and characters in a scene. * **Customizable prompts**: Allows tailored outputs through instructive, descriptive, or question-based prompts. * **Flexible text generation**: Supports various tasks, including summarization, chaptering, and open-ended text generation. * **Segment videos**: Extract structured, timestamped segments from your videos by defining custom segment types and fields, including lists of timestamped events inside each segment. * **Segmentation without a token limit**: Segments long videos and returns the segments it produced, even when the output is incomplete. **Use cases**: * **Robotics and physical AI**: Label actions, check quality, and predict teleoperated trajectories and next actions in egocentric footage. * **Content structuring**: Organize and structure content for e-learning platforms to improve usability. * **Product analysis**: Compare product images or extract product details and visible features. * **SEO optimization**: Optimize content to rank higher in search engine results. * **Highlight creation**: Create short, engaging video clips for media and broadcasting. * **Incident reporting**: Record and report incidents for security and law enforcement purposes. * **Generative video QA**: Check AI-generated clips against their prompt or flag visual defects before you use them. For details on how your usage is measured and billed, see the [Pricing](https://www.twelvelabs.io/pricing) page. # Workflow The platform provides two methods to analyze content. Choose the method that fits your use case: #### [Analyze videos](/v1.3/docs/guides/analyze-videos-and-images/videos) Upload a video as an asset or pass it inline, then analyze it synchronously or asynchronously. #### [Analyze images](/v1.3/docs/guides/analyze-videos-and-images/images) Upload your images as assets or pass them inline, and analyze them in a single request. **Customize text generation** You can configure the temperature to control output randomness, set the maximum response length, and request [structured JSON responses](/v1.3/docs/guides/analyze-videos-and-images/structured-responses) for programmatic processing. To extract timestamped segments with custom fields, see the [Segment videos](/v1.3/docs/guides/segment-videos) page. > Analyze videos and images to generate text based on their content, including quality control of AI-generated video, such as checking a clip against its prompt and detecting visual defects. ## Docs - [Analyze videos](https://docs.twelvelabs.io/docs/guides/analyze-videos-and-images/videos.md): Analyze videos to generate text based on their content, including quality control of AI-generated video, such as checking a clip against its prompt and detecting visual defects. - [Analyze images](https://docs.twelvelabs.io/docs/guides/analyze-videos-and-images/images.md): Analyze images to generate text based on their content. - [Structured responses](https://docs.twelvelabs.io/docs/guides/analyze-videos-and-images/structured-responses.md): Retrieve JSON responses from video analysis. Define schemas, stream results, and integrate into your applications. - [Tune the temperature](https://docs.twelvelabs.io/docs/guides/analyze-videos-and-images/tune-the-temperature.md): Control the randomness of the output. Balance creativity and determinism. - [Prompt engineering](https://docs.twelvelabs.io/docs/guides/analyze-videos-and-images/prompt-engineering.md): Craft effective prompts for video analysis. Learn techniques to improve output quality, relevance, and precision.