Sync analysis
This method analyzes a video or one or more images and returns the results directly in the response. Each request must contain a video or one or more images, but not both. You can use general analysis (prompt-based text generation) with either media type.
Input requirements
Videos
- Minimum duration: 1 second
- Maximum duration: 1 hour
- Formats: FFmpeg supported formats
- Resolution: 360x360 to 5184x2160 pixels
- Aspect ratio: Between 1:1 and 1:2.4, or between 2.4:1 and 1:1.
Images
- You can provide one to twenty images per request.
- Formats: JPEG, PNG, WebP, GIF, and BMP.
- Maximum size: 20 MB per image.
- Maximum pixel count: 16,777,216 pixels per image (width × height).
When to use this method:
- Analyze videos up to 1 hour, or analyze images
- Retrieve immediate results without polling for task completion
- Stream text fragments in real time for immediate processing and feedback
Do not use this method for:
- Videos longer than 1 hour. Use the
POSTmethod of the/analyze/tasksendpoint instead. - Video segmentation with custom segment definitions. Use the
POSTmethod of the/analyze/tasksendpoint instead.
On the Free plan, you have a total of 600 minutes (10 hours) shared across indexing, analysis, and segmentation. For details, see the Video hours and video count limits section.
This endpoint is rate-limited. For details, see the Rate limits page.
Authentication
Your API key.
You can find your API key on the API Keys page.
Request
The video understanding model to use for analysis.
pegasus1.6: For details about this version, see the Pegasus 1.6 page.pegasus1.5: For details about this version, see the Pegasus 1.5 page.
Default: pegasus1.5
An object specifying the source of the video content. Include exactly one source. Mutually exclusive with the image parameter.
A list of up to twenty objects containing the images to analyze. For each image, include exactly one source. Requires Pegasus 1.6. Using any other model returns a parameter_invalid error.
Mutually exclusive with the video and prompt_v2 parameters. The prompt parameter is required when you provide images.
A text prompt that guides the model on the desired format or content. To include reference images in your prompt, use the prompt_v2 parameter instead. Mutually exclusive with the prompt_v2 parameter.
Your prompts can be instructive or descriptive, or you can phrase them as questions. This text counts toward the context window.
A structured prompt that uses <@name> placeholders to reference images. Mutually exclusive with the prompt and image parameters.
The prompt text and reference images count toward the context window.
Controls the randomness of the text output.
Default: 0.2 Min: 0 Max: 1
Set this parameter to true to enable streaming responses in the NDJSON format.
Default: true
Specifies the format of the response. When you omit this parameter, the platform returns unstructured text. Only the json_schema type is supported for synchronous analysis.
Start of the analysis window, as an absolute timestamp in seconds, based on the internal metadata of the video. Use with end_time to analyze only a portion of the video.
- If omitted, defaults to the internal start time of the video.
- Most videos start at 0, but some (for example, from cameras or broadcast recordings) may have a non-zero start time. To find the value, run
ffprobe -v error -show_entries format=start_time,duration -of default=noprint_wrappers=1 your_video.mp4. - Must be less than
end_timeand the video duration. The window (end_time - start_time) must be at least 1 second.
End of the analysis window, as an absolute timestamp in seconds, based on the internal metadata of the video. Use with start_time to analyze only a portion of the video.
- If omitted, defaults to the internal start time of the video plus its duration.
- Most videos start at 0, but some (for example, from cameras or broadcast recordings) may have a non-zero start time. To find the value, run
ffprobe -v error -show_entries format=start_time,duration -of default=noprint_wrappers=1 your_video.mp4. - Must be greater than
start_timeand less than or equal to the video duration. The window (end_time - start_time) must be at least 1 second.
Response headers
The maximum number of requests you can make per rate limit window for this endpoint. For details, see the Rate limits page.
Response
When the value of the stream parameter is set to true, the platform provides a streaming response in the NDJSON format.
The stream contains the following types of events:
- Stream start
- Text generation
- Stream end
To integrate the response into your application, follow the guidelines below:
- Parse each line of the response as a separate JSON object.
- Check the
event_typefield to determine how to handle the event. - For
text_generationevents, process thetextfield as it arrives. Depending on your application’s requirements, this may involve displaying the text incrementally, storing it for later use, or performing any tasks. - Use the
stream_startandstream_endevents to manage the lifecycle of your streaming session.
When the value of the stream parameter is set to false, the response is as follows: