Skip to navigation

Create async embeddings

The EmbedClient.V2Client.TasksClient class provides methods to create embeddings asynchronously for audio, video, images, and documents.

Creating embeddings asynchronously requires three steps:

  1. Create a task using the create method. The platform returns a task ID.
  2. Poll for the status of the task using the retrieve method. Wait until the status is ready.
  3. Retrieve the embeddings from the response when the status is ready using the retrieve method.

Methods

List embedding tasks

Description: This method returns a list of the async embedding tasks in your account. The platform returns your async embedding tasks sorted by creation date, with the newest at the top of the list.

Notes
  • Embeddings are stored for seven days.
  • When you invoke this method without specifying the started_at and ended_at parameters, the platform returns all the async embedding tasks created within the last seven days.

Function signature and example:

def list(
self,
*,
started_at: typing.Optional[str] = None,
ended_at: typing.Optional[str] = None,
status: typing.Optional[str] = None,
page: typing.Optional[int] = None,
page_limit: typing.Optional[int] = None,
request_options: typing.Optional[RequestOptions] = None,
) -> SyncPager[MediaEmbeddingTask]

Parameters

NameTypeRequiredDescription
started_atstrNoRetrieve the embedding tasks that were created after the specified date and time, expressed in the RFC 3339 format (“YYYY-MM-DDTHH:mm:ssZ”).
ended_atstrNoRetrieve the embedding tasks that were created before the specified date and time, expressed in the RFC 3339 format (“YYYY-MM-DDTHH:mm:ssZ”).
statusstrNoFilter the embedding tasks by their current status. Values: processing, ready, or failed.
pageintNoA number that identifies the page to retrieve. Default: 1.
page_limitintNoThe number of items to return on each page. Default: 10. Max: 50.
request_optionsRequestOptionsNoRequest-specific configuration.

Return value

Returns a SyncPager[MediaEmbeddingTask] object that allows you to iterate through the paginated task results.

The SyncPager[T] class contains the following properties and methods:

NameTypeDescription
itemsOptional[List[T]]A list containing the current page of items. Can be None.
has_nextboolIndicates whether there is a next page to load.
get_nextOptional[Callable[[], Optional[SyncPager[T]]]]A callable function that retrieves the next page. Can be None.
responseOptional[BaseHttpResponse]The HTTP response object. Can be None.
next_page()Optional[SyncPager[T]]Calls get_next() if available and returns the next page object.
__iter__()Iterator[T]Allows iteration through all items across all pages using for loops.
iter_pages()Iterator[SyncPager[T]]Allows iteration through page objects themselves.

The MediaEmbeddingTask class contains the following properties:

NameTypeDescription
idOptional[str]The unique identifier of the embedding task.
model_nameOptional[str]The name of the video understanding model the platform used to create the embedding.
statusOptional[str]A string indicating the status of the embedding task. It can take one of the following values: processing, ready or failed.
created_atOptional[datetime]The date and time when the task was created.
updated_atOptional[datetime]The date and time when the task was last updated.
video_embeddingOptional[MediaEmbeddingTaskVideoEmbedding]An object containing the metadata associated with the embedding. See VideoEmbeddingMetadata for details.
audio_embeddingOptional[MediaEmbeddingTaskAudioEmbedding]An object containing the metadata associated with the embedding. See AudioEmbeddingMetadata for details.
document_embeddingOptional[MediaEmbeddingTaskDocumentEmbedding]An object containing the metadata associated with the embedding. Present only for document tasks created with Marengo 3.5. See DocumentEmbeddingMetadata for details.
image_embeddingOptional[MediaEmbeddingTaskImageEmbedding]An object containing the metadata associated with the embedding. Present only for image tasks created with Marengo 3.5. See ImageEmbeddingMetadata for details.

Each of the four objects above wraps a single metadata field. All four metadata classes inherit the following properties:

NameTypeDescription
input_urlOptional[str]The URL of the media file used to generate the embedding. Present if a URL was provided in the request.
input_filenameOptional[str]The name of the media file used to generate the embedding. Present if a file was provided in the request.

VideoEmbeddingMetadata

The VideoEmbeddingMetadata class contains the metadata associated with the embedding.

NameTypeDescription
video_clip_lengthOptional[float]The duration for each clip in seconds, as specified in the request. Note that the platform automatically truncates video segments shorter than 2 seconds. For a 31-second video divided into 6-second segments, the final 1-second segment will be truncated. This truncation only applies to the last segment if it does not meet the minimum length requirement of 2 seconds.
video_embedding_scopeOptional[List[str]]The scope you’ve specified in the request.
video_embedding_optionOptional[List[str]]The embedding_option values used to generate the embedding.
durationOptional[float]The total duration of the video in seconds.

AudioEmbeddingMetadata

The AudioEmbeddingMetadata class contains the metadata associated with the embedding.

NameTypeDescription
audio_embedding_optionOptional[List[str]]The type of the embedding. It can take one of the following values: ["audio"] or ["transcription"].
audio_embedding_scopeOptional[List[str]]The scope you’ve specified in the request.
durationOptional[float]The total duration of the audio in seconds.
start_offset_secOptional[float]The start offset in seconds from the beginning of the audio where processing should begin.
end_offset_secOptional[float]The end offset in seconds from the beginning of the audio where processing should end.

DocumentEmbeddingMetadata

The DocumentEmbeddingMetadata class contains the metadata associated with the embedding. Only Marengo 3.5 returns this object.

NameTypeDescription
document_embedding_optionOptional[List[str]]The embedding_option values used to generate the embedding.
document_embedding_scopeOptional[List[str]]The embedding_scope values used to generate the embedding.

ImageEmbeddingMetadata

The ImageEmbeddingMetadata class contains the metadata associated with the embedding. Only Marengo 3.5 returns this object.

NameTypeDescription
image_embedding_optionOptional[List[str]]The embedding_option values used to generate the embedding. Always ["visual"].
image_embedding_scopeOptional[List[str]]The embedding_scope values used to generate the embedding. Always ["asset"].

API Reference

List async embedding tasks

Create an async embedding task

Description: This method creates embeddings for audio, video, images, and documents asynchronously.

Use this method to embed content at scale, such as long files or the media files you want to make searchable. For a query, or for results you need in the same request, use the Create embeddings class instead.

The content this method accepts depends on the model. Both models embed audio and video. Marengo 3.5 also embeds images and documents. For the formats, resolutions, file sizes, and duration limits each model accepts, see the input requirements for Marengo 3.5 or Marengo 3.0.

Notes
  • Creating a task validates only basic metadata and playability, not the full file. A file can pass this check but still fail later during embedding. When you retrieve the results, check the status field. If it is failed, the error.message field contains the reason.
  • This method is rate-limited. With Marengo 3.5, the platform counts input tokens for each type of content. A task can exceed a limit before you see an error. For details, see Input token limits for embedding.
  • Embeddings are stored for seven days.

Function signature and example:

def create(
self,
*,
input_type: CreateAsyncEmbeddingRequestInputType,
model_name: CreateAsyncEmbeddingRequestModelName,
embedding_uncertainty: typing.Optional[bool] = OMIT,
embedding_dimension: typing.Optional[int] = OMIT,
audio: typing.Optional[AsyncAudioInputRequest] = OMIT,
video: typing.Optional[AsyncVideoInputRequest] = OMIT,
document: typing.Optional[AsyncDocumentInputRequest] = OMIT,
image: typing.Optional[AsyncImageInputRequest] = OMIT,
request_options: typing.Optional[RequestOptions] = None,
) -> TasksCreateResponse

Parameters

NameTypeRequiredDescription
input_typeCreateAsyncEmbeddingRequestInputTypeYesThe type of content for the embeddings. Values:
- audio: An audio file.
- video: A video file.
- document: A PDF, plain text, or Markdown file. Requires Marengo 3.5.
- image: An image file. Requires Marengo 3.5.
model_nameCreateAsyncEmbeddingRequestModelNameYesThe embedding model to use.

Values:
- marengo3.5: For details about this version, see the Marengo 3.5 page.
- marengo3.0: For details about this version, see the Marengo 3.0 page.
audioAsyncAudioInputRequestNoAudio input configuration. Required when input_type is audio. See AsyncAudioInputRequest for details.
videoAsyncVideoInputRequestNoVideo input configuration. Required when input_type is video. See AsyncVideoInputRequest for details.
documentAsyncDocumentInputRequestNoDocument input configuration. Required when input_type is document. Requires Marengo 3.5. See AsyncDocumentInputRequest for details.
imageAsyncImageInputRequestNoImage input configuration. Required when input_type is image. Requires Marengo 3.5. See AsyncImageInputRequest for details.
embedding_uncertaintyOptional[bool]NoSet this parameter to True to include a per-dimension uncertainty vector in the data[].embedding_uncertainty field of the result. The vector has the same length as the embedding array. A higher value indicates lower confidence in that dimension. Requires Marengo 3.5.

Requirements:
- For audio or video input, set the embedding_scope field to exclude asset. For example, set the video.embedding_scope field to ["clip"]. The field defaults to ["clip", "asset"], so the platform returns a 400 error if you keep the default. This requirement does not apply to image input.
- For a PDF document, the platform returns a 400 error regardless of the document.embedding_scope value.
- For a plain text or Markdown document, set the document.embedding_scope field to ["local"]. Any other value returns a 400 error.
embedding_dimensionOptional[int]NoThe number of dimensions for each embedding that the task produces, including the data[].embedding_uncertainty vector.

Marengo 3.5 produces Matryoshka embeddings: a shorter embedding consists of the first values of the full-length embedding. A 256-dimension embedding, for example, is the first 256 values of a 512-dimension embedding of the same content. Shorter embeddings reduce index size and speed up similarity search; longer embeddings produce higher retrieval quality.

Requirements:
- Requires Marengo 3.5. Setting this parameter when model_name is marengo3.0 returns a 400 error.
- Applies to the entire task: you cannot set it for a single input type or embedding.
- Set it once, when you create the task. To use a different value, create a new task.
- Use the same value across an index.

Default: 512
request_optionsRequestOptionsNoRequest-specific configuration.

AsyncAudioInputRequest

The AsyncAudioInputRequest class specifies the configuration for processing audio content. Required when input_type is audio.

Base64-encoded audio can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.

NameTypeRequiredDescription
media_sourceMediaSourceYesSpecifies the source of the audio file. See MediaSource for details.
start_secfloatNoThe start time in seconds for processing the audio file.

Use this parameter to process a portion of the audio file starting from a specific time.

Default: 0 (start from the beginning).
end_secfloatNoThe end time in seconds for processing the audio file.

Use this parameter to process a portion of the audio file ending at a specific time. The end time must be greater than the start time.

Default: End of the audio file
segmentationAsyncAudioInputRequestSegmentationNoSpecifies how the platform divides the audio into segments.

The structure of this object depends on the model version:

- With Marengo 3.5: Place your settings in the temporal object. Both strategies are available: dynamic divides the audio into variable-length segments that follow scene changes, and fixed divides it into equal-length segments. Default: temporal.dynamic, min_duration_sec: 2.
- With Marengo 3.0: Provide the settings directly in this object. Only fixed segmentation is available. Default: fixed, duration_sec: 6.

If you use a structure that does not match your model version, the platform returns a 400 error.

See AudioSegmentation and AsyncTemporalSegmentation for details.
embedding_optionList[str]NoThe types of embeddings you wish to generate.

Values:
- audio: Generates embeddings based on audio content (sounds, music, effects). With Marengo 3.5, this value includes speech, music, and non-dialog audio.
- transcription: Generates embeddings based on transcribed speech. Requires Marengo 3.0.

You can specify multiple values to generate different types of embeddings for the same audio.

Default: ["audio", "transcription"] for Marengo 3.0; ["audio"] for Marengo 3.5.
embedding_scopeList[str]NoThe scope for which you wish to generate embeddings.

Values:
- clip: Generates one embedding for each segment. Works with both Marengo 3.0 and Marengo 3.5.
- local: Generates one embedding for each segment. Equivalent to clip when using Marengo 3.5.
- asset: Generates one embedding for the entire audio file

You can specify multiple scopes to generate embeddings at different levels.

Default: ["clip", "asset"]
embedding_typeList[str]NoSpecifies how to structure the embedding. Include this parameter only when the embedding_option parameter contains at least two values.

Values:
- separate_embedding: Returns separate embeddings for each modality specified in the embedding_option parameter.
- fused_embedding: Returns a single combined embedding that integrates all modalities into one vector. With Marengo 3.5, this value requires the time_based_metadata field.

Specify both values to receive separate and fused embeddings in the same response.

Default: separate_embedding.
time_based_metadataOptional[List[TimeBasedMetadataEntry]]NoYour own time-aligned text, such as a stats feed or scene descriptions. The platform folds each entry into the fused embedding of the segments it overlaps in time, and it affects only that embedding. Requires the fused_embedding value in the embedding_type field. This field is supported only with Marengo 3.5. See TimeBasedMetadataEntry for details.

AsyncVideoInputRequest

The AsyncVideoInputRequest class specifies the configuration for processing video content. Required when input_type is video.

Base64-encoded video can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.

NameTypeRequiredDescription
media_sourceMediaSourceYesSpecifies the source of the video file. See MediaSource for details.
start_secfloatNoThe start time in seconds for processing the video file.

Use this parameter to process a portion of the video file starting from a specific time.

Default: 0 (start from the beginning)
end_secfloatNoThe end time in seconds for processing the video file.

Use this parameter to process a portion of the video file ending at a specific time. The end time must be greater than the start time.
Default: End of the video file
segmentationAsyncVideoInputRequestSegmentationNoSpecifies how the platform divides the video into segments.

The structure of this object depends on the model version:

- With Marengo 3.5: Place your settings in the temporal object. Both strategies are available: dynamic divides the video into variable-length segments that follow scene changes, and fixed divides it into equal-length segments. Default: temporal.dynamic, min_duration_sec: 2.
- With Marengo 3.0: Provide the settings directly in this object. Default: dynamic, min_duration_sec: 4.

If you use a structure that does not match your model version, the platform returns a 400 error.

See VideoSegmentation and AsyncTemporalSegmentation for details.
embedding_optionList[str]NoThe types of embeddings to generate for the video.

Values:
- visual: Generates embeddings based on visual content (scenes, objects, actions)
- audio: Generates embeddings based on audio content (sounds, music, effects). With Marengo 3.5, this value includes speech, music, and non-dialog audio.
- transcription: Generates embeddings based on transcribed speech. Requires Marengo 3.0.

You can specify multiple values to generate different types of embeddings for the same video.

Default: ["visual", "audio", "transcription"] for Marengo 3.0; ["visual", "audio"] for Marengo 3.5.
embedding_scopeList[str]NoThe scope for which you wish to generate embeddings.

Values:
- clip: Generates one embedding for each segment. Works with both Marengo 3.0 and Marengo 3.5.
- local: Generates one embedding for each segment. Equivalent to clip when using Marengo 3.5.
- asset: Generates one embedding for the entire video file. Use this scope for videos up to 10-30 seconds to maintain optimal performance.

You can specify multiple scopes to generate embeddings at different levels.

Default: ["clip", "asset"]
embedding_typeList[str]NoSpecifies how to structure the embedding. Include this parameter only when the embedding_option parameter contains at least two values.

Values:
- separate_embedding: Returns separate embeddings for each modality specified in the embedding_option parameter.
- fused_embedding: Returns a single combined embedding that integrates all modalities into one vector. With Marengo 3.5, this value requires the time_based_metadata field.

Specify both values to receive separate and fused embeddings in the same response.

Default: separate_embedding.
time_based_metadataOptional[List[TimeBasedMetadataEntry]]NoYour own time-aligned text, such as a stats feed or scene descriptions, which you can generate by segmenting a video with Pegasus. The platform folds each entry into the fused embedding of the segments it overlaps in time, and it affects only that embedding. Requires the fused_embedding value in the embedding_type field. This field is supported only with Marengo 3.5. See TimeBasedMetadataEntry for details.

TimeBasedMetadataEntry

One time-aligned metadata entry. The platform folds the text of the entry into the fused embedding of every segment that overlaps the time range of the entry.

Used by the AsyncAudioInputRequest.time_based_metadata and AsyncVideoInputRequest.time_based_metadata fields. Not applicable to the AsyncDocumentInputRequest class. Requires Marengo 3.5.

NameTypeRequiredDescription
startfloatYesThe start time of the entry in seconds, measured from the beginning of the asset. Set the same value in the end field for an event that happens at a single point in time, such as one entry in a stats feed.
endfloatYesThe end time of the entry in seconds, measured from the beginning of the asset.
textstrYesThe text to fold into the fused embedding of the overlapping segments.

AsyncDocumentInputRequest

The AsyncDocumentInputRequest class specifies the configuration for processing documents. Requires Marengo 3.5.

The platform accepts PDF (.pdf), plain text (.txt), and Markdown (.md) files. The decoded file can be up to 512 MB. It embeds a PDF file from its rendered pages or from its extracted text, and a plain text or Markdown file from its text.

A PDF file also has a page allowance: 64 pages for each MB of file size. A 0.5 MB file is allowed 64 pages, and a 4 MB file is allowed 256 pages. The quadrants strategy counts each page five times against this allowance. The platform checks the page count of the file against the allowance before processing the file. If the file exceeds the allowance, the platform creates the task and sets its status field to the failed value. The error.message field contains the page count and the allowance. Plain text and Markdown files have no page allowance.

The embedding_option and embedding_scope fields combine, and the platform supports the following combinations.

File typeembedding_optionembedding_scopeResult
PDFvisuallocalOne embedding for each rendered page. The default for PDF files.
PDFvisualassetOne embedding for the entire file.
PDFtextassetOne embedding for the extracted text of the entire file.
Plain text, MarkdowntextassetOne embedding for the entire file. The default for plain text and Markdown files.
Plain text, MarkdowntextlocalOne embedding for each chunk of whole sentences. Requires the segmentation.sequential field.

You can request more than one combination at a time. For example, embedding_scope: ["local", "asset"] on a PDF file returns the per-page embeddings and the whole-file embedding together. The platform pairs each value in one field with each value in the other. Each pair must appear in this table; if you send a pair outside it, the platform returns a 400 error. If you omit a field, the platform uses its default value. If you embed a PDF file with the embedding_option field set to ["text"], also set the embedding_scope field to ["asset"]. For PDF files, the default ["local"] pairs with only the visual option.

NameTypeRequiredDescription
media_sourceMediaSourceYesSpecifies the source of the document file. See MediaSource for details.
segmentationOptional[DocumentSegmentation]NoSpecifies how the platform divides your document before it generates embeddings. Use the spatial field to divide each rendered page of a PDF file. Use the sequential field to divide a plain text or Markdown file into chunks. See DocumentSegmentation for details.
embedding_optionOptional[List[str]]NoThe types of embeddings to generate for the document.

Values:
- visual: Generates embeddings from the rendered pages. Valid for PDF files.
- text: Generates embeddings from the text content. Valid for PDF, plain text, and Markdown files.

Default: ["visual"] for PDF files; ["text"] for plain text and Markdown files.
embedding_typeOptional[List[str]]NoSpecifies how to structure the embedding.

Values:
- separate_embedding: Returns one embedding per requested embedding_scope.
- fused_embedding: The platform returns a 400 error if you set this value. Documents have a single modality.

Default: separate_embedding.
embedding_scopeOptional[List[str]]NoThe scope for which you wish to generate embeddings.

Values:
- local: Returns one embedding for each part of the file. For a PDF file, each part is a rendered page, and you can divide each page further with the segmentation.spatial field. For a plain text or Markdown file, each part is a chunk of whole sentences, and the segmentation.sequential field is required.
- asset: Returns one embedding for the entire file.

Default: ["local"] for PDF files; ["asset"] for plain text and Markdown files.

DocumentSegmentation

The DocumentSegmentation class specifies how the platform divides your document before it generates embeddings. Requires Marengo 3.5.

Use the spatial field to divide each rendered page of a PDF file. Use the sequential field to divide a plain text or Markdown file into chunks.

Provide the field that matches your file. If you provide neither field, the platform returns a 400 error.

NameTypeRequiredDescription
spatialOptional[DocumentSpatialSegmentation]NoSpecifies how the platform divides each rendered page of a PDF file. See DocumentSpatialSegmentation for details.
sequentialOptional[DocumentSequentialSegmentation]NoSpecifies how the platform divides a plain text or Markdown file into chunks. See DocumentSequentialSegmentation for details.

DocumentSpatialSegmentation

The DocumentSpatialSegmentation class specifies how the platform divides each rendered page of a PDF file. Plain text and Markdown files have no rendered pages, so the platform returns a 400 error if you include this object.

This object requires the visual value in the embedding_option field and the local value in the embedding_scope field.

NameTypeRequiredDescription
strategyOptional[DocumentSpatialSegmentationStrategy]NoThe strategy for dividing each page.

Values:
- standard: Returns one embedding for each page.
- quadrants: Divides each page into a 2×2 grid. Returns five embeddings: one for the whole page, and one for each quarter. The data[].quadrant field identifies which quarter each embedding represents. This field is null on the whole-page embedding. This strategy uses five times as many tokens as the standard strategy. It also counts each page five times against the page allowance of the file.

Default: standard

DocumentSequentialSegmentation

The DocumentSequentialSegmentation class specifies how the platform divides a plain text or Markdown file into chunks. For a PDF file, the platform divides by page instead and returns a 400 error if you include this object.

This object requires the text value in the embedding_option field and the local value in the embedding_scope field.

NameTypeRequiredDescription
strategyDocumentSequentialSegmentationStrategyYesThe strategy for dividing the text into chunks. Always sentence, which groups whole sentences into each chunk, up to the number of sentences in the max_sentences field.
max_sentencesintYesThe maximum number of sentences in each chunk. This field has no default.

Choose a value small enough that every chunk fits in the context window of the model. If a chunk exceeds that window, the task fails, and the platform does not truncate it. How many sentences fit depends on the length of the sentences in your file.
overlap_sentencesOptional[int]NoThe number of sentences at the end of one chunk that the platform repeats at the start of the next. The overlap preserves the context of the previous chunk. This value must be less than the max_sentences value.

Default: 0

AsyncImageInputRequest

The AsyncImageInputRequest class specifies the configuration for processing image content. Required when input_type is image. Requires Marengo 3.5. The decoded file can be up to 32 MB.

For an image, the embedding_option, embedding_type, and embedding_scope fields each accept a single value; the platform returns a 400 error if you send any other value.

NameTypeRequiredDescription
media_sourceMediaSourceYesSpecifies the source of the image file. See MediaSource for details.
embedding_optionOptional[List[str]]NoThe type of embedding to generate for the image. Always visual.
embedding_typeOptional[List[str]]NoSpecifies how to structure the embedding. Always separate_embedding.
embedding_scopeOptional[List[str]]NoThe scope for which to generate embeddings. Always asset, which produces one embedding for the entire image.

MediaSource

The MediaSource class specifies the source of the media file. Provide exactly one of the following:

NameTypeRequiredDescription
base_64_stringstrNoThe base64-encoded media data. Encoding grows the payload by about a third, so the string you send is larger than the original file.

The maximum size depends on the input type and the model. The description of the field that contains this media source states the limit where it differs; for the formats and sizes each model accepts, see the input requirements for Marengo 3.5 or Marengo 3.0.
urlstrNoThe publicly accessible URL of the media file. Use direct links to raw media files. Video hosting platforms and cloud storage sharing links are not supported.
asset_idstrNoThe unique identifier of an asset from a direct or multipart upload. The asset status must be ready. Use assets.retrieve to check the status.

AudioSegmentation

The AudioSegmentation class specifies how the platform divides the audio into segments using fixed-length intervals.

NameTypeRequiredDescription
strategyAudioSegmentationStrategyYesThe segmentation strategy. Value: fixed.
fixedAudioSegmentationFixedYesConfiguration for fixed segmentation.

This object is required when the strategy field is fixed. See AudioSegmentationFixed for details.

AudioSegmentationFixed

The AudioSegmentationFixed class configures fixed-length segmentation for audio.

NameTypeRequiredDescription
duration_secintYesThe duration in seconds for each segment. The platform divides the audio into segments of this exact length. The final segment may be shorter if the audio duration is not evenly divisible.

Min: 2.
Max: 10.

Example: With duration_sec: 5, a 12-second audio file produces segments: [0-5s], [5-10s], [10-12s].

VideoSegmentation

The VideoSegmentation type specifies how the platform divides the video into segments. Use one of the following:

Fixed segmentation: Divides the video into equal-length segments:

NameTypeRequiredDescription
strategyLiteral["fixed"]YesThe segmentation strategy. Value: fixed.
fixedVideoSegmentationFixedFixedYesConfiguration for fixed segmentation. See VideoSegmentationFixedFixed for details.

Dynamic segmentation: Divides the video into adaptive segments based on scene changes:

NameTypeRequiredDescription
strategyLiteral["dynamic"]YesThe segmentation strategy. Value: dynamic.
dynamicVideoSegmentationDynamicDynamicYesConfiguration for dynamic segmentation. See VideoSegmentationDynamicDynamic for details.

VideoSegmentationFixedFixed

The VideoSegmentationFixedFixed class configures fixed-length segmentation for video.

NameTypeRequiredDescription
duration_secintYesThe duration in seconds for each segment.

The platform divides the video into segments of this exact length. The final segment may be shorter if the video duration is not evenly divisible.

Min: 2.
Max: 10.

Example: With duration_sec: 5, a 12-second video produces segments: [0-5s], [5-10s], [10-12s].

VideoSegmentationDynamicDynamic

The VideoSegmentationDynamicDynamic class configures dynamic segmentation for video based on scene changes.

NameTypeRequiredDescription
min_duration_secintYesThe minimum duration in seconds for each segment.

The platform divides the video into segments that are at least this long. Segments adapt to scene changes and content boundaries and may be longer than the minimum.

Min: 2.
Max: 5.

Example: With min_duration_sec: 3, segments might be: [0-3.2s], [3.2-7.8s], [7.8-12.1s]

AsyncTemporalSegmentation

The AsyncTemporalSegmentation class wraps your settings in a temporal object. Use with Marengo 3.5.

NameTypeRequiredDescription
temporalTemporalSegmentationYesSpecifies how the platform divides the file into segments. See TemporalSegmentation for details.

TemporalSegmentation

The TemporalSegmentation type specifies how the platform divides the file into segments. The strategy field selects one variant:

Dynamic segmentation: Creates variable-length segments that align with scene or content boundaries. Use this for content-aware segmentation.

NameTypeRequiredDescription
strategyLiteral["dynamic"]YesMust be dynamic. Identifies this as the content-aware segmentation variant.
dynamicTemporalSegmentationDynamicDynamicYesConfiguration for dynamic segmentation. This object is required when strategy is dynamic. See TemporalSegmentationDynamicDynamic for details.

Fixed segmentation: Creates equal-length segments. Use this for consistent timing.

NameTypeRequiredDescription
strategyLiteral["fixed"]YesMust be fixed. Identifies this as the equal-length segmentation variant.
fixedTemporalSegmentationFixedFixedYesConfiguration for fixed segmentation. This object is required when strategy is fixed. See TemporalSegmentationFixedFixed for details.

TemporalSegmentationDynamicDynamic

The TemporalSegmentationDynamicDynamic class configures dynamic segmentation. This object is required when strategy is dynamic.

NameTypeRequiredDescription
min_duration_secintYesThe minimum duration in seconds for each segment.

The platform divides the file into segments that are at least this long. Segments adapt to scene changes and content boundaries and may be longer than the minimum.

Min: 2.
Max: 5.

TemporalSegmentationFixedFixed

The TemporalSegmentationFixedFixed class configures fixed segmentation. This object is required when strategy is fixed.

NameTypeRequiredDescription
duration_secintYesThe duration in seconds for each segment. The platform divides the file into segments of this exact length. The final segment may be shorter if the duration is not evenly divisible.

Min: 2.
Max: 10.

Return value

Returns a TasksCreateResponse object containing the task details.

The TasksCreateResponse class contains the following properties:

NameTypeDescription
idstrThe unique identifier of the embedding task.
statusLiteral["processing"]The initial status of the embedding task. Value: processing.
dataOptional[List[EmbeddingData]]Array of embedding results. Present when status is ready; null when status is processing or failed.
metadataOptional[TasksCreateResponseMetadata]Metadata about the task you created. The platform sets the length of your embeddings when it creates the task, and the embedding_dimension field contains that length. Only Marengo 3.5 returns this object.

The TasksCreateResponseMetadata class contains the following properties:

NameTypeDescription
embedding_dimensionOptional[EmbeddingDimension]The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field.

API Reference

Create an async embedding task

Retrieve task status and results

Description: This method retrieves the status and the results of an async embedding task.

Invoke this method repeatedly until the status field is ready or failed. When the status is ready, use the embeddings from the response. When the status is failed, the error.message field contains the reason.

Note

Embeddings are stored for seven days.

Function signature and example:

def retrieve(
self,
task_id: str,
*,
request_options: typing.Optional[RequestOptions] = None
) -> EmbeddingTaskResponse

Parameters

NameTypeRequiredDescription
task_idstrYesThe unique identifier of the embedding task.
request_optionsRequestOptionsNoRequest-specific configuration.

Return value

Returns an EmbeddingTaskResponse object containing the task status and results.

The EmbeddingTaskResponse class contains the following properties:

NameTypeDescription
idstrThe unique identifier of the embedding task.
statusEmbeddingTaskResponseStatusThe current status of the task.

Values:
- processing: The platform is creating the embeddings
- ready: Processing is complete. Embeddings are available in the data field
- failed: The task failed. The data field is null, and the error.message field contains the reason
created_atOptional[datetime]The date and time when the task was created.
updated_atOptional[datetime]The date and time when the task was last updated.
dataOptional[List[EmbeddingData]]An object containing the embedding results, or null otherwise.
usageOptional[EmbeddingUsage]Token counts for the request. Only Marengo 3.5 returns this field. See EmbeddingUsage for details.
metadataOptional[EmbeddingTaskMediaMetadata]Metadata for the media input. See EmbeddingTaskMediaMetadata for details.
errorOptional[EmbeddingTaskResponseError]An object describing why the embedding task failed. Present only when status is failed. Omitted otherwise.

The EmbeddingData class contains the following properties:

NameTypeDescription
embeddingList[float]The embedding vector for the content.
embedding_uncertaintyOptional[List[float]]A per-dimension uncertainty vector with the same length as the embedding array. A higher value indicates lower confidence in that dimension. Present when the request sets embedding_uncertainty: true. Only Marengo 3.5 returns this field.
embedding_optionOptional[EmbeddingDataEmbeddingOption]The type of the embedding.

Values:
- visual: Embedding based on visual content (a video, a page of a PDF file, or an image embedded asynchronously).
- audio: Embedding based on audio content.
- transcription: Embedding based on transcribed speech. Returned only for content embedded with Marengo 3.0.
- text: Embedding based on the text content of a PDF, plain text, or Markdown file embedded asynchronously.
- fused: Embedding based on a combination of the modalities specified in the request. The platform returns this embedding only for video and audio input, and only when the embedding_type parameter includes the fused_embedding value.
- null: For text embeddings and images embedded synchronously.
embedding_scopeOptional[EmbeddingDataEmbeddingScope]The scope for which the embedding was generated.

Values:
- clip: Embedding for a segment. For video and audio input, one embedding per detected segment.
- page: Embedding for one page of a PDF file embedded asynchronously, or for one quadrant of a page when the request sets the document.segmentation.spatial.strategy field to the quadrants strategy. With that strategy, five entries share the same scope and page numbers, so read the quadrant field to tell them apart: the whole-page entry has no quadrant value.
- chunk: Embedding for one chunk of whole sentences of a plain text or Markdown file embedded asynchronously. Read the chunk_index field for the position of the chunk in the file.
- asset: Embedding for the entire file. For video and audio input, use this scope for content up to 10-30 seconds to maintain optimal performance.
- null: For text embeddings and images embedded synchronously.

When you request the local scope, the platform returns clip for audio and video, page for PDF files, and chunk for plain text and Markdown files. For audio, video, and document input, the metadata.embedding_scopes field contains the scopes you requested.
start_secOptional[float]The start time in seconds for this segment. This field is null for text and image embeddings.
end_secOptional[float]The end time in seconds for this segment. This field is null for text and image embeddings.
start_page_numberOptional[int]The first page this embedding covers, counting from 1. The platform returns this field only for page-level embeddings of a PDF file, and null in every other case.
quadrantOptional[str]The quarter of the page this embedding covers. The platform returns this field only when the request sets the document.segmentation.spatial.strategy field to the quadrants strategy, and only on the four quadrant embeddings of a page. This field is null on the whole-page embedding and in every other case.
chunk_indexOptional[int]The position of this chunk in the file, counting from 0. The platform returns this field only on chunk-scope embeddings of a plain text or Markdown file, and null in every other case. Read this field rather than the position of the entry in the data array, which provides no ordering guarantee.
end_page_numberOptional[int]The last page this embedding covers, counting from 1 and including that page. This field matches the start_page_number field when the embedding covers a single page. The platform returns this field only for page-level embeddings of a PDF file, and null in every other case.

EmbeddingUsage

The EmbeddingUsage class provides token counts for the request. Only Marengo 3.5 returns this object.

NameTypeDescription
input_tokensDict[str, int]The number of tokens the request used. Each key names a type of content the request processed, and each value is the token count for that content.

The platform reports the content types your request actually used. Each key is one of the following: video, audio, image, document, or text. Read the keys the response returns rather than assuming a fixed set.
truncatedboolWhether the input was truncated to fit within the token limit.
truncation_reasonLiteral["model_context_window"]The reason the input was truncated. Present only when the truncated field is true.

Values:
- model_context_window: The input exceeded the context window of the model.

EmbeddingTaskMediaMetadata

The EmbeddingTaskMediaMetadata type provides metadata for the media input. The input_type field selects one variant:

Audio: Metadata for audio embeddings.

NameTypeDescription
input_typeLiteral["audio"]The type of the input content. Value: audio.
input_urlOptional[str]The publicly accessible URL for the audio file.
input_filenameOptional[str]The name of the audio file.
embedding_optionsList[str]The embedding_option values used to generate the embedding.
embedding_scopesList[EmbeddingAudioMetadataEmbeddingScopesItem]The embedding_scope values used to generate the embedding.
embedding_dimensionOptional[int]The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field.
durationfloatThe duration of the audio in seconds.
start_offset_secOptional[float]The start offset in seconds.
end_offset_secOptional[float]The end offset in seconds.

Video: Metadata for video embeddings.

NameTypeDescription
input_typeLiteral["video"]The type of the input content. Value: video.
input_urlOptional[str]The publicly accessible URL for the video file.
input_filenameOptional[str]The name of the video file.
clip_lengthOptional[int]Length of each video clip in seconds. Only available for fixed segmentation.
embedding_scopesList[EmbeddingVideoMetadataEmbeddingScopesItem]The embedding_scope values used to generate the embedding.
embedding_dimensionOptional[int]The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field.
embedding_optionsList[str]The embedding_option values used to generate the embedding.
durationfloatThe duration of the video in seconds.
start_offset_secOptional[float]The start offset in seconds.
end_offset_secOptional[float]The end offset in seconds.

Document: Metadata for document embeddings. Only Marengo 3.5 returns this object.

NameTypeDescription
input_typeLiteral["document"]The type of the input content. Value: document.
input_urlOptional[str]The publicly accessible URL for the document file.
input_filenameOptional[str]The name of the document file.
embedding_optionsOptional[List[str]]The embedding_option values used to generate the embedding.
embedding_scopesOptional[List[AsyncDocumentMetadataEmbeddingScopesItem]]The embedding_scope values used to generate the embedding.
embedding_dimensionOptional[int]The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field.

Image: Metadata for image embeddings. Only Marengo 3.5 returns this object.

NameTypeDescription
input_typeLiteral["image"]The type of the input content. Value: image.
input_urlOptional[str]The publicly accessible URL for the image file.
input_filenameOptional[str]The name of the image file.
embedding_optionsOptional[List[str]]The embedding_option values used to generate the embedding. Always ["visual"].
embedding_scopesOptional[List[AsyncImageMetadataEmbeddingScopesItem]]The embedding_scope values used to generate the embedding. Always ["asset"].
embedding_dimensionOptional[int]The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field.

The EmbeddingTaskResponseError class contains the following property:

NameTypeDescription
messagestrA human-readable message that describes why the task failed. Possible values:
- “The embedding service is temporarily unstable. Please try again later.”
- “The embedding task failed. Please try again later.”
- “The embedding task could not complete processing.”
- “The embedding request was invalid.”
- “The media exceeds the maximum allowed size for embedding.”
- “Content exceeds the model’s context window.”
- “The embedding result is no longer available and cannot be retrieved.”
- “failed to fetch or materialize source media”
- “We could not process your media for embedding. Please verify the input file and try again.” For the steps to fix the file, see the How do I fix a file that could not be processed for embedding? section on the Frequently asked questions page.

When the platform can measure the overage in pages, bytes, or seconds, the message includes those numbers in place of the fixed sentence for that limit. For example: “Document (17 pages) exceeds maximum limit of 6 pages”, “File size (33.6 MB) exceeds maximum limit of 32 MB”, or “Duration (35.0 seconds) exceeds maximum limit of 30 seconds (0 minutes)”.

API Reference

Retrieve task status and results