Skip to navigation

Create an async embedding task

This method creates embeddings for audio, video, images, and documents asynchronously.

Use this method to embed content at scale, such as long files or the media files you want to make searchable. For a query, or for results you need in the same request, use the POST method of the /embed-v2 endpoint instead.

The content this method accepts depends on the model. Both models embed audio and video. Marengo 3.5 also embeds images and documents: PDF, plain text, and Markdown files. For the formats, resolutions, file sizes, and duration limits each model accepts, see the input requirements for Marengo 3.5 or Marengo 3.0.

Creating embeddings asynchronously requires three steps:

  1. Create a task using this method. The platform returns a task identifier.
  2. Poll for the status of the task using the GET method of the /embed-v2/tasks/{task_id} endpoint. Wait until the status is ready.
  3. Retrieve the embeddings from the response when the status is ready using the GET method of the /embed-v2/tasks/{task_id} endpoint.
Notes
  • Creating a task validates only basic metadata and, for audio and video sources, playability, not the full file. A file can pass this check but still fail later during embedding. When you retrieve the results, check the status field. If it is failed, the error.message field contains the reason.
  • This method is rate-limited. With Marengo 3.5, the platform counts input tokens for each type of content. A task can exceed a limit before you see an error. For details, see Input token limits for embedding.
  • Embeddings are stored for seven days.

Authentication

x-api-keystring

Your API key.

Note

You can find your API key on the API Keys page.

Request

This endpoint expects an object.
input_typeenumRequired

The type of content for the embeddings.

Values:

  • audio: An audio file.
  • video: A video file.
  • document: A PDF, plain text, or Markdown file. Requires Marengo 3.5.
  • image: An image file. Requires Marengo 3.5.
Allowed values:
model_nameenumRequiredDefaults to marengo3.0

The embedding model to use.

Values:

  • marengo3.5: For details about this version, see the Marengo 3.5 page.
  • marengo3.0: For details about this version, see the Marengo 3.0 page.
Allowed values:
embedding_uncertaintybooleanOptionalDefaults to false

Set this parameter to true to include a per-dimension uncertainty vector in the data[].embedding_uncertainty field of the result. The vector has the same length as the embedding array. A higher value indicates lower confidence in that dimension. Requires Marengo 3.5.

Requirements:

  • For audio or video input, set the embedding_scope field to exclude asset. For example, set the video.embedding_scope field to ["clip"]. The field defaults to ["clip", "asset"], so the platform returns a 400 error if you keep the default. This requirement does not apply to image input.
  • For a PDF document, the platform returns a 400 error regardless of the document.embedding_scope value.
  • For a plain text or Markdown document, set the document.embedding_scope field to ["local"]. Any other value returns a 400 error.
embedding_dimensionenumOptional

The number of dimensions for each embedding that the task produces, including the data[].embedding_uncertainty vector.

Marengo 3.5 produces Matryoshka embeddings: a shorter embedding consists of the first values of the full-length embedding. A 256-dimension embedding, for example, is the first 256 values of a 512-dimension embedding of the same content. Shorter embeddings reduce index size and speed up similarity search; longer embeddings produce higher retrieval quality.

Requirements:

  • Requires Marengo 3.5. Setting this parameter with model_name: marengo3.0 returns a 400 error.
  • Applies to the entire task: you cannot set it for a single input type or embedding.
  • Set it once, when you create the task. To use a different value, create a new task.
  • Use the same value across an index.

Default: 512

Allowed values:
audioobjectOptional

This field is required if the input_type parameter is audio.

Base64-encoded audio can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.

videoobjectOptional

This field is required if the input_type parameter is video.

Base64-encoded video can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.

documentobjectOptional

This field is required if the input_type parameter is document. Requires Marengo 3.5.

The platform accepts PDF (.pdf), plain text (.txt), and Markdown (.md) files. The decoded file can be up to 512 MB. It embeds a PDF file from its rendered pages or from its extracted text, and a plain text or Markdown file from its text.

A PDF file also has a page allowance: 64 pages for each MB of file size. A 0.5 MB file is allowed 64 pages, and a 4 MB file is allowed 256 pages. The quadrants strategy counts each page five times against this allowance. The platform checks the page count of the file against the allowance before processing the file. If the file exceeds the allowance, the platform creates the task and sets its status field to the failed value. The error.message field contains the page count and the allowance. Plain text and Markdown files have no page allowance.

The embedding_option and embedding_scope fields combine, and the platform supports the following combinations.

File typeembedding_optionembedding_scopeResult
PDFvisuallocalOne embedding for each rendered page. The default for PDF files.
PDFvisualassetOne embedding for the entire file.
PDFtextassetOne embedding for the extracted text of the entire file.
Plain text, MarkdowntextassetOne embedding for the entire file. The default for plain text and Markdown files.
Plain text, MarkdowntextlocalOne embedding for each chunk of whole sentences. Requires the segmentation.sequential field.

You can request more than one combination at a time. For example, embedding_scope: ["local", "asset"] on a PDF file returns the per-page embeddings and the whole-file embedding together. The platform pairs each value in one field with each value in the other. Each pair must appear in this table; if you send a pair outside it, the platform returns a 400 error. If you omit a field, the platform uses its default value. If you embed a PDF file with embedding_option: ["text"], also set embedding_scope: ["asset"]. For PDF files, the default ["local"] pairs with only the visual option.

imageobjectOptional

This field is required if the input_type parameter is image. Requires Marengo 3.5. The decoded file can be up to 32 MB. For an image, the embedding_option, embedding_type, and embedding_scope fields each accept a single value; the platform returns a 400 error if you send any other value.

Response headers

LocationstringOptional
URL to poll for task status and results
X-Ratelimit-DimensionsstringOptional

A comma-separated, alphabetically sorted list of the rate limits the request was measured against. Only the limits that applied to the request appear.

Each entry is a label. Replace <label> with an entry to read the values for that limit: X-Ratelimit-<label>-Limit, X-Ratelimit-<label>-Remaining, and X-Ratelimit-<label>-Reset. This list keeps the mixed-case spelling of each label, such as InputToken-Video. The platform sends the headers as X-Ratelimit-Inputtoken-Video-Limit. HTTP header names are case-insensitive, so the two spellings match.

The /embed-v2 and /embed-v2/tasks endpoints report every limit through this family. They do not send the aggregate X-Ratelimit-Limit, X-Ratelimit-Remaining, X-Ratelimit-Used, or X-Ratelimit-Reset headers that other endpoints send.

X-Ratelimit-Request-Limitinteger

The maximum number of requests you can make per rate limit window for this endpoint. For details, see the Rate limits page.

X-Ratelimit-Request-Remaininginteger
The number of requests remaining in the current rate limit window for this endpoint.
X-Ratelimit-Request-Resetinteger
The time at which the current request rate limit window resets, expressed in UTC epoch seconds.
X-Ratelimit-Inputtoken-Video-Limitinteger
The maximum number of input tokens per rate limit window for the type of content named in this header.
X-Ratelimit-Inputtoken-Video-Remaininginteger

The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.

X-Ratelimit-Inputtoken-Video-Resetinteger
The time at which the current input token rate limit window resets for the type of content named in this header, expressed in UTC epoch seconds.
X-Ratelimit-Inputtoken-Audio-Limitinteger
The maximum number of input tokens per rate limit window for the type of content named in this header.
X-Ratelimit-Inputtoken-Audio-Remaininginteger

The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.

X-Ratelimit-Inputtoken-Audio-Resetinteger
The time at which the current input token rate limit window resets for the type of content named in this header, expressed in UTC epoch seconds.
X-Ratelimit-Inputtoken-Image-Limitinteger
The maximum number of input tokens per rate limit window for the type of content named in this header.
X-Ratelimit-Inputtoken-Image-Remaininginteger

The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.

X-Ratelimit-Inputtoken-Image-Resetinteger
The time at which the current input token rate limit window resets for the type of content named in this header, expressed in UTC epoch seconds.
X-Ratelimit-Inputtoken-Document-Limitinteger
The maximum number of input tokens per rate limit window for the type of content named in this header.
X-Ratelimit-Inputtoken-Document-Remaininginteger

The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.

X-Ratelimit-Inputtoken-Document-Resetinteger
The time at which the current input token rate limit window resets for the type of content named in this header, expressed in UTC epoch seconds.
X-Ratelimit-Inputtoken-Text-Limitinteger
The maximum number of input tokens per rate limit window for the type of content named in this header.
X-Ratelimit-Inputtoken-Text-Remaininginteger

The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.

X-Ratelimit-Inputtoken-Text-Resetinteger
The time at which the current input token rate limit window resets for the type of content named in this header, expressed in UTC epoch seconds.

Response

An embedding task has successfully been created.
_idstring
The unique identifier of the embedding task
statusenum
The initial status of the embedding task.
Allowed values:
datalist of objects or nullOptional

An array of embedding results when status is ready, or null when status is processing or failed.

metadataobjectOptional

Metadata about the task you created. The platform sets the length of your embeddings when it creates the task, and the embedding_dimension field contains that length. Only Marengo 3.5 returns it.

Errors

400
Bad Request Error
429
Too Many Requests Error
500
Internal Server Error