Create an async embedding task
This method creates embeddings for audio, video, images, and documents asynchronously.
Use this method to embed content at scale, such as long files or the media files you want to make searchable. For a query, or for results you need in the same request, use the POST method of the /embed-v2 endpoint instead.
The content this method accepts depends on the model. Both models embed audio and video. Marengo 3.5 also embeds images and documents: PDF, plain text, and Markdown files. For the formats, resolutions, file sizes, and duration limits each model accepts, see the input requirements for Marengo 3.5 or Marengo 3.0.
Creating embeddings asynchronously requires three steps:
- Create a task using this method. The platform returns a task identifier.
- Poll for the status of the task using the
GETmethod of the/embed-v2/tasks/{task_id}endpoint. Wait until the status isready. - Retrieve the embeddings from the response when the status is
readyusing theGETmethod of the/embed-v2/tasks/{task_id}endpoint.
- Creating a task validates only basic metadata and, for audio and video sources, playability, not the full file. A file can pass this check but still fail later during embedding. When you retrieve the results, check the
statusfield. If it isfailed, theerror.messagefield contains the reason. - This method is rate-limited. With Marengo 3.5, the platform counts input tokens for each type of content. A task can exceed a limit before you see an error. For details, see Input token limits for embedding.
- Embeddings are stored for seven days.
Authentication
Your API key.
You can find your API key on the API Keys page.
Request
The type of content for the embeddings.
Values:
audio: An audio file.video: A video file.document: A PDF, plain text, or Markdown file. Requires Marengo 3.5.image: An image file. Requires Marengo 3.5.
The embedding model to use.
Values:
marengo3.5: For details about this version, see the Marengo 3.5 page.marengo3.0: For details about this version, see the Marengo 3.0 page.
Set this parameter to true to include a per-dimension uncertainty vector in the data[].embedding_uncertainty field of the result. The vector has the same length as the embedding array. A higher value indicates lower confidence in that dimension. Requires Marengo 3.5.
Requirements:
- For audio or video input, set the
embedding_scopefield to excludeasset. For example, set thevideo.embedding_scopefield to["clip"]. The field defaults to["clip", "asset"], so the platform returns a400error if you keep the default. This requirement does not apply to image input. - For a PDF document, the platform returns a
400error regardless of thedocument.embedding_scopevalue. - For a plain text or Markdown document, set the
document.embedding_scopefield to["local"]. Any other value returns a400error.
The number of dimensions for each embedding that the task produces, including the data[].embedding_uncertainty vector.
Marengo 3.5 produces Matryoshka embeddings: a shorter embedding consists of the first values of the full-length embedding. A 256-dimension embedding, for example, is the first 256 values of a 512-dimension embedding of the same content. Shorter embeddings reduce index size and speed up similarity search; longer embeddings produce higher retrieval quality.
Requirements:
- Requires Marengo 3.5. Setting this parameter with
model_name: marengo3.0returns a400error. - Applies to the entire task: you cannot set it for a single input type or embedding.
- Set it once, when you create the task. To use a different value, create a new task.
- Use the same value across an index.
Default: 512
This field is required if the input_type parameter is audio.
Base64-encoded audio can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.
This field is required if the input_type parameter is video.
Base64-encoded video can be up to 36 MB decoded. For a larger file, provide a URL or an asset identifier.
This field is required if the input_type parameter is document. Requires Marengo 3.5.
The platform accepts PDF (.pdf), plain text (.txt), and Markdown (.md) files. The decoded file can be up to 512 MB. It embeds a PDF file from its rendered pages or from its extracted text, and a plain text or Markdown file from its text.
A PDF file also has a page allowance: 64 pages for each MB of file size. A 0.5 MB file is allowed 64 pages, and a 4 MB file is allowed 256 pages. The quadrants strategy counts each page five times against this allowance. The platform checks the page count of the file against the allowance before processing the file. If the file exceeds the allowance, the platform creates the task and sets its status field to the failed value. The error.message field contains the page count and the allowance. Plain text and Markdown files have no page allowance.
The embedding_option and embedding_scope fields combine, and the platform supports the following combinations.
You can request more than one combination at a time. For example, embedding_scope: ["local", "asset"] on a PDF file returns the per-page embeddings and the whole-file embedding together. The platform pairs each value in one field with each value in the other. Each pair must appear in this table; if you send a pair outside it, the platform returns a 400 error. If you omit a field, the platform uses its default value. If you embed a PDF file with embedding_option: ["text"], also set embedding_scope: ["asset"]. For PDF files, the default ["local"] pairs with only the visual option.
This field is required if the input_type parameter is image. Requires Marengo 3.5. The decoded file can be up to 32 MB. For an image, the embedding_option, embedding_type, and embedding_scope fields each accept a single value; the platform returns a 400 error if you send any other value.
Response headers
A comma-separated, alphabetically sorted list of the rate limits the request was measured against. Only the limits that applied to the request appear.
Each entry is a label. Replace <label> with an entry to read the values for that limit: X-Ratelimit-<label>-Limit, X-Ratelimit-<label>-Remaining, and X-Ratelimit-<label>-Reset. This list keeps the mixed-case spelling of each label, such as InputToken-Video. The platform sends the headers as X-Ratelimit-Inputtoken-Video-Limit. HTTP header names are case-insensitive, so the two spellings match.
The /embed-v2 and /embed-v2/tasks endpoints report every limit through this family. They do not send the aggregate X-Ratelimit-Limit, X-Ratelimit-Remaining, X-Ratelimit-Used, or X-Ratelimit-Reset headers that other endpoints send.
The maximum number of requests you can make per rate limit window for this endpoint. For details, see the Rate limits page.
The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.
The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.
The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.
The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.
The number of input tokens remaining in the current rate limit window for the type of content named in this header. A synchronous response already counts the tokens the request used. This value can be 0 on a response that succeeded. A 202 response does not yet count the task you submitted.
Response
An array of embedding results when status is ready, or null when status is processing or failed.
Metadata about the task you created. The platform sets the length of your embeddings when it creates the task, and the embedding_dimension field contains that length. Only Marengo 3.5 returns it.