> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.twelvelabs.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.twelvelabs.io/_mcp/server.

# Create sync embeddings

> Create embeddings synchronously.

The `EmbedClient.V2Client` class provides methods to create embeddings synchronously for multimodal content, returning embeddings immediately in the response.

# Methods

## Create sync embeddings

**Description**: This method synchronously creates embeddings for multimodal content and returns the results immediately in the response.

Use this method to embed a query for retrieving matching content. With Marengo 3.5, audio and video can be up to 30 seconds. With Marengo 3.0, they can be up to 10 minutes. For longer content, use the [`embed.v_2.tasks.create`](/v1.3/sdk-reference/python/create-embeddings-v-2/create-async-embeddings#create-an-async-embedding-task) method instead.

The content this method accepts depends on the model. With Marengo 3.5, this method accepts only the `multi_input` input type; provide text, images, audio, or video as media sources. With Marengo 3.0, use the individual input types. For the formats, resolutions, file sizes, and duration limits each model accepts, see the input requirements for [Marengo 3.5](/v1.3/docs/concepts/models/marengo/marengo-3-5#input-requirements) or [Marengo 3.0](/v1.3/docs/concepts/models/marengo/marengo-3-0#input-requirements).

> **Note**
>
> This method is rate-limited. With Marengo 3.5, the platform counts input tokens for each type of content. A request can exceed a limit before you see an error. For details, see [Input token limits for embedding](/v1.3/docs/get-started/rate-limits#input-token-limits-for-embedding).

**Function signature and example**

**`Function signature`**

```python Function signature
def create(
    self,
    *,
    input_type: CreateEmbeddingsRequestInputType,
    model_name: CreateEmbeddingsRequestModelName,
    auto_truncate: typing.Optional[bool] = OMIT,
    embedding_uncertainty: typing.Optional[bool] = OMIT,
    embedding_dimension: typing.Optional[int] = OMIT,
    text: typing.Optional[TextInputRequest] = OMIT,
    image: typing.Optional[ImageInputRequest] = OMIT,
    text_image: typing.Optional[TextImageInputRequest] = OMIT,
    audio: typing.Optional[AudioInputRequest] = OMIT,
    video: typing.Optional[VideoInputRequest] = OMIT,
    multi_input: typing.Optional[MultiInputRequest] = OMIT,
    request_options: typing.Optional[RequestOptions] = None,
) -> EmbeddingSuccessResponse
```

**`Create text embeddings`**

```python Create text embeddings
from twelvelabs import TwelveLabs, MultiInputRequest

response = client.embed.v_2.create(
    input_type="multi_input",
    model_name="marengo3.5",
    multi_input=MultiInputRequest(
        input_text="<YOUR_TEXT>",
    ),
)

print(f"Number of embeddings: {len(response.data)}")
for embedding_data in response.data:
    print(f"Embedding dimensions: {len(embedding_data.embedding)}")
    print(f"First 10 values: {embedding_data.embedding[:10]}")
```

**`Create image embeddings`**

```python Create image embeddings
from twelvelabs import TwelveLabs, MultiInputRequest, MultiInputMediaSource

response = client.embed.v_2.create(
    input_type="multi_input",
    model_name="marengo3.5",
    multi_input=MultiInputRequest(
        media_sources=[
            MultiInputMediaSource(
                media_type="image",
                url="<YOUR_IMAGE_URL>",
            ),
        ],
    ),
)

print(f"Number of embeddings: {len(response.data)}")
for embedding_data in response.data:
    print(f"Embedding dimensions: {len(embedding_data.embedding)}")
    print(f"First 10 values: {embedding_data.embedding[:10]}")
```

**`Create video embeddings`**

```python Create video embeddings
from twelvelabs import TwelveLabs, MultiInputRequest, MultiInputMediaSource

response = client.embed.v_2.create(
    input_type="multi_input",
    model_name="marengo3.5",
    multi_input=MultiInputRequest(
        media_sources=[
            MultiInputMediaSource(
                media_type="video",
                url="<YOUR_VIDEO_URL>",
            ),
        ],
    ),
)

print(f"Number of embeddings: {len(response.data)}")
for embedding_data in response.data:
    print(f"Embedding dimensions: {len(embedding_data.embedding)}")
    print(f"First 10 values: {embedding_data.embedding[:10]}")
```

**`Create embeddings from multiple images`**

```python Create embeddings from multiple images
from twelvelabs import TwelveLabs, MultiInputRequest, MultiInputMediaSource

response = client.embed.v_2.create(
    input_type="multi_input",
    model_name="marengo3.5",
    multi_input=MultiInputRequest(
        media_sources=[
            MultiInputMediaSource(
                media_type="image",
                url="<YOUR_IMAGE_URL_1>",
            ),
            MultiInputMediaSource(
                media_type="image",
                url="<YOUR_IMAGE_URL_2>",
            ),
        ],
    ),
)

print(f"Number of embeddings: {len(response.data)}")
for embedding_data in response.data:
    print(f"Embedding dimensions: {len(embedding_data.embedding)}")
    print(f"First 10 values: {embedding_data.embedding[:10]}")
```

**`Create embeddings from images and text`**

```python Create embeddings from images and text
from twelvelabs import TwelveLabs, MultiInputRequest, MultiInputMediaSource

response = client.embed.v_2.create(
    input_type="multi_input",
    model_name="marengo3.5",
    multi_input=MultiInputRequest(
        input_text="<@image1> and <@image2> show contrasting scenes",
        media_sources=[
            MultiInputMediaSource(
                name="image1",
                media_type="image",
                url="<YOUR_IMAGE_URL_1>",
            ),
            MultiInputMediaSource(
                name="image2",
                media_type="image",
                url="<YOUR_IMAGE_URL_2>",
            ),
        ],
    ),
)

print(f"Number of embeddings: {len(response.data)}")
for embedding_data in response.data:
    print(f"Embedding dimensions: {len(embedding_data.embedding)}")
    print(f"First 10 values: {embedding_data.embedding[:10]}")
```

### Parameters

| Name                    | Type                               | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ----------------------- | ---------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `input_type`            | `CreateEmbeddingsRequestInputType` | Yes      | The type of content for the embeddings. Values: - `multi_input`: Text and up to 10 media sources, combined into a single embedding. To reference a specific media source from your text, use a placeholder in the format `&lt;@name&gt;`, where `name` matches the `name` field of a media source. Marengo 3.5 accepts images, video, audio, and documents as media sources. Marengo 3.0 accepts images. - `audio`: An audio file. Requires Marengo 3.0. - `video`: A video file. Requires Marengo 3.0. - `image`: An image file. Requires Marengo 3.0. - `text`: Text input. Requires Marengo 3.0. - `text_image`: Text and an image. Requires Marengo 3.0.                                                                                                                                                                                                             |
| `model_name`            | `CreateEmbeddingsRequestModelName` | Yes      | The embedding model to use. Values: - `marengo3.5`: For details about this version, see the [Marengo 3.5](/v1.3/docs/concepts/models/marengo/marengo-3-5) page. - `marengo3.0`: For details about this version, see the [Marengo 3.0](/v1.3/docs/concepts/models/marengo/marengo-3-0) page.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| `text`                  | `TextInputRequest`                 | No       | Text input configuration. Required when `input_type` is `text`. See [TextInputRequest](#textinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| `image`                 | `ImageInputRequest`                | No       | Image input configuration. Required when `input_type` is `image`. See [ImageInputRequest](#imageinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| `text_image`            | `TextImageInputRequest`            | No       | Combined text and image input configuration. Required when `input_type` is `text_image`. See [TextImageInputRequest](#textimageinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `audio`                 | `AudioInputRequest`                | No       | Audio input configuration. Required when `input_type` is `audio`. See [AudioInputRequest](#audioinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| `video`                 | `VideoInputRequest`                | No       | Video input configuration. Required when `input_type` is `video`. See [VideoInputRequest](#videoinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| `multi_input`           | `MultiInputRequest`                | No       | Text and media source configuration. Required when `input_type` is `multi_input`. See [MultiInputRequest](#multiinputrequest) for details.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| `auto_truncate`         | `Optional[bool]`                   | No       | Controls the behavior of the platform when the text in your request exceeds 2,000 tokens. Requires Marengo 3.5. Values: - `False`: The platform returns a `400` error. - `True`: Truncate your text to fit the limit, and set the `usage.truncated` field to `True` in the response.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `embedding_uncertainty` | `Optional[bool]`                   | No       | Set this parameter to `True` to include a per-dimension uncertainty vector in the `data[].embedding_uncertainty` field of the response. The vector has the same length as the `embedding` array. A higher value indicates lower confidence in that dimension. Requires Marengo 3.5. **Requirements**: - Set this parameter to `True` only for a text-only or media-only request. If you combine text with media sources, the platform returns a `400` error. - The platform returns a `400` error if your request includes a document, whether PDF, plain text, or Markdown.                                                                                                                                                                                                                                                                                             |
| `embedding_dimension`   | `Optional[int]`                    | No       | The number of dimensions for each embedding in the response, including the [`data[].embedding_uncertainty`](/v1.3/api-reference/create-embeddings-v2/create-embeddings#response.body.data.embedding-uncertainty) vector. Marengo 3.5 produces Matryoshka embeddings: a shorter embedding consists of the first values of the full-length embedding. A 256-dimension embedding, for example, is the first 256 values of a 512-dimension embedding of the same content. Shorter embeddings reduce index size and speed up similarity search; longer embeddings produce higher retrieval quality. **Requirements**: - Requires Marengo 3.5. Setting this parameter when `model_name` is `marengo3.0` returns a `400` error. - Applies to the entire request: you cannot set it for a single input type or embedding. - Use the same value across an index. **Default**: 512 |
| `request_options`       | `RequestOptions`                   | No       | Request-specific configuration.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |

#### TextInputRequest

The `TextInputRequest` class specifies the configuration for processing text content. Required when `input_type` is `text`.

| Name         | Type  | Required | Description                                                                           |
| ------------ | ----- | -------- | ------------------------------------------------------------------------------------- |
| `input_text` | `str` | Yes      | The text for which you wish to create an embedding. The maximum length is 500 tokens. |

#### ImageInputRequest

The `ImageInputRequest` class specifies the configuration for processing image content. Required when `input_type` is `image`. Requires Marengo 3.0. The decoded file can be up to 32 MB.

| Name           | Type          | Required | Description                                                                          |
| -------------- | ------------- | -------- | ------------------------------------------------------------------------------------ |
| `media_source` | `MediaSource` | Yes      | Specifies the source of the image file. See [MediaSource](#mediasource) for details. |

#### TextImageInputRequest

The `TextImageInputRequest` class specifies the configuration for processing combined text and image content. Required when `input_type` is `text_image`. Requires Marengo 3.0. The decoded file can be up to 32 MB.

| Name           | Type          | Required | Description                                                                           |
| -------------- | ------------- | -------- | ------------------------------------------------------------------------------------- |
| `media_source` | `MediaSource` | Yes      | Specifies the source of the image file. See [MediaSource](#mediasource) for details.  |
| `input_text`   | `str`         | Yes      | The text for which you wish to create an embedding. The maximum length is 500 tokens. |

#### AudioInputRequest

The `AudioInputRequest` class specifies the configuration for processing audio content. Required when `input_type` is `audio`. Requires Marengo 3.0. The decoded file can be up to 36 MB.

| Name               | Type                | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| ------------------ | ------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `media_source`     | `MediaSource`       | Yes      | Specifies the source of the audio file. See [MediaSource](#mediasource) for details.                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `start_sec`        | `float`             | No       | The start time in seconds for processing the audio file. Use this parameter to process a portion of the audio file starting from a specific time. **Default**: 0 (start from the beginning).                                                                                                                                                                                                                                                                                                              |
| `end_sec`          | `float`             | No       | The end time in seconds for processing the audio file. Use this parameter to process a portion of the audio file ending at a specific time. The end time must be greater than the start time. **Default**: End of the audio file                                                                                                                                                                                                                                                                          |
| `segmentation`     | `AudioSegmentation` | No       | Specifies how the platform divides the audio into segments. See [AudioSegmentation](#audiosegmentation) for details.                                                                                                                                                                                                                                                                                                                                                                                      |
| `embedding_option` | `List[str]`         | No       | The types of embeddings you wish to generate. **Values**: - `audio`: Generates embeddings based on audio content (sounds, music, effects) - `transcription`: Generates embeddings based on transcribed speech You can specify multiple values to generate different types of embeddings for the same audio. **Default**: `["audio", "transcription"]`                                                                                                                                                     |
| `embedding_scope`  | `List[str]`         | No       | The scope for which you wish to generate embeddings. **Values**: - `clip`: Generates one embedding for each segment - `asset`: Generates one embedding for the entire audio file You can specify multiple scopes to generate embeddings at different levels. **Default**: `["clip", "asset"]`                                                                                                                                                                                                             |
| `embedding_type`   | `List[str]`         | No       | Specifies how to structure the embedding. Include this parameter only when the `embedding_option` parameter contains at least two values. **Values**: - `separate_embedding`: Returns separate embeddings for each modality specified in the `embedding_option` parameter. - `fused_embedding`: Returns a single combined embedding that integrates all modalities into one vector. Specify both values to receive separate and fused embeddings in the same response. **Default**: `separate_embedding`. |

#### VideoInputRequest

The `VideoInputRequest` class specifies the configuration for processing video content. Required when `input_type` is `video`. Requires Marengo 3.0. The decoded file can be up to 36 MB.

| Name               | Type                | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| ------------------ | ------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `media_source`     | `MediaSource`       | Yes      | Specifies the source of the video file. See [MediaSource](#mediasource) for details.                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `start_sec`        | `float`             | No       | The start time in seconds for processing the video file. Use this parameter to process a portion of the video file starting from a specific time. **Default**: 0 (start from the beginning)                                                                                                                                                                                                                                                                                                               |
| `end_sec`          | `float`             | No       | The end time in seconds for processing the video file. Use this parameter to process a portion of the video file ending at a specific time. The end time must be greater than the start time. **Default**: End of the video file                                                                                                                                                                                                                                                                          |
| `segmentation`     | `VideoSegmentation` | No       | Specifies how the platform divides the video into segments. See [VideoSegmentation](#videosegmentation) for details.                                                                                                                                                                                                                                                                                                                                                                                      |
| `embedding_option` | `List[str]`         | No       | The types of embeddings to generate for the video. **Values**: - `visual`: Generates embeddings based on visual content (scenes, objects, actions) - `audio`: Generates embeddings based on audio content (sounds, music, effects) - `transcription`: Generates embeddings based on transcribed speech You can specify multiple values to generate different types of embeddings for the same video. **Default**: `["visual", "audio", "transcription"]`                                                  |
| `embedding_scope`  | `List[str]`         | No       | The scope for which you wish to generate embeddings. **Values**: - `clip`: Generates one embedding for each segment - `asset`: Generates one embedding for the entire video file. Use this scope for videos up to 10-30 seconds to maintain optimal performance. You can specify multiple scopes to generate embeddings at different levels. **Default**: `["clip", "asset"]`                                                                                                                             |
| `embedding_type`   | `List[str]`         | No       | Specifies how to structure the embedding. Include this parameter only when the `embedding_option` parameter contains at least two values. **Values**: - `separate_embedding`: Returns separate embeddings for each modality specified in the `embedding_option` parameter. - `fused_embedding`: Returns a single combined embedding that integrates all modalities into one vector. Specify both values to receive separate and fused embeddings in the same response. **Default**: `separate_embedding`. |

#### MultiInputRequest

The `MultiInputRequest` class specifies the configuration for processing text and media sources. Required when `input_type` is `multi_input`.

Marengo 3.5 accepts images, video, audio, and documents as media sources. Marengo 3.0 accepts images.

Include text in the `input_text` field when you combine media sources of different types. For example, if you combine an image and a video without text, the platform returns a `400` error. Media sources of the same type do not require text.

**Document sources**

The platform embeds a plain text or Markdown document as text. Combining content into a single embedding requires at least one image, video, or audio source; if the content is all text, the platform returns a `400` error. A plain text or Markdown document combined with the `input_text` field, and two plain text documents, are both all-text content. Send a single text document as your only source, use `input_text` on its own, or add an image, video, or audio source.

A PDF document must be your only source. You cannot combine it with any other media source, including another document, or with the `input_text` field; the platform returns a `400` error if you do. To combine document content with an image, video, or audio source, send a plain text or Markdown document instead of a PDF file.

| Name            | Type                          | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| --------------- | ----------------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `input_text`    | `str`                         | No       | Text to include in the embedding. With Marengo 3.5, the text cannot exceed 2,000 tokens. Use the `auto_truncate` parameter to control the behavior of the platform when your text exceeds it. **Usage options**: - Provide text without media sources to create a text-only embedding. - Combine text with media sources to add context. Example: "A person cooking." - Use media source references to describe relationships between specific media sources. The format is `&lt;@name&gt;`, where `name` matches the `name` field of a media source. Example: "A person wearing \<@outfit> and holding \<@accessory>." - Omit this field to create an embedding from media sources only. |
| `media_sources` | `List[MultiInputMediaSource]` | No       | An array of up to 10 media sources to include in the embedding. Omit it to create a text-only embedding from the `input_text` field. The platform processes media sources in the order they appear in the array. If you use media source references in the `input_text` parameter, each must have a corresponding media source with a matching `name` field. If a reference has no match, the request fails.                                                                                                                                                                                                                                                                              |

#### MediaSource

The `MediaSource` class specifies the source of the media file. Provide exactly one of the following:

| Name             | Type  | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| ---------------- | ----- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `base_64_string` | `str` | No       | The base64-encoded media data. Encoding grows the payload by about a third, so the string you send is larger than the original file. The maximum size depends on the input type and the model. The description of the field that contains this media source states the limit where it differs; for the formats and sizes each model accepts, see the input requirements for [Marengo 3.5](/v1.3/docs/concepts/models/marengo/marengo-3-5#input-requirements) or [Marengo 3.0](/v1.3/docs/concepts/models/marengo/marengo-3-0#input-requirements). |
| `url`            | `str` | No       | The publicly accessible URL of the media file. Use direct links to raw media files. Video hosting platforms and cloud storage sharing links are not supported.                                                                                                                                                                                                                                                                                                                                                                                    |
| `asset_id`       | `str` | No       | The unique identifier of an asset from a [direct](/v1.3/sdk-reference/python/upload-content/direct-uploads) or [multipart](/v1.3/sdk-reference/python/upload-content/multipart-uploads) upload. The asset status must be `ready`. Use [`assets.retrieve`](/v1.3/sdk-reference/python/manage-assets#retrieve-an-asset) to check the status.                                                                                                                                                                                                        |

#### MultiInputMediaSource

A class specifying a media source for multi-input embeddings. You must provide exactly one of the `url`, `base_64_string`, or `asset_id` fields. With Marengo 3.5, each media source can be up to 32 MB, whichever of the three fields you use. Audio and video can be up to 30 seconds. If your content exceeds either limit, the platform returns either a `400` or a `413` error. A PDF file also has a page allowance: 16 pages for each MB of file size. A 0.5 MB file is allowed 16 pages, and a 4 MB file is allowed 64 pages. The platform checks the page count of the file against the allowance before processing the file. If the file exceeds the allowance, the platform returns a `413` error. The error message includes the page count and the allowance. Plain text and Markdown files have no page allowance.

| Name             | Type                             | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                    |
| ---------------- | -------------------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `name`           | `str`                            | No       | The unique identifier for this media source. This field is required when `input_text` references this media source.                                                                                                                                                                                                                                                                                            |
| `media_type`     | `MultiInputMediaSourceMediaType` | No       | The type of media. **Values**: - `image`: An image file. Works with both Marengo 3.0 and Marengo 3.5. - `video`: A video file. Requires Marengo 3.5. - `audio`: An audio file. Requires Marengo 3.5. - `document`: A PDF (`.pdf`), plain text (`.txt`), or Markdown (`.md`) file. Requires Marengo 3.5. For the rules on combining a document with other sources, see [MultiInputRequest](#multiinputrequest). |
| `url`            | `str`                            | No       | The publicly accessible URL of the media file. Use direct links to raw files. Media hosting platforms and cloud storage sharing links are not supported.                                                                                                                                                                                                                                                       |
| `base_64_string` | `str`                            | No       | The base64-encoded media data.                                                                                                                                                                                                                                                                                                                                                                                 |
| `asset_id`       | `str`                            | No       | The unique identifier of an asset from a [direct](/v1.3/sdk-reference/python/upload-content/direct-uploads) or [multipart](/v1.3/sdk-reference/python/upload-content/multipart-uploads) upload.                                                                                                                                                                                                                |

#### AudioSegmentation

The `AudioSegmentation` class specifies how the platform divides the audio into segments using fixed-length intervals.

| Name       | Type                     | Required | Description                                                                                                                                                            |
| ---------- | ------------------------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `strategy` | `Literal["fixed"]`       | Yes      | The segmentation strategy. Value: `fixed`.                                                                                                                             |
| `fixed`    | `AudioSegmentationFixed` | Yes      | Configuration for fixed segmentation. This object is required when the `strategy` field is `fixed`. See [AudioSegmentationFixed](#audiosegmentationfixed) for details. |

#### AudioSegmentationFixed

The `AudioSegmentationFixed` class configures fixed-length segmentation for audio.

| Name           | Type  | Required | Description                                                                                                                                                                                                                                                                                                                            |
| -------------- | ----- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `duration_sec` | `int` | Yes      | The duration in seconds for each segment. The platform divides the audio into segments of this exact length. The final segment may be shorter if the audio duration is not evenly divisible. **Min**: `2`. **Max**: `10`. **Example**: With `duration_sec: 5`, a 12-second audio file produces segments: \[0-5s], \[5-10s], \[10-12s]. |

#### VideoSegmentation

The `VideoSegmentation` type specifies how the platform divides the video into segments. Use one of the following:

**Fixed segmentation**: Divides the video into equal-length segments:

| Name       | Type                          | Required | Description                                                                                                        |
| ---------- | ----------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------ |
| `strategy` | `Literal["fixed"]`            | Yes      | The segmentation strategy. Value: `fixed`.                                                                         |
| `fixed`    | `VideoSegmentationFixedFixed` | Yes      | Configuration for fixed segmentation. See [VideoSegmentationFixedFixed](#videosegmentationfixedfixed) for details. |

**Dynamic segmentation**: Divides the video into adaptive segments based on scene changes:

| Name       | Type                              | Required | Description                                                                                                                  |
| ---------- | --------------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `strategy` | `Literal["dynamic"]`              | Yes      | The segmentation strategy. Value: `dynamic`.                                                                                 |
| `dynamic`  | `VideoSegmentationDynamicDynamic` | Yes      | Configuration for dynamic segmentation. See [VideoSegmentationDynamicDynamic](#videosegmentationdynamicdynamic) for details. |

#### VideoSegmentationFixedFixed

The `VideoSegmentationFixedFixed` class configures fixed-length segmentation for video.

| Name           | Type  | Required | Description                                                                                                                                                                                                                                                                                                                       |
| -------------- | ----- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `duration_sec` | `int` | Yes      | The duration in seconds for each segment. The platform divides the video into segments of this exact length. The final segment may be shorter if the video duration is not evenly divisible. **Min**: `2`. **Max**: `10`. **Example**: With `duration_sec: 5`, a 12-second video produces segments: \[0-5s], \[5-10s], \[10-12s]. |

#### VideoSegmentationDynamicDynamic

The `VideoSegmentationDynamicDynamic` class configures dynamic segmentation for video based on scene changes.

| Name               | Type  | Required | Description                                                                                                                                                                                                                                                                                                                                         |
| ------------------ | ----- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `min_duration_sec` | `int` | Yes      | The minimum duration in seconds for each segment. The platform divides the video into segments that are at least this long. Segments adapt to scene changes and content boundaries and may be longer than the minimum. **Min**: `2`. **Max**: `5`. **Example**: With `min_duration_sec: 3`, segments might be: \[0-3.2s], \[3.2-7.8s], \[7.8-12.1s] |

### Return value

Returns an `EmbeddingSuccessResponse` object containing the embedding results.

The `EmbeddingSuccessResponse` class contains the following properties:

| Name       | Type                               | Description                                                                                                                                                                                |
| ---------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `data`     | `List[EmbeddingData]`              | Array of embedding results.                                                                                                                                                                |
| `usage`    | `Optional[EmbeddingUsage]`         | Token counts for the request. Only Marengo 3.5 returns this field. See [EmbeddingUsage](#embeddingusage) for details.                                                                      |
| `metadata` | `Optional[EmbeddingMediaMetadata]` | Metadata for the media input. Available for the `image`, `text_image`, `audio`, `video`, and `multi_input` input types. See [EmbeddingMediaMetadata](#embeddingmediametadata) for details. |

The `EmbeddingData` class contains the following properties:

| Name                    | Type                                     | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ----------------------- | ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `embedding`             | `List[float]`                            | The embedding vector for the content.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `embedding_uncertainty` | `Optional[List[float]]`                  | A per-dimension uncertainty vector with the same length as the `embedding` array. A higher value indicates lower confidence in that dimension. Present when the request sets `embedding_uncertainty: true`. Only Marengo 3.5 returns this field.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `embedding_option`      | `Optional[EmbeddingDataEmbeddingOption]` | The type of the embedding. **Values**: - `visual`: Embedding based on visual content (a video, a page of a PDF file, or an image embedded asynchronously). - `audio`: Embedding based on audio content. - `transcription`: Embedding based on transcribed speech. Returned only for content embedded with Marengo 3.0. - `text`: Embedding based on the text content of a PDF, plain text, or Markdown file embedded asynchronously. - `fused`: Embedding based on a combination of the modalities specified in the request. The platform returns this embedding only for video and audio input, and only when the `embedding_type` parameter includes the `fused_embedding` value. - `null`: For text embeddings and images embedded synchronously.                                                                                                                                                                                     |
| `embedding_scope`       | `Optional[EmbeddingDataEmbeddingScope]`  | The scope for which the embedding was generated. **Values**: - `clip`: Embedding for a segment. For video and audio input, one embedding per detected segment. - `page`: Embedding for one page of a PDF file embedded asynchronously, or for one quadrant of a page when the request sets `document.segmentation.spatial.strategy` to `quadrants`. With that strategy, five entries share the same scope and page numbers, so read the `quadrant` field to tell them apart: the whole-page entry has no `quadrant` value. - `chunk`: Embedding for one chunk of whole sentences of a plain text or Markdown file embedded asynchronously. Read the `chunk_index` field for the position of the chunk in the file. - `asset`: Embedding for the entire file. For video and audio input, use this scope for content up to 10-30 seconds to maintain optimal performance. - `null`: For text embeddings and images embedded synchronously. |
| `start_sec`             | `Optional[float]`                        | The start time in seconds for this segment. This field is `null` for text and image embeddings.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| `end_sec`               | `Optional[float]`                        | The end time in seconds for this segment. This field is `null` for text and image embeddings.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `start_page_number`     | `Optional[int]`                          | The first page this embedding covers, counting from 1. The platform returns this field only for page-level embeddings of a PDF file, and `null` in every other case.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `quadrant`              | `Optional[str]`                          | The quarter of the page this embedding covers. The platform returns this field only when the request sets `document.segmentation.spatial.strategy` to `quadrants`, and only on the four quadrant embeddings of a page. This field is `null` on the whole-page embedding and in every other case.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `chunk_index`           | `Optional[int]`                          | The position of this chunk in the file, counting from 0. The platform returns this field only on `chunk`-scope embeddings of a plain text or Markdown file, and `null` in every other case. Read this field rather than the position of the entry in the `data` array, which provides no ordering guarantee.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| `end_page_number`       | `Optional[int]`                          | The last page this embedding covers, counting from 1 and including that page. This field matches the `start_page_number` field when the embedding covers a single page. The platform returns this field only for page-level embeddings of a PDF file, and `null` in every other case.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |

#### EmbeddingUsage

The `EmbeddingUsage` class provides token counts for the request. Only Marengo 3.5 returns this object.

| Name           | Type             | Description                                                                                                                                                                                                                                                                                                                                                              |
| -------------- | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `input_tokens` | `Dict[str, int]` | The number of tokens the request used. Each key names a type of content the request processed, and each value is the token count for that content. The platform reports only the content types your request actually used. The possible keys are `video`, `audio`, `image`, `document`, and `text`. Read the keys the response returns rather than assuming a fixed set. |
| `truncated`    | `bool`           | Whether the input was truncated to fit within the token limit.                                                                                                                                                                                                                                                                                                           |

#### EmbeddingMediaMetadata

The `EmbeddingMediaMetadata` type provides metadata for the media input. Available for the `image`, `text_image`, `audio`, `video`, and `multi_input` input types. The `input_type` field selects one variant:

**Image**: Metadata for image embeddings.

| Name             | Type               | Description                                     |
| ---------------- | ------------------ | ----------------------------------------------- |
| `input_type`     | `Literal["image"]` | The type of the input content. Value: `image`.  |
| `input_url`      | `Optional[str]`    | The publicly accessible URL for the image file. |
| `input_filename` | `Optional[str]`    | The name of the image file.                     |

**Text and image**: Metadata for text-image embeddings.

| Name             | Type                    | Description                                         |
| ---------------- | ----------------------- | --------------------------------------------------- |
| `input_type`     | `Literal["text_image"]` | The type of the input content. Value: `text_image`. |
| `input_url`      | `Optional[str]`         | The publicly accessible URL for the image file.     |
| `input_filename` | `Optional[str]`         | The name of the image file.                         |

**Audio**: Metadata for audio embeddings.

| Name                  | Type                                              | Description                                                                                        |
| --------------------- | ------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `input_type`          | `Literal["audio"]`                                | The type of the input content. Value: `audio`.                                                     |
| `input_url`           | `Optional[str]`                                   | The publicly accessible URL for the audio file.                                                    |
| `input_filename`      | `Optional[str]`                                   | The name of the audio file.                                                                        |
| `embedding_options`   | `List[str]`                                       | The `embedding_option` values used to generate the embedding.                                      |
| `embedding_scopes`    | `List[EmbeddingAudioMetadataEmbeddingScopesItem]` | The `embedding_scope` values used to generate the embedding.                                       |
| `embedding_dimension` | `Optional[int]`                                   | The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field. |
| `duration`            | `float`                                           | The duration of the audio in seconds.                                                              |
| `start_offset_sec`    | `Optional[float]`                                 | The start offset in seconds.                                                                       |
| `end_offset_sec`      | `Optional[float]`                                 | The end offset in seconds.                                                                         |

**Video**: Metadata for video embeddings.

| Name                  | Type                                              | Description                                                                                        |
| --------------------- | ------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `input_type`          | `Literal["video"]`                                | The type of the input content. Value: `video`.                                                     |
| `input_url`           | `Optional[str]`                                   | The publicly accessible URL for the video file.                                                    |
| `input_filename`      | `Optional[str]`                                   | The name of the video file.                                                                        |
| `clip_length`         | `Optional[int]`                                   | Length of each video clip in seconds. Only available for fixed segmentation.                       |
| `embedding_scopes`    | `List[EmbeddingVideoMetadataEmbeddingScopesItem]` | The `embedding_scope` values used to generate the embedding.                                       |
| `embedding_dimension` | `Optional[int]`                                   | The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field. |
| `embedding_options`   | `List[str]`                                       | The `embedding_option` values used to generate the embedding.                                      |
| `duration`            | `float`                                           | The duration of the video in seconds.                                                              |
| `start_offset_sec`    | `Optional[float]`                                 | The start offset in seconds.                                                                       |
| `end_offset_sec`      | `Optional[float]`                                 | The end offset in seconds.                                                                         |

**Multi-input**: Metadata for multi-input embeddings.

| Name                  | Type                     | Description                                                                                        |
| --------------------- | ------------------------ | -------------------------------------------------------------------------------------------------- |
| `input_type`          | `Literal["multi_input"]` | The type of the input content. Value: `multi_input`.                                               |
| `embedding_dimension` | `Optional[int]`          | The number of dimensions for each embedding in this response. Only Marengo 3.5 returns this field. |

### API Reference

[Create sync embeddings](/v1.3/api-reference/create-embeddings-v2/create-embeddings)

### Related guide

* [Embed a query](/v1.3/docs/guides/create-embeddings/query)