TATECHATLAS
◎ English
Artificial intelligence / Guide

Indexing multimodal content: vectorization of images and text in Azure AI Search

A guide to setting up multimodal search in Azure AI Search using vectorization of text and images. Describes steps for creating an index, selecting models (e.g., CLIP), methods for generating embeddings, and executing hybrid queries to improve relevance.

On this page

To implement multimodal search in Azure AI Search, create an index with vector fields, use models like OpenAI CLIP to vectorize text and images, choose a method for embedding generation (inside or outside the indexer), and perform hybrid queries combining full-text and vector search for similarity-based retrieval.

Understanding multimodal search in Azure AI Search

Multimodal search in Azure AI Search enables finding similar content regardless of type - text or image - by converting data into numerical vectors (embeddings) where vector proximity reflects semantic similarity. For example, the query "puppy" can return dog images even without textual labels.

This approach relies on vector search, supporting not only text-to-text but also text-to-image matching, which is valuable in hybrid scenarios. Vectorization captures conceptual similarity, such as between "dog" and "canine", or between text and its corresponding image.

Preparing the index for vector data

To store vector representations, the Azure AI Search index must include fields of type Edm.Single with the vectorSearchProfile attribute. These fields hold arrays of numbers - embeddings. The vector length depends on the model used (e.g., 512 for CLIP).

from azure.search.documents.indexes import SearchIndexClient
from azure.core.credentials import AzureKeyCredential
from azure.search.documents.indexes.models import (
    SearchIndex,
    SearchField,
    SearchFieldDataType,
    VectorSearch,
    HnswVectorSearchAlgorithmConfiguration
)

# Create client
client = SearchIndexClient(endpoint="https://<service>.search.windows.net", credential=AzureKeyCredential("<key>"))

# Define fields
fields = [
    SearchField(name="id", type=SearchFieldDataType.String, key=True),
    SearchField(name="text", type=SearchFieldDataType.String),
    SearchField(name="image_vector", type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
                vector_search_dimensions=512, vector_search_profile_name="myHnswProfile"),
    SearchField(name="text_vector", type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
                vector_search_dimensions=512, vector_search_profile_name="myHnswProfile")
]

# Configure vector search
vector_search = VectorSearch(
    algorithm_configurations=[
        HnswVectorSearchAlgorithmConfiguration(name="myHnswProfile", kind="hnsw")
    ]
)

# Create index
index = SearchIndex(name="multimodal-index", fields=fields, vector_search=vector_search)
client.create_or_update_index(index)

Choosing a model for image vectorization

To convert images into vectors, use the Foundry Tools Image Retrieval Vectorize Image API or multimodal models like OpenAI CLIP. These models are trained on text-image pairs and generate a shared vector space.

For instance, sending a dog image via the API returns a 512-dimensional vector stored in the image_vector field, enabling comparison with text queries based on vector proximity.

Choosing a model for text vectorization

For text vectorization, suitable models include Azure OpenAI's text-embedding-ada-002 or SBERT. They transform text descriptions (e.g., "puppy in a meadow") into vectors of the same dimensionality as image vectors, ensuring compatibility in a shared vector space.

Embedding generation can be done externally (using OpenAI SDK) or integrated within the indexer by specifying an Azure OpenAI endpoint and key.

import openai
openai.api_type = "azure"
openai.api_key = "<key>"
openai.api_base = "https://<resource>.openai.azure.com/"
openai.api_version = "2023-05-15"

response = openai.Embedding.create(input="puppy", engine="text-embedding-ada-002")
text_embedding = response['data'][0]['embedding']

Integrated vs external vectorization

Azure AI Search supports two approaches: internal vectorization via indexer and external vectorization with precomputed embeddings. Internal vectorization requires a skillset with a built-in vectorizer and configuration of an Azure OpenAI resource.

External vectorization offers more control and suits large datasets. Embeddings are computed beforehand and uploaded with documents. This method is necessary when using models not directly supported by the indexer.

Creating a multimodal index

A multimodal index combines text, images, and their vector representations in one schema. Fields like text and image_vector allow both full-text and vector search, forming the basis for hybrid queries.

Each document includes an ID, source text or image description, and corresponding vectors. All data is indexed together, ensuring consistency during search.

Executing a multimodal search query

Search is performed via a hybrid query containing both search (for text) and vectorQueries (for vectors). For example, the text query "puppy" is vectorized, and the resulting vector is compared against image_vector in the index.

Full-text and vector search results are combined using Reciprocal Rank Fusion (RRF), improving relevance. The response provides a single ranked list of documents based on overall similarity.

import requests

url = "https://<service>.search.windows.net/indexes/multimodal-index/docs/search?api-version=2023-11-01"
headers = {"Content-Type": "application/json", "api-key": "<key>"}
payload = {
    "search": "puppy",
    "vectorQueries": [
        {
            "kind": "vector", 
            "vector": text_embedding, 
            "k": 10, 
            "fields": "image_vector"
        }
    ],
    "select": "id,text",
    "top": 5
}

response = requests.post(url, json=payload, headers=headers)
print(response.json())

Evaluating and improving search results

Search quality is evaluated using metrics like precision and recall. Test with labeled query-relevant image pairs. Precision measures the proportion of correctly retrieved results among all returned.

To improve relevance, use hybrid search with semantic ranking (queryType=semantic). Tune parameters like k and oversampling in vector queries, and test filtering strategies.

from sklearn.metrics import precision_score

# Example evaluation (hypothetical values)
y_true = [1, 0, 1, 1, 0]  # 1 = relevant, 0 = not
y_pred = [1, 1, 1, 0, 0]  # predicted results
precision = precision_score(y_true, y_pred, average='binary')
print(f"Precision: {precision:.2f}")

Things to check

  • Index created with vector fields
  • Vector search profile (HNSW) specified
  • Text and image embeddings have the same dimensionality
  • Hybrid query includes search and vectorQueries
  • Results merged via RRF

Vector search functionality is unavailable for services created before January 1, 2019. Embedding generation may incur costs when using Azure OpenAI or Foundry API.

Sources

  1. Microsoft Learn: vector search ↗
  2. Microsoft Learn: hybrid search ↗
  3. scikit-learn: precision_score ↗
Back to top ↑