TATECHATLAS
◎ English
Artificial intelligence

Configuring Hybrid Search with Vector and Text Fields in Azure AI Search

Step-by-step guide to creating an index, generating embeddings, and executing hybrid queries that combine full-text and vector search using Azure AI Search.

On this page

Hybrid search is configured by adding vector fields to the index schema, generating embeddings for textual content, and forming a query with search and vectorQueries parameters. Results are merged using Reciprocal Rank Fusion (RRF), and optionally a semantic reranker can be enabled for re-ranking. Key steps: index schema definition, embedding generation, k and oversampling configuration, query construction with filters and facets, optional semantic reranker inclusion, testing, and monitoring quotas and costs.

1. Index Schema Definition with Vector and Text Fields

To begin, an Azure AI Search index is created containing both regular text fields for full-text search and vector fields for semantic search. Vector fields store numeric embeddings derived from document text. The index must have at least one searchable text field and one vector field. When creating an index via the portal or REST API, vector fields are declared with type "Edm.Single" and the "searchable": true attribute, but they are not used for ordinary full-text search; their role is vector similarity.

An example from Microsoft Learn documentation shows a hybrid query where the search parameter is combined with vectorQueries pointing to DescriptionVector and Description_frVector fields. This allows a single query to perform both keyword search and semantic search, with results merged by the RRF algorithm.

Important: vector fields cannot be used directly in filters. For metadata (e.g., category, geography) a separate text or numeric field marked with filterable and/or facetable attributes is required.

POST https://my-service.search.windows.net/indexes/hotels-vector-quickstart/docs/search?api-version=2026-04-01\ncontent-type: application/JSON\n{\n  "count": true,\n  "search": "historic hotel walk to restaurants and shopping",\n  "select": "HotelId, HotelName, Category, Description, Address/City, Address/StateProvince",\n  "filter": "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",\n  "vectorFilterMode": "postFilter",\n  "facets": ["Address/StateProvince"],\n  "vectorQueries": [{\n    "kind": "vector",\n    "vector": [0.1,0.2,…],\n    "k": 50,\n    "fields": "DescriptionVector",\n    "exhaustive": true,\n    "oversampling": 20\n  }],\n  "skip": 0,\n  "top": 10,\n  "queryType": "semantic",\n  "queryLanguage": "en-us",\n  "semanticConfiguration": "my-semantic-config"\n}

2. Generating or Importing Embeddings for Text Fields

Embeddings can be generated in two ways: using Azure AI Search built-in capabilities (indexer pipeline with Azure OpenAI) or external models (OpenAI embeddings, SBERT, etc.) and then loading vectors directly into the index. In the first case, a skillset is configured that automatically chunks text and calls the embedding model. In the second case, vectors are formed externally and passed via API when creating or updating documents.

Documentation recommends using Azure OpenAI for embedding generation, specifically the text-embedding-ada-002 model. Embeddings must have the same dimensionality (e.g., 1536 for Ada), and this value is specified when creating the vector field.

When importing data via the "Import data" wizard, an option for vectorization can be selected, and the service will automatically generate embeddings for the loaded content if a suitable data source and skillset are configured.

/* Example call to Azure OpenAI for obtaining an embedding */\nPOST https://my-openai-resource.openai.azure.com/openai/deployments/text-embedding-ada-002/embeddings?api-version=2024-02-15\nHeaders: api-key: YOUR_API_KEY\n{\n  "input": "historic hotel walk to restaurants and shopping"\n}

3. Configuring Vector Search Parameters (k, oversampling, exhaustive)

The k parameter in vectorQueries specifies how many nearest neighbors will be returned for each vector. It is recommended to set k >= 50 if a semantic reranker is used, so it has enough candidates for re-ranking.

The oversampling parameter specifies an additional percentage of candidates to be extracted beyond k, which helps improve result quality when using HNSW indexes. Typical values from 10 to 50 give a good balance between latency and accuracy.

The exhaustive parameter indicates whether to use a full scan for finding nearest neighbors. Setting exhaustive: true guarantees finding the true k nearest neighbors, but increases latency. By default false, and for most scenarios HNSW is sufficient.

In the example from the hybrid-search-overview document, two vectorQueries are shown: one with exhaustive=true and oversampling=20, another with exhaustive=false and oversampling=10. Configuration depends on index size and relevancy requirements.

Note: increasing k and oversampling increases accuracy but also query latency. Always test values on a representative data sample.

"vectorQueries": [{\n    "kind": "vector",\n    "vector": <array> ,\n    "k": 50,\n    "fields": "DescriptionVector",\n    "exhaustive": true,\n    "oversampling": 20\n }]

4. Forming a Hybrid Query with search and vectorQueries

A hybrid query combines the search parameter (full-text query) and one or more vectorQueries. The request is sent as a single HTTP POST message to the index endpoint with api-version=2026-04-01 (or current version).

The server executes full-text and vector searches in parallel, then merges results using the Reciprocal Rank Fusion (RRF) algorithm. RRF assigns each document a combined rank based on its position in both result lists, allowing combination of precise keyword search and semantic similarity advantages.

In the query, queryType: semantic can be specified to enable the semantic reranker, which applies machine reading to RRF results and re-ranks them based on query context. This is particularly useful for conceptual queries where meaning matters more than word matching.

Filters and facets are applied to fields other than vector fields. For example, geospatial filter geo.distance or category filters work on the merged result after RRF. It is important to test vectorFilterMode behavior (preFilter vs postFilter), as the order of filter application affects performance and returned document set.

A complete query example is provided in section 1 and includes facets, filter, vectorQueries and semanticConfiguration.

POST https://my-service.search.windows.net/indexes/hotels-vector-quickstart/docs/search?api-version=2026-04-01\ncontent-type: application/JSON\n{\n  "count": true,\n  "search": "historic hotel walk to restaurants and shopping",\n  "select": "HotelId, HotelName, Category, Description, Address/City, Address/StateProvince",\n  "filter": "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",\n  "vectorFilterMode": "postFilter",\n  "facets": ["Address/StateProvince"],\n  "vectorQueries": [{\n    "kind": "vector",\n    "vector": [0.1,0.2,…],\n    "k": 50,\n    "fields": "DescriptionVector",\n    "exhaustive": true,\n    "oversampling": 20\n  }],\n  "skip": 0,\n  "top": 10,\n  "queryType": "semantic",\n  "queryLanguage": "en-us",\n  "semanticConfiguration": "my-semantic-config"\n}

5. Applying Filters and Facets on Non-Vector Fields

After RRF merging, filters and facets operate on the final document set. This allows preserving existing search functionality: geospatial filters, attribute filters, facet buckets for navigation.

Filters in a hybrid query are applied after full-text and vector search are executed (default postFilter), but can be configured as vectorFilterMode: preFilter if documents need to be excluded before vector search. Testing shows postFilter often gives better relevance, as the filter does not restrict candidates before vector search.

Facets (facets) can be used to split results by fields, e.g., Address/StateProvince, as in the example. Facet results reflect distribution among merged results, not only among one search component.

Important to remember that vector fields cannot be filtered directly. Any selection conditions by metadata must rely on separate text or numeric fields marked with appropriate attributes when creating the index.

"filter": "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",\n  "facets": ["Address/StateProvince"]

6. Optionally: Enabling Semantic Reranker for Re-ranking

Semantic reranker is available by adding queryType: semantic and specifying semanticConfiguration in the request. The reranker uses machine reading to analyze fused RRF results and re-rank them based on semantic correspondence to the query.

This improves search quality for queries where conceptual proximity matters (e.g., synonyms, multilingual search). Benchmarks show that hybrid search with semantic reranker provides significant relevance improvement compared to pure vector or keyword search.

Semantic reranker configuration is set at the index level (section "semantic configurations"). In the request, the configuration name is indicated, as well as the query language (queryLanguage).

If the semantic reranker is not needed, a request can be sent without queryType: semantic. In this case results are ranked only based on RRF, combining BM25 (for text) and HNSW/eKNN (for vectors).

 "queryType": "semantic",\n  "queryLanguage": "en-us",\n  "semanticConfiguration": "my-semantic-config"

7. Testing and Iterating k, oversampling and Filter Behavior

After deployment, it is necessary to experiment with k, oversampling and filter behavior (preFilter/postFilter) to find the optimal speed-accuracy ratio for your domain.

Test different k values (e.g., 10, 50, 100) and oversampling (10, 20, 50) on a representative query sample. Pay attention to changes in @search.rerankerScore and document order in the response.

Compare results with and without semantic reranker. If k is too small, the semantic reranker will not receive enough candidates, and re-ranking will be less effective.

Monitor query latency: increasing k and oversampling increases response time. Find the minimum k that provides acceptable relevance when using semantic reranker.

Use Azure Portal or REST API debugging tools: parameter count: true allows seeing total match count, and @search.rerankerScore field shows the semantic reranker evaluation.

/* Example checking results with count */\n{\n  "count": true,\n  "search": "...",\n  "vectorQueries": [{"kind": "vector", "vector": [...], "k": 50, ...}],\n  "top": 10\n}

8. Deployment and Monitoring Vector Index Quotas and Costs

Ensure your search service was created after April 3, 2024 - such services offer higher vector index quotas. If the service is older, it can be updated to obtain larger quotas.

Generating embeddings through Azure OpenAI or other models incurs charges from the model provider. These costs should be considered when planning data volume and update frequency.

Monitor vector index usage via Azure Monitor metrics: number of vector fields, dimensionality, document count. Exceeding quotas can lead to indexing or query errors.

For hybrid search, monitor query latency: large k and oversampling increase load. If latency becomes unacceptable, reduce oversampling or recheck HNSW configuration.

It is recommended to set up quota exceeded alerts and regularly check index state, especially after large document batches.

/* Checking service version and quotas */\nGET https://my-service.search.windows.net?api-version=2026-04-01\nHeaders: api-key: YOUR_API_KEY

Things to check

  • Index contains at least one searchable text field and one vector field with matching embedding dimensionality.
  • Embeddings are generated and loaded into vector fields (via indexer or direct API).
  • Hybrid query includes search and vectorQueries with correct k and fields values.
  • When using semantic reranker, queryType: semantic and semanticConfiguration are set.
  • Filters and facets reference text/numeric fields, not vector fields directly.
  • Service version is not earlier than April 3, 2024 for increased vector quotas.
  • Cost monitoring for embeddings and query latency is enabled.

Vector fields cannot be used directly in filters. For metadata (e.g., category, geography) a separate text or numeric field marked with filterable and/or facetable attributes is required.

Sources

  1. Microsoft Learn: vector search ↗
  2. Microsoft Learn: hybrid search ↗
  3. scikit-learn: precision_score ↗
Back to top ↑