TATECHATLAS
◎ English
Artificial intelligence / Guide

Filtered Vector Search Strategy: Applying Filters to Vector Queries in Azure AI Search

Filtered vector search in Azure AI Search combines vector similarity with metadata filtering. Pre-filtering reduces the candidate set before scoring, potentially lowering latency but risking recall loss. Post-filtering preserves recall but may return fewer results. Testing both modes is essential to find the right trade-off for your data.

On this page

Filtered vector search in Azure AI Search lets you attach a filter expression to a vector query. The filter targets non-vector metadata fields, not the vector field itself. The engine can apply the filter before or after vector similarity scoring. Pre-filtering narrows the candidate set before nearest-neighbor search, which can silently drop relevant documents that fail the metadata condition but may lower latency for selective filters. Post-filtering computes similarity over the full index and then removes non-matching results, potentially returning fewer than k documents. The choice is a strategy decision: post-filtering is often recommended when using semantic ranker because it needs a larger candidate pool, but the documentation advises testing both modes for your data. To implement it, you must design filterable text or numeric fields alongside your vector fields, and set the vectorFilterMode parameter in the query request.

What Filtered Vector Search Means in Azure AI Search

Filtered vector search in Azure AI Search means attaching a filter expression to a query that also contains a vector query. The filter operates on text or numeric fields in the index schema, never on the vector field itself. This lets you include or exclude documents based on metadata criteria while still performing nearest-neighbor similarity search over embeddings. The engine can process the filter either before or after the vector query execution, which changes which documents enter the similarity-scoring pipeline.

Pre-filtering reduces the candidate set before kNN scoring, which can lower latency when the filter is highly selective. However, because it may silently drop conceptually relevant documents, the documentation advises benchmarking both pre-filter and post-filter modes with representative queries to confirm the trade-off for your specific data distribution.

Pre-Filter Versus Post-Filter: The Central Decision

The search engine can apply the filter before or after executing the vector query. This is a strategy decision, not a syntax choice, and it directly affects the candidate pool size. In pre-filtering, the filter runs first and reduces the set of documents that the vector search will score. In post-filtering, the vector search runs over the entire index, and the filter removes non-matching results afterward. The choice influences recall, precision, and the number of results returned. The documentation does not prescribe a universal best mode; it advises testing both to determine which works for your scenario.

Designing Filterable Metadata Fields

Because vector fields are not filterable, every attribute you want to constrain at query time must be stored in a separate non-vector field that is marked as filterable in the index schema. For example, if you want to filter by category, location, or date, you need to create text or numeric fields for those properties and set their filterable attribute to true. The filter expression in the query references these fields, not the vector field. Without them, filtered vector search is impossible. This schema design step is a prerequisite that must be completed before any query can use a filter alongside a vector search.

The vectorFilterMode Parameter and a Concrete Query

The hybrid search documentation provides a concrete query that demonstrates filtered vector search in practice. The request includes a filter expression using geo.distance to find hotels within 300 kilometers of a point, and sets vectorFilterMode to postFilter. The query also contains a full-text search, two vector queries targeting different vector fields, facets, and semantic ranking. The filter is geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300, and vectorFilterMode is postFilter. This means the engine first finds the 50 nearest neighbors by embedding similarity across all hotels, then removes those outside the radius. The result set may contain fewer than 50 documents. The vector queries specify k=50 and oversampling to improve recall. The example shows how filter, vector queries, and other features coexist in a single request.

POST https://{{searchServiceName}}.search.windows.net/indexes/hotels-vector-quickstart/docs/search?api-version=2026-04-01
content-type: application/JSON

{
  "count": true,
  "search": "historic hotel walk to restaurants and shopping",
  "select": "HotelId, HotelName, Category, Description, Address/City, Address/StateProvince",
  "filter": "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
  "vectorFilterMode": "postFilter",
  "facets": ["Address/StateProvince"],
  "vectorQueries": [
    {
      "kind": "vector",
      "vector": [<array of embeddings>],
      "k": 50,
      "fields": "DescriptionVector",
      "exhaustive": true,
      "oversampling": 20
    },
    {
      "kind": "vector",
      "vector": [<array of embeddings>],
      "k": 50,
      "fields": "Description_frVector",
      "exhaustive": false,
      "oversampling": 10
    }
  ],
  "skip": 0,
  "top": 10,
  "queryType": "semantic",
  "queryLanguage": "en-us",
  "semanticConfiguration": "my-semantic-config"
}

Interaction with Semantic Ranker

When using semantic ranker, the documentation suggests post-filtering as a starting point. The reason is that semantic ranker needs a larger candidate pool to read and reorder. Pre-filtering might reduce the pool too much, limiting the ranker's effectiveness. However, this is not a strict rule; the source explicitly recommends testing to confirm which behavior is best for your queries. The semantic ranker operates on the merged results from full-text and vector search, and the filter can be applied before or after that merge, affecting the input to the ranker. If you set k to 50 for vector queries when using semantic ranker, you maximize the inputs available for reranking.

Trade-offs: Recall Loss Under Pre-Filtering

Pre-filtering narrows the candidate set before vector similarity is computed. This can silently drop conceptually relevant documents that fail the metadata condition, even if they are very similar to the query vector. For example, a hotel that is a perfect semantic match but just outside the geographic radius would be excluded before scoring. Post-filtering computes similarity over the full index and then removes non-matching results, which can return fewer than k items but ensures that the similarity ranking considered all documents. The implication is that the choice between pre-filter and post-filter trades recall against result count and latency. Pre-filtering preserves k and may reduce latency for selective filters but risks missing relevant documents; post-filtering preserves recall but may return fewer results and require more computation.

Combining Filters with Facets and Scoring Profiles

Filters and facets target data structures within the index that are distinct from the inverted indexes used for full-text search and the vector indexes used for vector search. When filters and faceted operations execute, the search engine can apply the operational result to the hybrid search results in the response. This means a filter can coexist with faceted navigation and scoring profiles without conflicting with vector ranking. Facets are computed over the final result set after filtering and merging. Scoring profiles can be applied to text fields, but explicit sort orders override relevance-ranked results, so avoid orderby if you want similarity and BM25 relevance to drive the ranking.

Applicability Limits and When to Reconsider the Approach

The primary constraints are that vector fields themselves are not filterable, so every filter criterion must live in a separate field. Post-filtering can yield fewer results than requested k, which may require handling in your application. Pre-filtering can reduce recall but may improve latency for highly selective filters. The documentation recommends testing to confirm which mode is best rather than assuming a single correct answer. If your filter criteria are highly selective and you need exact k results, you might need to adjust oversampling or reconsider the filter design. Testing should include measuring recall, precision, and latency under realistic workloads. If no combination of filter mode and oversampling gives acceptable results, consider redesigning the filter to be less restrictive or moving some constraints into the vector embedding itself through metadata-aware embedding strategies.

Things to check

  • Verify that every filter criterion in the query corresponds to a non-vector field marked as filterable in the index schema.
  • Confirm that the vectorFilterMode parameter is set to either preFilter or postFilter in the hybrid query request.
  • Benchmark both pre-filter and post-filter modes with representative queries to determine the optimal balance of recall, latency, and result count for your specific dataset.
  • Check that post-filtered results may contain fewer documents than the requested k; handle this in your application logic.
  • Ensure that facets and scoring profiles are applied to the hybrid result set and do not interfere with vector ranking.

Vector fields themselves cannot be used in filter expressions; every filter criterion requires a dedicated non-vector field. Post-filtering can return fewer documents than the requested k because similarity is computed over the full index and then results are removed. Pre-filtering can silently reduce recall by shrinking the candidate pool before nearest-neighbor search runs, though it may lower latency for selective filters. The documentation recommends testing both modes rather than prescribing one as universally correct; the best choice depends on your data distribution and query patterns.

Sources

  1. Microsoft Learn: vector search ↗
  2. Microsoft Learn: hybrid search ↗
  3. scikit-learn: precision_score ↗
Back to top ↑