Learn how to effectively evaluate the precision and recall of AI models, especially when working with small question datasets. This guide covers the concepts, calculation methods, and practical considerations using scikit-learn.
ai / precision / recall / evaluation / small dataset / machine learning / scikit-learn / classification
Step-by-step guide to creating an index, generating embeddings, and executing hybrid queries that combine full-text and vector search using Azure AI Search.
azure-ai-search / hybrid-search / vector-search / semantic-reranker / index-schema / embeddings
This article explains how to choose effective chunk boundaries and overlap in retrieval-augmented generation (RAG) systems, comparing fixed-size and document-aware strategies. Using a hypothetical support note example, it demonstrates the risks of splitting across semantic units and shows how metadata and overlap can preserve context. The guidance is based on Microsoft Azure documentation for chunking in vector search and RAG workflows.
RAG / chunking / Azure AI Search / information retrieval / document processing
Compare three similarity calculations on small vectors and see when normalization changes the ranking.
ai
Explore the principles of semantic search, its reliance on embeddings, and the reasons why text that appears similar might not provide the desired answer. Learn about vector search, hybrid search, and their applications in information retrieval.
semantic search / vector search / embeddings / AI / information retrieval / natural language processing / hybrid search / Azure AI Search
A technical guide explaining why syntactically correct JSON from Large Language Models (LLMs) is insufficient for production systems and how JSON Schema provides the necessary structural and type guarantees.
LLM / JSON / JSON Schema / Data Validation / Software Engineering / Python