मल्टीमोडल कॉंटेंट इंडेक्सिंग: एज़्यूर एआई सर्च में छवि और पाठ के वेक्टराइज़ेशन
एज़्यूर एआई सर्च में मल्टीमोडल सर्च सेट अप करने की गाइड, पाठ और छवियों के वेक्टराइज़ेशन का उपयोग करके. इंडेक्स बनाना, मॉडल का चयन (उदाहरण के लिए CLIP), एम्बेडिंग्स के लिए विधि का चयन और हाइब्रिड क्वेरीज़ का उपयोग करके रिलेवेंस को सुधारने के कदमों का वर्णन करता है.
इस पृष्ठ पर
संक्षिप्त उत्तर
सीधे टेक्स्ट से चित्र खोजने के लिए एक संगत multimodal मॉडल के जोड़े वाले टेक्स्ट और चित्र encoder उपयोग करें। समान vector लंबाई field के लिए जरूरी है, लेकिन स्वतंत्र रूप से प्रशिक्षित spaces को समान नहीं बनाती। चित्र vectors को vector field और captions को searchable text field में रखें। Query vector संबंधित text encoder से बनाएं। Hybrid ranking सूचियां मिलाती है, फिर भी प्रतिनिधि queries पर गुणवत्ता मापना आवश्यक है।
एज़्यूर एआई सर्च में मल्टीमोडल सर्च के बारे में
एज़्यूर एआई सर्च में मल्टीमोडल सर्च पाठ या छवि के प्रकार को নির্বाह करके समान कॉंटेंट खोजने की अनुमति देता है, जो डेटा को नंबरिकल वेक्टर्स (एम्बेडिंग्स) में परिवर्तित करके, जहां वेक्टर निकटता सेमांटिक समानता को प्रतिबिंबित करती है. उदाहरण के लिए, क्वेरी 'पप्पी' पाठ के बिना छवि लेबल के डॉग छवियों को लौटाता है.
यह विधि वेक्टर सर्च पर आधारित है, जो केवल पाठ-से-पाठ के साथ-साथ पाठ-से-छवि मैचिंग भी समर्थन करती है, जो हाइब्रिड सीनारियो में मूल्यवान है. वेक्टराइज़ेशन कॉन्सेपटुअल समानता को पकड़ती है, जैसे 'डॉग' और 'कैनाइन' के बीच या पाठ और इसकी संबंधित छवि के बीच.
वेक्टर डेटा के लिए इंडेक्स तैयार करना
Vector field का प्रकार Collection(Edm.Single) है, scalar Edm.Single नहीं। Dimensions चुने encoder के वास्तविक output से मिलनी चाहिए। उदाहरण में 512 components माने गए हैं; यह सभी CLIP या Azure मॉडल की सामान्य dimension नहीं है। python -m pip install azure-search-documents से dependency स्थापित करें और HnswAlgorithmConfiguration तथा VectorSearchProfile वाला SDK उपयोग करें। उदाहरण named algorithm को profile से जोड़ता है और text पर पूर्ण पाठ खोज सक्षम करता है। यह स्थानीय index definition बनाता है जिसका अपेक्षित output ['id', 'text', 'image_vector'] है; Azure resource या documents नहीं बनाता। Deployment के लिए configured service और authenticated index-management client चाहिए।
from azure.search.documents.indexes.models import (
SearchIndex, SearchField, SearchFieldDataType, VectorSearch,
HnswAlgorithmConfiguration, VectorSearchProfile,
)
dimensions = 512 # Example only: use the chosen encoder's actual output length.
profile = "multimodal-profile"
index = SearchIndex(
name="multimodal-index",
fields=[
SearchField(name="id", type=SearchFieldDataType.String, key=True),
SearchField(name="text", type=SearchFieldDataType.String, searchable=True),
SearchField(
name="image_vector",
type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
searchable=True,
vector_search_dimensions=dimensions,
vector_search_profile_name=profile,
),
],
vector_search=VectorSearch(
algorithms=[HnswAlgorithmConfiguration(name="hnsw-config")],
profiles=[VectorSearchProfile(name=profile, algorithm_configuration_name="hnsw-config")],
),
)
print([field.name for field in index.fields])छवि वेक्टराइज़ेशन के लिए मॉडल का चयन करना
साथ प्रशिक्षित टेक्स्ट और चित्र encoder चुनें तथा model version और preprocessing बनाए रखें। CLIP ऐसे मॉडल का एक परिवार है; अलग checkpoints की vector लंबाई अलग हो सकती है। किसी स्वतंत्र text-only मॉडल का query vector केवल समान लंबाई से image space के अनुकूल नहीं हो जाता। Model बदलने पर सामान्यतः indexed vectors दोबारा बनाने पड़ते हैं। केवल field dimension बदलना पुराने vectors का अर्थ नहीं बदलता।
पाठ वेक्टराइज़ेशन के लिए मॉडल का चयन करना
दो वैध डिजाइनों में अंतर करें। Direct image search में paired text encoder की query को image embeddings से तुलना करते हैं। दूसरे तरीके में चित्र captions बनाकर captions और queries को एक ही text मॉडल से encode करते हैं। तब खोज descriptions पर होती है और उनमें न लिखे दृश्य विवरण खो सकते हैं। Text-only ada-002 या SBERT vector की CLIP image vector से सीधे तुलना न करें। लंबाई मिलाने से coordinate का अर्थ समान नहीं होता।
इंटिग्रेटेड वर्सा एक्स्टर्नल वेक्ट्राइज़ेशन
Ingestion के दौरान indexer skillset सामग्री निकाल और embeddings बना सकता है। Query-time vectorizer अलग configuration है जो query input को vector बनाता है; embedding skill अकेला query vectorization नहीं स्थापित करता। External vectorization में document vectors upload से पहले और query vectors application में बनते हैं। दोनों तरीकों में हर space के model, preprocessing और dimensions संगत होने चाहिए। बाद में update के लिए इनका version भी रखें।
एक मल्टीमोडल इंडेक्स की रचना करना
एक मल्टीमोडल इंडेक्स पाठ, छवियों, और उनके वेक्ट्रल प्रतिनिधित्वों को एक शेमा में जोड़ता है। पाठ और इमेज_वेक्ट्र जैसी फ़ील्डें पूर्ण-टेक्स्ट और वेक्ट्र सर्च दोनों के लिए अनुमति प्रदान करती हैं, जो हाइब्रिड क्वेरीज़ के लिए आधार बनाती हैं।
एक मल्टीमोडल सर्च क्वेरी की एक्जीक्यूशन
नीचे का helper request body बनाता है; यह पूरा network client नहीं है और embedding नहीं बनाता। Paired text encoder का असली vector दें, फिर service endpoint, supported REST API version और credentials से body भेजें। लंबाई की जांच schema mismatch पकड़ती है, model compatibility सिद्ध नहीं करती। Text branch searchable text पर और vector branch image_vector पर खोजती है। RRF ranks मिलाता है, बेहतर relevance की गारंटी नहीं देता।
def hybrid_payload(query_vector):
# Precondition: query_vector comes from the paired text encoder
# compatible with the indexed image encoder; length alone is insufficient.
if len(query_vector) != 512:
raise ValueError("Expected the example encoder's 512 components")
return {
"search": "puppy",
"vectorQueries": [{
"kind": "vector", "vector": list(query_vector),
"k": 10, "fields": "image_vector",
}],
"select": "id,text", "top": 5,
}सर्च परिणामों की एवल्यूएशन और सुधार
लेबल वाले query-result pairs पर text-only, vector-only और hybrid retrieval तुलना करें। Semantic ranking के लिए supported content तथा अलग configuration चाहिए। इसी dataset पर k और filters चुनें। Oversampling समर्थित compressed-vector rescoring configuration से जुड़ा है, uncompressed index का सार्वभौमिक विकल्प नहीं। नीचे precision की शैक्षिक गणना चुने binary decisions को मापती है, खोज गुणवत्ता के हर पहलू को नहीं। असली query relevance के लिए उपयुक्त labels भी चाहिए।
from sklearn.metrics import precision_score
# Example evaluation (hypothetical values)
y_true = [1, 0, 1, 1, 0] # 1 = relevant, 0 = not
y_pred = [1, 1, 1, 0, 0] # predicted results
precision = precision_score(y_true, y_pred, average='binary')
print(f"Precision: {precision:.2f}")क्या जाँचें
- Paired चित्र और टेक्स्ट encoder का space तथा version संगत है
- Dimensions field से मेल खाती हैं; लंबाई अकेले compatibility नहीं सिद्ध करती
- Vector field मौजूदा profile और algorithm से जुड़ा है; captions searchable हैं
- Query vectorization ingestion skills से अलग configured है
- गुणवत्ता labeled queries पर तुलना होती है, सुधार मान नहीं लेते
उपयोग की सीमाएँ
किसी तारीख से पहले बने हर service को असमर्थ मानने के बजाय उसके वास्तविक capabilities देखें। कुछ पुराने services vector index नहीं जोड़ सकते और migration चाहिए। Embeddings, image descriptions, indexing और optional semantic ranking की अलग लागत हो सकती है। Snippets schema और request की बनावट समझाते हैं; वे retrieval quality प्रमाणित नहीं करते। मूल्यांकन के लिए model, documents और queries चाहिए।