Skip to Content

Hybrid Retrieval

Overview

v1.1.0 replaces the PostgreSQL hybrid_search() SQL function with Qdrant server-side Reciprocal Rank Fusion (RRF) over two vector legs:

LegTechniqueDimensionNotes
DenseCosine similarity3072text-embedding-3-large via OpenAI API
SparseBM25-style hashed TF65 536 hash indicesHashedSparseEncoder (in-process, no model)
FusionRRF (models.Fusion.RRF)—Qdrant server-side, no Python ranking code

Module: shared/qdrant/client.py

shared/qdrant/ ├── client.py ← QdrantVectorStore + QdrantHybridRetriever └── __init__.py

QdrantVectorStore

Thin async adapter over qdrant-client. One instance per process, connected to QDRANT_URL.

class QdrantVectorStore: COLLECTION = "knowledge_chunks" async def ensure_collection(self) -> None: # Creates collection if absent; recreates on dimension mismatch # HNSW config: m=16, ef_construct=128 # Payload indexes: user_id (tenant), is_global, document_id, chunk_index async def upsert_chunks(self, items: list[dict]) -> int: # Writes dense + sparse PointStruct batch to Qdrant async def hybrid_search( self, query_dense: list[float], query_sparse_indices: list[int], query_sparse_values: list[float], user_id: str, top_k: int = 5, ) -> list[RetrievedChunk]: ... async def delete_document(self, document_id: str) -> None: ... async def delete_user_data(self, user_id: str) -> None: ...

QdrantHybridRetriever

High-level retriever used by the chat pipeline:

class QdrantHybridRetriever: store: QdrantVectorStore embedder: CloudEmbeddingGateway # dense embeddings sparse: HashedSparseEncoder # sparse encoding async def search( self, query: str, user_id: str, top_k: int = 20, document_id: Optional[uuid.UUID] = None, ) -> list[SearchResult]: # 1. dense = embedder.embed_query(query) → 3072-dim vector # 2. s_idx, s_val = sparse.encode(query) → hashed BM25 indices # 3. hits = store.hybrid_search(dense, sparse) → Qdrant RRF fusion # 4. Optionally backfill document_title from PostgreSQL if missing in payload # 5. Return List[SearchResult]

Hybrid Search — Request Structure

Qdrant’s query_points with server-side RRF:

response = await client.query_points( collection_name="knowledge_chunks", prefetch=[ # Leg 1: Dense cosine Prefetch(query=query_dense, using="dense", limit=top_k * 4, filter=tenant_filter), # Leg 2: Sparse BM25 Prefetch(query=SparseVector(indices=q_idx, values=q_val), using="sparse", limit=top_k * 4, filter=tenant_filter), ], query=FusionQuery(fusion=Fusion.RRF), # server-side RRF limit=top_k, with_payload=True, )

Tenant filter: each search is scoped to user_id == <current_user> OR is_global == true — users only retrieve their own documents plus shared global documents.


Pipeline: QdrantHybridRetriever.search()

Collection Schema

FieldTypeConfiguration
dense (vector)float[]Size: 3072, Distance: Cosine, HNSW m=16 ef_construct=128
sparse (vector)SparseVectorBM25-style hash indices (dim=65536)
user_id (payload)keywordTenant index (is_tenant=true) for partition performance
is_global (payload)boolIndexed for fast global doc filtering
document_id (payload)keywordIndexed for per-document deletion
chunk_index (payload)integerOrdered chunk position within document
content (payload)textRaw chunk text (not indexed)
document_title (payload)textDenormalized for zero-join retrieval
source_url (payload)textOrigin URL or repo path

Retrieval in the Chat Pipeline

The features/chat/pipeline.py (QueryPipeline) calls retriever.search() when use_rag=true:

chunks = await retriever.search( query=user_message, user_id=user_id, top_k=config.top_k or 20, ) context = format_citations(chunks) # [1] content [source: title]

The formatted context is injected into the LLM prompt alongside user memory and conversation history.

Last updated on