Skip to Content
Data ArchitectureVector & AI Schema

Vector & AI Schema

Hybrid Vector Architecture: Qdrant + PostgreSQL 18

OpenTier v1.1.0 utilizes a decoupled, high-performance hybrid storage architecture:

  • Qdrant Vector Database: Dedicated system of record for high-dimensional vector embeddings and sub-millisecond retrieval. Stores dense embeddings (3072d via text-embedding-3-large) paired with sparse BM25 lexical token vectors for Reciprocal Rank Fusion (RRF).
  • PostgreSQL 18: System of record for relational metadata, user ownership, document sources, credit balances, and transaction audit ledgers.

Qdrant Collection Schema (knowledge_chunks)

Vector FieldTypeDescription
denseDense Vector (size = 3072, Cosine)Captures deep semantic meaning and cross-lingual conceptual similarity.
sparseSparse Vector (Indices + Values)Captures exact keyword, symbol, and BM25 token matches via HashedSparseEncoder.

HNSW Vector Index Parameters

  • m = 16: Number of bi-directional edges per node in the HNSW graph.
  • ef_construct = 128: Search depth during index construction.
  • Distance metric: Cosine.

Payload Indexes (Hardware-Level Tenant Partitioning)

Qdrant payload fields are indexed to enable sub-millisecond pre-filtering prior to vector similarity calculations:

  • user_id: Keyword index (is_tenant = true) for hardware-level partition co-location per user.
  • is_global: Boolean index for system-wide shared knowledge base resources.
  • document_id: Keyword index for cascade deletion and document grouping.
  • chunk_index: Integer index for chronological ordering.

Qdrant Point Payload Structure

{ "id": "c7a8e23b-6e1b-4f51-87a4-4f510a72cb3e", # Chunk UUID "vector": { "dense": [0.0124, -0.0451, ..., 0.0892], # 3072 dimensions "sparse": { "indices": [1042, 45120, 58912], # Hashed unigram tokens "values": [0.851, 1.204, 0.652] # Sublinear TF weights } }, "payload": { "user_id": "9b1deb4d-3b7d-4bad-9bdd-2b0d7b3dcb6d", "is_global": false, "document_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479", "chunk_index": 0, "content": "Raw markdown chunk content...", "document_title": "Architecture Overview", "source_url": "https://github.com/Celestial-0/OpenTier", "metadata": { "author": "OpenTier", "section": "System Topology" } } }

Relational Catalog Model (PostgreSQL 18)

documents Table

Stores root document metadata and complete raw source text:

  • user_id: VARCHAR string matching session user_id for loose coupling.
  • document_type: Enum string (text, website, github_repo, file).
  • is_global: When true, indexed in Qdrant as accessible to all users.
  • source_url: URL or repository origin path.
  • metadata: JSONB containing author, commit hash, or crawl depth parameters.

ingestion_jobs Table

Tracks asynchronous background processing state in worker.py:

  • status: queued → processing → completed or failed.
  • progress_percent: Real-time completion tracker for UI progress bars.
  • errors: JSONB array capturing chunking or scraping error details.

End-to-End Ingestion Lifecycle

Last updated on