Skip to Content
Intelligence Service (Python)LLM Gateway & Catalog

LLM Gateway, Model Catalog & Memory

Architecture Overview

v1.1.0 replaces the single LLMClient with a three-layer system:

shared/llm/ ├── gateway.py ← LLM request dispatcher (chat + stream) ├── catalog.py ← DB-backed model registry with TTL cache ├── embeddings.py ← Cloud embedding gateway + sparse encoder ├── ports.py ← ModelEntry / ProviderEntry dataclasses └── tokenizer.py ← Token counting utilities

Provider credentials are never stored in environment variables. They are sealed at rest in PostgreSQL using AES-256-GCM (ENCRYPTION_MASTER_KEY) and decrypted on-demand by shared/crypto/secrets.py.


ModelCatalog (shared/llm/catalog.py)

DB-backed registry over the model_providers and models tables. Maintains an in-process TTL cache (30 s + random jitter) so model resolution stays off the critical request path.

class ModelCatalog: async def resolve_chat( self, slug: Optional[str] = None ) -> tuple[ModelEntry, list[ModelEntry]]: # Returns (primary_model, fallback_chain[]) # Falls back to the DB-configured default when slug is None async def default_embedding(self) -> ModelEntry: # Returns the active embedding model entry async def provider_api_key(self, provider: ProviderEntry) -> Optional[str]: # Decrypt-on-demand; plaintext never touches logs or caches

Fallback chain: each ModelEntry can have a fallback_slug. resolve_chat() walks the chain until it finds a model with no further fallback or encounters a cycle, returning the full ordered list so the gateway can retry in sequence.

Per-Request Model Override

ChatConfig.model (Protobuf optional string) overrides the default for a single request:

message ChatConfig { optional float temperature = 1; optional int32 max_tokens = 2; optional bool use_rag = 3; optional string model = 4; // model slug override optional int32 context_limit = 5; }

The features/chat/grpc_servicer.py extracts config.model and passes it to ModelCatalog.resolve_chat(slug).


LLMGateway (shared/llm/gateway.py)

Dispatches chat and streaming requests to any OpenAI-compatible provider (Google Gemini, OpenAI, Ollama, vLLM, Azure OpenAI):

class LLMGateway: async def chat( self, messages: list[dict], model_entry: ModelEntry, api_key: str, config: ChatConfig | None = None, ) -> str: ... # full response text async def stream( self, messages: list[dict], model_entry: ModelEntry, api_key: str, config: ChatConfig | None = None, ) -> AsyncGenerator[str, None]: ... # yields tokens

Supported providers (configured via model_providers table slug):

Provider SlugAPI CompatibilityNotes
googleOpenAI-compatible via Gemini RESTbase_url → Gemini endpoint
openaiOpenAI nativehttps://api.openai.com/v1
ollamaOpenAI-compatiblebase_url=http://host.docker.internal:11434/v1
customAny OpenAI-compatible endpointAdmin-configured base_url

Retry Logic

Uses tenacity with exponential backoff on rate-limit and timeout errors (up to 3 attempts, 2 s → 4 s → 8 s). Retries walk the fallback chain when the primary model fails.


CloudEmbeddingGateway (shared/llm/embeddings.py)

Generates dense vector embeddings via any OpenAI-compatible embedding endpoint. Default: text-embedding-3-large → 3072 dimensions.

class CloudEmbeddingGateway: dims: int # 3072 default async def embed_query(self, text: str) -> list[float]: # Single embedding for retrieval queries async def embed_documents(self, texts: list[str]) -> list[list[float]]: # Batched (64 per API call) for ingestion

Constructed by build_embedding_gateway() which reads the active embedding model from ModelCatalog and decrypts its provider API key.

HashedSparseEncoder — BM25-style Sparse Leg

Dependency-free sparse encoder producing SparseVector-compatible (indices, values) pairs for Qdrant’s sparse slot. Runs entirely in-process — no network call, no model download.

class HashedSparseEncoder: dim: int = 65_536 # hash space def encode(self, text: str) -> tuple[list[int], list[float]]: # Tokenise (lowercase alphanum), sublinear TF weighting, SHA-1 hashing # Returns (sorted unique indices, weights)

Prompt Construction (prompts/chat_system.md)

System prompt is externalized to server/intelligence/prompts/chat_system.md (loaded at startup by prompts/loader.py). The full prompt is assembled from 4 components at request time:

  1. System prompt — loaded from chat_system.md
  2. User memory — from user_memories.memory (persistent cross-session context)
  3. Retrieved context — top-k chunks from Qdrant hybrid search, formatted as numbered citations
  4. Conversation history — prior chat_messages rows (multi-turn coherence)

ChatConfig.use_rag = false skips the Qdrant retrieval step — the LLM operates in pure conversation mode with no knowledge base context.


User Memory (features/chat/memory.py)

Persistent per-user memory stored in PostgreSQL:

CREATE TABLE user_memories ( user_id VARCHAR(255) PRIMARY KEY, memory TEXT NOT NULL DEFAULT '', metadata JSONB NOT NULL DEFAULT '{}', updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW() );

Update flow (v1.1.0 — async):

  1. After each chat turn, Intelligence publishes a chat.completed.v1 event to Redis CHAT_EVENTS stream
  2. The memory consumer group in worker.py picks it up
  3. handle_memory_extraction() generates a memory summary via LLM and upserts user_memories

This decouples memory extraction from the hot chat path — TTFT is not affected by memory writes.


LLM Provider Management

[!IMPORTANT] No LLM API keys go in .env. Set ENCRYPTION_MASTER_KEY and configure all providers via the Admin Dashboard (/dashboard → Admin → Models).

The database stores:

  • model_providers — slug, display name, base URL, AES-256-GCM encrypted API key blob
  • models — slug, provider FK, token pricing, default flags, fallback chain slug

Tokens vs. Centralized Platform Credits

OpenTier enforces a strict separation between provider workload units and platform billing:

  1. Native Raw Tokens (Provider Level):

    • AI providers (Google Gemini, OpenAI, Anthropic, Ollama) measure inference strictly in raw input/prompt tokens and output/completion tokens.
    • The intelligence runtime receives or counts these exact raw tokens and stores them in chat_messages metadata and Redis event envelopes.
    • External providers have zero awareness of OpenTier credits.
  2. OpenTier Credits (Centralized Platform Currency):

    • Credits are OpenTier’s unified internal billing currency, denominated at **1 credit = 0.001USD∗∗(1,000credits=0.001 USD** (1,000 credits = 1.00).
    • The models catalog columns input_cost_per_mtok and output_cost_per_mtok define the credit rate per 1,000,000 provider tokens: Credits Billed=Tokensin×input_cost_per_mtok+Tokensout×output_cost_per_mtok1,000,000\text{Credits Billed} = \frac{\text{Tokens}_{\text{in}} \times \text{input\_cost\_per\_mtok} + \text{Tokens}_{\text{out}} \times \text{output\_cost\_per\_mtok}}{1,000,000}
    • For example, with Gemini 3.5 Flash Lite (0.10input/0.10 input / 0.40 output per 1M tokens), the catalog sets rates to 100.0 and 400.0 credits/MTok. A turn with 1,500 input tokens and 300 output tokens bills: (1,5001,000,000×100)+(3001,000,000×400)=0.15+0.12=0.27 credits\left(\frac{1,500}{1,000,000} \times 100\right) + \left(\frac{300}{1,000,000} \times 400\right) = 0.15 + 0.12 = 0.27 \text{ credits}
    • Pre-stream reservations (credit_holds), ledger records (credit_transactions), and user balances (user_credit_balances) operate exclusively in credits, while preserving the raw token audit counts on every row.
Last updated on