LLM Gateway, Model Catalog & Memory
Architecture Overview
v1.1.0 replaces the single LLMClient with a three-layer system:
shared/llm/
├── gateway.py ← LLM request dispatcher (chat + stream)
├── catalog.py ← DB-backed model registry with TTL cache
├── embeddings.py ← Cloud embedding gateway + sparse encoder
├── ports.py ← ModelEntry / ProviderEntry dataclasses
└── tokenizer.py ← Token counting utilitiesProvider credentials are never stored in environment variables. They are sealed at rest in PostgreSQL using AES-256-GCM (ENCRYPTION_MASTER_KEY) and decrypted on-demand by shared/crypto/secrets.py.
ModelCatalog (shared/llm/catalog.py)
DB-backed registry over the model_providers and models tables. Maintains an in-process TTL cache (30 s + random jitter) so model resolution stays off the critical request path.
class ModelCatalog:
async def resolve_chat(
self, slug: Optional[str] = None
) -> tuple[ModelEntry, list[ModelEntry]]:
# Returns (primary_model, fallback_chain[])
# Falls back to the DB-configured default when slug is None
async def default_embedding(self) -> ModelEntry:
# Returns the active embedding model entry
async def provider_api_key(self, provider: ProviderEntry) -> Optional[str]:
# Decrypt-on-demand; plaintext never touches logs or cachesFallback chain: each ModelEntry can have a fallback_slug. resolve_chat() walks the chain until it finds a model with no further fallback or encounters a cycle, returning the full ordered list so the gateway can retry in sequence.
Per-Request Model Override
ChatConfig.model (Protobuf optional string) overrides the default for a single request:
message ChatConfig {
optional float temperature = 1;
optional int32 max_tokens = 2;
optional bool use_rag = 3;
optional string model = 4; // model slug override
optional int32 context_limit = 5;
}The features/chat/grpc_servicer.py extracts config.model and passes it to ModelCatalog.resolve_chat(slug).
LLMGateway (shared/llm/gateway.py)
Dispatches chat and streaming requests to any OpenAI-compatible provider (Google Gemini, OpenAI, Ollama, vLLM, Azure OpenAI):
class LLMGateway:
async def chat(
self,
messages: list[dict],
model_entry: ModelEntry,
api_key: str,
config: ChatConfig | None = None,
) -> str: ... # full response text
async def stream(
self,
messages: list[dict],
model_entry: ModelEntry,
api_key: str,
config: ChatConfig | None = None,
) -> AsyncGenerator[str, None]: ... # yields tokensSupported providers (configured via model_providers table slug):
| Provider Slug | API Compatibility | Notes |
|---|---|---|
google | OpenAI-compatible via Gemini REST | base_url → Gemini endpoint |
openai | OpenAI native | https://api.openai.com/v1 |
ollama | OpenAI-compatible | base_url=http://host.docker.internal:11434/v1 |
custom | Any OpenAI-compatible endpoint | Admin-configured base_url |
Retry Logic
Uses tenacity with exponential backoff on rate-limit and timeout errors (up to 3 attempts, 2 s → 4 s → 8 s). Retries walk the fallback chain when the primary model fails.
CloudEmbeddingGateway (shared/llm/embeddings.py)
Generates dense vector embeddings via any OpenAI-compatible embedding endpoint. Default: text-embedding-3-large → 3072 dimensions.
class CloudEmbeddingGateway:
dims: int # 3072 default
async def embed_query(self, text: str) -> list[float]:
# Single embedding for retrieval queries
async def embed_documents(self, texts: list[str]) -> list[list[float]]:
# Batched (64 per API call) for ingestionConstructed by build_embedding_gateway() which reads the active embedding model from ModelCatalog and decrypts its provider API key.
HashedSparseEncoder — BM25-style Sparse Leg
Dependency-free sparse encoder producing SparseVector-compatible (indices, values) pairs for Qdrant’s sparse slot. Runs entirely in-process — no network call, no model download.
class HashedSparseEncoder:
dim: int = 65_536 # hash space
def encode(self, text: str) -> tuple[list[int], list[float]]:
# Tokenise (lowercase alphanum), sublinear TF weighting, SHA-1 hashing
# Returns (sorted unique indices, weights)Prompt Construction (prompts/chat_system.md)
System prompt is externalized to server/intelligence/prompts/chat_system.md (loaded at startup by prompts/loader.py). The full prompt is assembled from 4 components at request time:
- System prompt — loaded from
chat_system.md - User memory — from
user_memories.memory(persistent cross-session context) - Retrieved context — top-k chunks from Qdrant hybrid search, formatted as numbered citations
- Conversation history — prior
chat_messagesrows (multi-turn coherence)
ChatConfig.use_rag = false skips the Qdrant retrieval step — the LLM operates in pure conversation mode with no knowledge base context.
User Memory (features/chat/memory.py)
Persistent per-user memory stored in PostgreSQL:
CREATE TABLE user_memories (
user_id VARCHAR(255) PRIMARY KEY,
memory TEXT NOT NULL DEFAULT '',
metadata JSONB NOT NULL DEFAULT '{}',
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);Update flow (v1.1.0 — async):
- After each chat turn, Intelligence publishes a
chat.completed.v1event to RedisCHAT_EVENTSstream - The
memoryconsumer group inworker.pypicks it up handle_memory_extraction()generates a memory summary via LLM and upsertsuser_memories
This decouples memory extraction from the hot chat path — TTFT is not affected by memory writes.
LLM Provider Management
[!IMPORTANT] No LLM API keys go in
.env. SetENCRYPTION_MASTER_KEYand configure all providers via the Admin Dashboard (/dashboard → Admin → Models).
The database stores:
model_providers— slug, display name, base URL, AES-256-GCM encrypted API key blobmodels— slug, provider FK, token pricing, default flags, fallback chain slug
Tokens vs. Centralized Platform Credits
OpenTier enforces a strict separation between provider workload units and platform billing:
-
Native Raw Tokens (Provider Level):
- AI providers (Google Gemini, OpenAI, Anthropic, Ollama) measure inference strictly in raw input/prompt tokens and output/completion tokens.
- The intelligence runtime receives or counts these exact raw tokens and stores them in
chat_messagesmetadata and Redis event envelopes. - External providers have zero awareness of OpenTier credits.
-
OpenTier Credits (Centralized Platform Currency):
- Credits are OpenTier’s unified internal billing currency, denominated at **1 credit = 1.00).
- The
modelscatalog columnsinput_cost_per_mtokandoutput_cost_per_mtokdefine the credit rate per 1,000,000 provider tokens: - For example, with Gemini 3.5 Flash Lite (0.40 output per 1M tokens), the catalog sets rates to
100.0and400.0credits/MTok. A turn with 1,500 input tokens and 300 output tokens bills: - Pre-stream reservations (
credit_holds), ledger records (credit_transactions), and user balances (user_credit_balances) operate exclusively in credits, while preserving the raw token audit counts on every row.