Skip to Content

Memory Subsystem

Overview

In OpenTier v1.1.0, user memory is a persistent, personalized memory layer that spans across conversations while being extracted asynchronously via Redis Streams.

Memory extraction is intentionally decoupled from the hot streaming chat path:

  1. During chat, the user’s existing memory is fetched with a quick indexed lookup and injected into the prompt.
  2. Upon completion of a chat turn, the intelligence service publishes an event to Redis Streams.
  3. A background worker picks up the event, executes LLM-driven memory extraction, and atomically upserts the memory ledger.

This architecture ensures zero latency impact on Time-to-First-Token (TTFT) and token streaming throughput.


Data Model (user_memories)

The transactional system of record for memory is PostgreSQL 18:

CREATE TABLE IF NOT EXISTS user_memories ( user_id VARCHAR(255) PRIMARY KEY, memory TEXT NOT NULL DEFAULT '', metadata JSONB NOT NULL DEFAULT '{}', updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW() ); CREATE INDEX IF NOT EXISTS idx_user_memories_updated_at ON user_memories(updated_at);

Schema Attributes

ColumnTypeDescription
user_idVARCHAR(255)Primary key identifying the user.
memoryTEXTCompact, synthesized text representation of user preferences, context, and facts.
metadataJSONBStructured metadata (e.g. extraction count, model used, topic tags).
updated_atTIMESTAMPTZTimestamp of the most recent memory update.

Asynchronous Memory Extraction Flow


Worker Consumer Handler (features/chat/memory.py)

The worker registers consumer group memory listening to stream Streams.CHAT_EVENTS:

# features/chat/memory.py async def handle_memory_extraction(envelope: EventEnvelope, msg_id: str) -> None: """Extract user preferences and facts asynchronously after chat turn completes.""" payload = envelope.payload user_id = payload.get("user_id") user_message = payload.get("user_message") assistant_message = payload.get("assistant_message") if not user_id or not user_message or not assistant_message: return # 1. Fetch current memory current_memory = await repository.get_user_memory(user_id) # 2. Run summarizer LLM prompt updated_memory = await extract_memory_delta( existing=current_memory, user_message=user_message, assistant_message=assistant_message, ) # 3. Upsert back to PostgreSQL if updated_memory and updated_memory != current_memory: await repository.upsert_user_memory(user_id, updated_memory)

Prompt Injection & Context Assembly

During a chat request in features/chat/pipeline.py, the stored memory is loaded and provided to the system prompt template:

<!-- prompts/chat_system.md --> You are OpenTier, an intelligent AI assistant. {% if user_memory %} <user_memory> {{ user_memory }} </user_memory> Use the user memory context above to tailor responses to the user's preferences, ongoing projects, and style without explicitly repeating it unless relevant. {% endif %}

Safety & Size Boundaries

  • Length Capping: Memory updates are capped to prevent context window explosion.
  • Deduplication: Sublinear extraction prompts ensure repetitive queries do not pollute memory with duplicated statements.
  • Isolation: Tenant filtering ensures user_id boundaries are strictly enforced across memory queries and storage.
Last updated on