Memory Subsystem
Overview
In OpenTier v1.1.0, user memory is a persistent, personalized memory layer that spans across conversations while being extracted asynchronously via Redis Streams.
Memory extraction is intentionally decoupled from the hot streaming chat path:
- During chat, the user’s existing memory is fetched with a quick indexed lookup and injected into the prompt.
- Upon completion of a chat turn, the intelligence service publishes an event to Redis Streams.
- A background worker picks up the event, executes LLM-driven memory extraction, and atomically upserts the memory ledger.
This architecture ensures zero latency impact on Time-to-First-Token (TTFT) and token streaming throughput.
Data Model (user_memories)
The transactional system of record for memory is PostgreSQL 18:
CREATE TABLE IF NOT EXISTS user_memories (
user_id VARCHAR(255) PRIMARY KEY,
memory TEXT NOT NULL DEFAULT '',
metadata JSONB NOT NULL DEFAULT '{}',
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX IF NOT EXISTS idx_user_memories_updated_at ON user_memories(updated_at);Schema Attributes
| Column | Type | Description |
|---|---|---|
user_id | VARCHAR(255) | Primary key identifying the user. |
memory | TEXT | Compact, synthesized text representation of user preferences, context, and facts. |
metadata | JSONB | Structured metadata (e.g. extraction count, model used, topic tags). |
updated_at | TIMESTAMPTZ | Timestamp of the most recent memory update. |
Asynchronous Memory Extraction Flow
Worker Consumer Handler (features/chat/memory.py)
The worker registers consumer group memory listening to stream Streams.CHAT_EVENTS:
# features/chat/memory.py
async def handle_memory_extraction(envelope: EventEnvelope, msg_id: str) -> None:
"""Extract user preferences and facts asynchronously after chat turn completes."""
payload = envelope.payload
user_id = payload.get("user_id")
user_message = payload.get("user_message")
assistant_message = payload.get("assistant_message")
if not user_id or not user_message or not assistant_message:
return
# 1. Fetch current memory
current_memory = await repository.get_user_memory(user_id)
# 2. Run summarizer LLM prompt
updated_memory = await extract_memory_delta(
existing=current_memory,
user_message=user_message,
assistant_message=assistant_message,
)
# 3. Upsert back to PostgreSQL
if updated_memory and updated_memory != current_memory:
await repository.upsert_user_memory(user_id, updated_memory)Prompt Injection & Context Assembly
During a chat request in features/chat/pipeline.py, the stored memory is loaded and provided to the system prompt template:
<!-- prompts/chat_system.md -->
You are OpenTier, an intelligent AI assistant.
{% if user_memory %}
<user_memory>
{{ user_memory }}
</user_memory>
Use the user memory context above to tailor responses to the user's preferences,
ongoing projects, and style without explicitly repeating it unless relevant.
{% endif %}Safety & Size Boundaries
- Length Capping: Memory updates are capped to prevent context window explosion.
- Deduplication: Sublinear extraction prompts ensure repetitive queries do not pollute memory with duplicated statements.
- Isolation: Tenant filtering ensures
user_idboundaries are strictly enforced across memory queries and storage.
Last updated on