Chat Message: End-to-End Runtime Flow
This page traces a single streaming chat turn from the user submission to the final rendered token on screen, detailing the gRPC streaming bridge, Qdrant hybrid retrieval, and asynchronous background event offloading via Redis Streams.
Full Sequence Diagram
Storage & Network Boundaries Per Request
In v1.1.0, heavy compute and writes are strictly partitioned:
| Target | Operations Per Turn | Synchronous / Async | Purpose |
|---|---|---|---|
| PostgreSQL 18 | 4 queries | Synchronous | 1 session auth + 1 conv check + 1 user msg insert + 1 assistant msg insert. |
| Qdrant | 1 vector query | Synchronous | Sub-millisecond dense (3072d) + sparse (BM25) server-side RRF hybrid retrieval. |
| Redis 8 Streams | 1 XADD | Synchronous fire-and-forget | Enqueues chat.completed.v1 event; terminates gRPC obligation. |
| Background Workers | 2 consumer groups | Asynchronous | Offloaded credit deduction & LLM memory extraction execute outside user turn latency. |
Latency Budget & TTFT
| Stage | Typical Latency | P99 Latency | Architecture Driver |
|---|---|---|---|
| Browser → Rust Gateway | <1 ms | 5 ms | Zero-cost Axum routing. |
| Session & Ownership Lookup | 1–3 ms | 8 ms | Single-query composite index on sessions + conversations. |
| gRPC Channel dispatch | <0.5 ms | 2 ms | High-throughput HTTP/2 multiplexing via tonic. |
| In-memory Prompt & User Memory | 1–3 ms | 6 ms | Indexed lookup on user_memories. |
| Qdrant Hybrid Retrieval | 8–18 ms | 45 ms | Server-side RRF fusion over HNSW & sparse indices. |
| LLM Time to First Token (TTFT) | 150–600 ms | 1500 ms | External provider latency (Gemini, OpenAI). |
| Total Time to First Token | ~165–630 ms | ~1560 ms | User experiences sub-second response start. |
| Token Streaming Cadence | 15–30 ms/token | 80 ms/token | Buffered SSE chunking over HTTP/2. |
| Post-Stream Cleanup & SSE Close | 3–6 ms | 12 ms | DB insert + Redis event dispatch. |
Failure Recovery & Resilience
- Mid-Stream Provider Interruption:
- If the LLM provider cuts off mid-generation,
Pythoncatches the connection drop, emits a terminalChatStreamChunk(error="upstream_timeout")withis_final=true, and commits whatever partial assistant message was received. - The browser receives an SSE error notification, allowing the user to regenerate the message without losing the session.
- If the LLM provider cuts off mid-generation,
- Worker Event Retries & DLQ:
- If the background worker fails to deduct credits or extract memory (e.g. database lock), Redis Streams retries up to 3 times before routing the message to
opentier:intel:chat_events:dlq. The chat turn itself is never failed or stalled.
- If the background worker fails to deduct credits or extract memory (e.g. database lock), Redis Streams retries up to 3 times before routing the message to
Last updated on