Skip to Content
Runtime FlowChat Message End-to-End

Chat Message: End-to-End Runtime Flow

This page traces a single streaming chat turn from the user submission to the final rendered token on screen, detailing the gRPC streaming bridge, Qdrant hybrid retrieval, and asynchronous background event offloading via Redis Streams.


Full Sequence Diagram


Storage & Network Boundaries Per Request

In v1.1.0, heavy compute and writes are strictly partitioned:

TargetOperations Per TurnSynchronous / AsyncPurpose
PostgreSQL 184 queriesSynchronous1 session auth + 1 conv check + 1 user msg insert + 1 assistant msg insert.
Qdrant1 vector querySynchronousSub-millisecond dense (3072d) + sparse (BM25) server-side RRF hybrid retrieval.
Redis 8 Streams1 XADDSynchronous fire-and-forgetEnqueues chat.completed.v1 event; terminates gRPC obligation.
Background Workers2 consumer groupsAsynchronousOffloaded credit deduction & LLM memory extraction execute outside user turn latency.

Latency Budget & TTFT

StageTypical LatencyP99 LatencyArchitecture Driver
Browser → Rust Gateway<1 ms5 msZero-cost Axum routing.
Session & Ownership Lookup1–3 ms8 msSingle-query composite index on sessions + conversations.
gRPC Channel dispatch<0.5 ms2 msHigh-throughput HTTP/2 multiplexing via tonic.
In-memory Prompt & User Memory1–3 ms6 msIndexed lookup on user_memories.
Qdrant Hybrid Retrieval8–18 ms45 msServer-side RRF fusion over HNSW & sparse indices.
LLM Time to First Token (TTFT)150–600 ms1500 msExternal provider latency (Gemini, OpenAI).
Total Time to First Token~165–630 ms~1560 msUser experiences sub-second response start.
Token Streaming Cadence15–30 ms/token80 ms/tokenBuffered SSE chunking over HTTP/2.
Post-Stream Cleanup & SSE Close3–6 ms12 msDB insert + Redis event dispatch.

Failure Recovery & Resilience

  1. Mid-Stream Provider Interruption:
    • If the LLM provider cuts off mid-generation, Python catches the connection drop, emits a terminal ChatStreamChunk(error="upstream_timeout") with is_final=true, and commits whatever partial assistant message was received.
    • The browser receives an SSE error notification, allowing the user to regenerate the message without losing the session.
  2. Worker Event Retries & DLQ:
    • If the background worker fails to deduct credits or extract memory (e.g. database lock), Redis Streams retries up to 3 times before routing the message to opentier:intel:chat_events:dlq. The chat turn itself is never failed or stalled.
Last updated on