Platform Architecture
Chatty is a multi-tenant AI customer-support platform with a Next.js frontend, FastAPI application services, Supabase-backed persistence, streaming chat, RAG, voice, booking, omnichannel webhooks, and MCP access. The managed-Supabase profile is the default for both Chatty Cloud and self-hosted application containers.
This document breaks down the request flow, data boundaries, and deployment security model without implying that a self-hosted deployment must replace Supabase or use a specific cloud vendor.
High-Level Topology
Chatty follows a two-service application boundary: the public frontend calls the API over HTTPS, and the API talks to managed Supabase plus optional providers. The same containers can run on a VPS, Railway, Render, or a Heroku-style container host:
Core Infrastructure Layers
Asynchronous ASGI Engine
Built on FastAPI and Uvicorn inside the API container. Every network I/O call (database queries, vector similarity, optional Redis counters, LLM tokens) is non-blocking.
Vector RAG Pipeline
Powered by Supabase PostgreSQL with pgvector. Uses 768-dimensional embeddings generated by Google's text-embedding-004 model with indexed HNSW cosine distance search.
Persistent Edge Rate Limiting
Optional distributed rate limiting and ephemeral state backed by Upstash Redis. The managed-Supabase profile works without provisioning a local Redis service.
Multi-Channel Adapters
Native support for the Embed Widget, React SDK, WordPress plugin, WhatsApp Business API, and Model Context Protocol (MCP).
The Request Lifecycle
When a visitor interacts with a Chatty widget or API client, the request traverses multiple verification and processing stages:
1. Ingress & Origin Allowlist
Public widget endpoints check the incoming HTTP Origin header against the bot's configured domain allowlist. If a domain is not authorized, the backend immediately terminates the connection with 403 Forbidden and injects Content-Security-Policy: frame-ancestors 'none' to block unauthorized iframe embedding.
2. Distributed rate limiting (optional)
To protect backend workers from volumetric abuse, deployments can use the Upstash Redis token bucket integration. Counters are partitioned by client IP (for widget traffic) or SHA-256 API key hash (for developer API calls). The managed-Supabase default still runs without a locally provisioned Redis.
3. Session Resolution & Human Takeover Check
The conversation session UUID is retrieved from PostgreSQL. If the session has ai_paused: true (e.g. a customer support agent initiated a human takeover), the AI pipeline is bypassed, and the visitor message is routed directly to the Live Inbox dashboard.
4. Hybrid Knowledge Retrieval (RAG)
If knowledge sources exist, Chatty performs hybrid retrieval:
- Semantic Vector Search: Computes cosine distance against
knowledge_chunksusingpgvectorembeddings (<=>operator). - Keyword Filtering: Matches high-confidence lexical tokens.
- Top results are deduplicated, reranked, and formatted with document metadata (source URL, file title, page number) into the LLM system prompt.
5. Non-Blocking SSE Streaming
The orchestrator dispatches the assembled context to Google Gemini with streaming enabled. Token generators yield text/event-stream chunks directly through FastAPI's StreamingResponse, achieving Time-To-First-Token (TTFT) under 400ms.
Concurrency & Database Pooling
Database connection exhaustion is the most common failure point in AI applications that handle long-running LLM streams. Chatty mitigates this with a robust pooling architecture:
[API container instances]
โ
โโโ Worker 1 โโโ
โโโ Worker 2 โโโผโโ> [Async SQLAlchemy Engine]
โโโ Worker N โโโ โ
โผ
[Application connection pool]
- pool_size: 20
- max_overflow: 44
- pool_timeout: 30s
- pool_recycle: 1800s
- pool_pre_ping: true
โ
โผ
[Supabase Transaction Pooler]
(Port 6543)
Security Boundaries & Isolation
Chatty enforces enterprise security boundaries at every level of the stack:
| Boundary | Mechanism | Purpose |
|---|---|---|
| Tenant Isolation | PostgreSQL Row-Level Security (RLS) | Restricts database records strictly to the authenticated user_id or bot owner. |
| SSRF Prevention | Pre-flight DNS resolution & IP pinning | Prevents Server-Side Request Forgery against cloud metadata (169.254.169.254) and private RFC-1918 subnets. |
| Upload Safety | Stream-capped chunked readers (20MB) | Enforces maximum upload byte caps before writing to disk, preventing memory exhaustion and DoS. |
| JWT Verification | Asymmetric JWKS public key validation | Verifies Supabase authentication tokens cryptographically with cached public key rotation. |
| BYOK Encryption | AES-256-GCM symmetric encryption | Encrypts customer-supplied API keys (OpenAI, Anthropic, Gemini) before database persistence. |
For detailed security specifications, vulnerability reporting, and GDPR compliance, visit our Security & Privacy reference.