Real-Time Voice Agent
Chatty embeds an enterprise WebRTC voice pipeline that turns your AI assistant into a speaking conversational phone and browser agent with human-like latency under 500ms.
Visitors can tap the Call button in the chat widget header or visit the full-screen /voice-demo portal to talk naturally, interrupt the bot at any moment (barge-in), and receive answers sourced from your knowledge base.
The WebRTC Audio Pipeline
Core Audio Capabilities
Instant Interruption (Barge-In)
Silero Voice Activity Detection (VAD) detects when the visitor starts speaking and cuts off the bot's speech output instantly, mirroring natural human conversation.
Knowledge Base Grounding
The voice worker has full access to the bot's vector embeddings, knowledge collections, and scheduling tools during the live call.
Sub-500ms Latency
Streamlined WebRTC audio frames combined with Cartesia Sonic streaming TTS produce ultra-low latency voice output.
Live Call Transcripts
Full audio transcripts and speaker diarization are logged in real-time to the Chatty Live Inbox for agent review.
Generating Client Room Tokens
To initiate a voice call from your frontend or custom mobile application, request a secure LiveKit connection token from the Chatty API:
curl -X POST https://api.chatty.personaliai.com/api/voice/token \
-H "Authorization: Bearer YOUR_BOT_ID" \
-H "Content-Type: application/json" \
-d '{
"bot_id": "YOUR_BOT_ID",
"visitor_id": "visitor_user_123"
}'Response
{
"token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
"server_url": "wss://livekit.chatty.personaliai.com",
"room_name": "call_bot_123_visitor_user_123"
}Connect to the server_url using LiveKit Client SDKs (livekit-client in JavaScript/TypeScript, Swift for iOS, Kotlin for Android).