C
Chatty

Real-Time Voice Agent

Chatty embeds an enterprise WebRTC voice pipeline that turns your AI assistant into a speaking conversational phone and browser agent with human-like latency under 500ms.

Visitors can tap the Call button in the chat widget header or visit the full-screen /voice-demo portal to talk naturally, interrupt the bot at any moment (barge-in), and receive answers sourced from your knowledge base.


The WebRTC Audio Pipeline


Core Audio Capabilities

⚡

Instant Interruption (Barge-In)

Silero Voice Activity Detection (VAD) detects when the visitor starts speaking and cuts off the bot's speech output instantly, mirroring natural human conversation.

🧠

Knowledge Base Grounding

The voice worker has full access to the bot's vector embeddings, knowledge collections, and scheduling tools during the live call.

⏱️

Sub-500ms Latency

Streamlined WebRTC audio frames combined with Cartesia Sonic streaming TTS produce ultra-low latency voice output.

📝

Live Call Transcripts

Full audio transcripts and speaker diarization are logged in real-time to the Chatty Live Inbox for agent review.


Generating Client Room Tokens

To initiate a voice call from your frontend or custom mobile application, request a secure LiveKit connection token from the Chatty API:

curl -X POST https://api.chatty.personaliai.com/api/voice/token \
  -H "Authorization: Bearer YOUR_BOT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "bot_id": "YOUR_BOT_ID",
    "visitor_id": "visitor_user_123"
  }'

Response

{
  "token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
  "server_url": "wss://livekit.chatty.personaliai.com",
  "room_name": "call_bot_123_visitor_user_123"
}

Connect to the server_url using LiveKit Client SDKs (livekit-client in JavaScript/TypeScript, Swift for iOS, Kotlin for Android).