Voice Transcription
/api/widget/transcribeTranscribe recorded voice clips into clean, punctuated text. The audio is processed server-side using Google Gemini's native audio understanding pipeline, supporting multi-language speech recognition, technical terminology, and automatic noise suppression.
Endpoint Specifications
- Content-Type:
multipart/form-data - Supported Audio Formats:
audio/webm,audio/wav,audio/mpeg(MP3),audio/ogg,audio/x-m4a - Maximum Audio Duration: Up to 3 minutes (max 25 MB)
Request Form Fields
filefilebodyrequiredAudio recording file or binary blob from the visitor's microphone.
bot_idstringbodyBot identifier. When provided, the transcriber applies vocabulary hints derived from the bot's custom knowledge base and primary language settings.
languagestringbodydefault: autoISO 639-1 language code hint (e.g. "en", "es", "fr", "de"). Defaults to "auto" for automatic language identification.
Response
textstringThe transcribed text, with capitalization and punctuation.
duration_secondsnumberDetected audio length in seconds.
languagestringDetected or confirmed language code.
Browser Microphone Pipeline
Here is a complete browser implementation that captures microphone audio via the HTML5 MediaRecorder API and posts it directly to /api/widget/transcribe:
class ChattyVoiceRecorder {
private mediaRecorder: MediaRecorder | null = null;
private audioChunks: Blob[] = [];
async startRecording(): Promise<void> {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
this.audioChunks = [];
this.mediaRecorder = new MediaRecorder(stream, { mimeType: "audio/webm" });
this.mediaRecorder.ondataavailable = (event) => {
if (event.data.size > 0) {
this.audioChunks.push(event.data);
}
};
this.mediaRecorder.start();
}
async stopAndTranscribe(botId: string): Promise<string> {
return new Promise((resolve, reject) => {
if (!this.mediaRecorder) {
return reject(new Error("Recorder not initialized"));
}
this.mediaRecorder.onstop = async () => {
const audioBlob = new Blob(this.audioChunks, { type: "audio/webm" });
const formData = new FormData();
formData.append("file", audioBlob, "voice_input.webm");
formData.append("bot_id", botId);
try {
const res = await fetch("https://api.chatty.personaliai.com/api/widget/transcribe", {
method: "POST",
body: formData,
});
if (!res.ok) {
throw new Error(`Transcription failed with status ${res.status}`);
}
const data = await res.json();
resolve(data.text);
} catch (err) {
reject(err);
} finally {
// Clean up microphone tracks
this.mediaRecorder?.stream.getTracks().forEach((track) => track.stop());
}
};
this.mediaRecorder.stop();
});
}
}Examples
cURL
curl -X POST https://api.chatty.personaliai.com/api/widget/transcribe \
-F "file=@/path/to/voice_note.webm;type=audio/webm" \
-F "bot_id=YOUR_BOT_ID" \
-F "language=en"200 OK Response:
{
"text": "Could you please schedule a demonstration for Tuesday afternoon at three?",
"duration_seconds": 4.85,
"language": "en"
}Once transcribed, the resulting text string can be immediately submitted to /api/widget/chat/stream or /api/v1/chat.