C
Chatty

📊 Chatbot Performance Scorecard

Most AI evaluations measure raw model benchmarks: token generation speed, perplexity, or MMLU scores. They fail to answer the most crucial business question:

"Is your AI chatbot actually converting visitors and solving customer problems reliably?"

The Chatty Performance Scorecard grades chatbots on an A+ to F scale across five core operational pillars, providing engineering, compliance, and marketing teams with a complete operational health assessment in under 60 seconds.


🏆 The Five Operational Pillars

target

1. Conversion & Revenue Intent (25%)

Measures lead capture rate, meeting booking rate, and proactive campaign CTRs. Target: ≥ 15% session-to-lead conversion.

message-square

2. Engagement & Deflection (20%)

Measures autonomous resolution without human intervention or ticket escalation. Target: ≥ 85% deflection rate.

shield-check

3. Accuracy & Grounding (25%)

Measures RAG precision, anti-hallucination guardrails, and valid tool executions. Target: ≥ 98% execution consistency.

zap

4. Reliability & Latency (15%)

Measures p95 first-token latency, uptime, and fallback resilience. Target: p95 < 1,500ms for chat, < 500ms for voice.

smile

5. Customer Satisfaction & CSAT (15%)

Measures verified post-chat 1–5 star ratings and positive sentiment balance. Target: ≥ 4.5/5.0 average CSAT score.


📈 Letter Grade Scale

GradeComposite ScoreOperational StatusInterpretation
A+93 – 100%ExceptionalWorld-class conversion and autonomous deflection. Scale ready.
A90 – 92.9%ExcellentStrong lead generation, low hallucination risk, high customer satisfaction.
A-87 – 89.9%Very GoodSolid business metrics with minor optimization opportunities in edge cases.
B+83 – 86.9%GoodReliable everyday performance; review unhandled queries to boost deflection.
B80 – 82.9%CompetentFunctional customer support bot; needs proactive campaign tuning.
C65 – 74.9%Needs AttentionElevated human escalation rate or noticeable latency bottlenecks.
D55 – 64.9%PoorHigh error rates or lack of relevant knowledge documents.
F< 55%FailingCritical misconfiguration; bot is causing customer friction.

🛠️ Three Ways to Run an Audit

1. Terminal CLI (Instant Evaluation)

Chatty includes a standalone auditor script that prints a rich terminal scorecard:

# Run with synthetic benchmark sample:
python scripts/audit_bot.py --sample
 
# Audit your live bot instance:
python scripts/audit_bot.py --bot-id <YOUR_BOT_UUID> --days 30
 
# Export markdown report to file:
python scripts/audit_bot.py --bot-id <YOUR_BOT_UUID> --format markdown --output audit-report.md

2. Dashboard REST API

Integrate scorecard reporting into internal dashboards or CI/CD pipelines:

GET /api/admin/analytics/scorecard?bot_id=YOUR_BOT_ID&days=30
Authorization: Bearer <SUPABASE_JWT_OR_API_KEY>

3. Model Context Protocol (MCP) Inside Claude or Cursor

Ask your AI coding assistant:

"Run an audit on my bot 8f9024b1-e25c-4122-8d77-a82a6fce921b and explain the scorecard."

The MCP client invokes get_bot_performance_scorecard or reads chatty://bots/{bot_id}/scorecard automatically.


🚀 Optimization Playbook