📊 Chatbot Performance Scorecard
Most AI evaluations measure raw model benchmarks: token generation speed, perplexity, or MMLU scores. They fail to answer the most crucial business question:
"Is your AI chatbot actually converting visitors and solving customer problems reliably?"
The Chatty Performance Scorecard grades chatbots on an A+ to F scale across five core operational pillars, providing engineering, compliance, and marketing teams with a complete operational health assessment in under 60 seconds.
🏆 The Five Operational Pillars
1. Conversion & Revenue Intent (25%)
Measures lead capture rate, meeting booking rate, and proactive campaign CTRs. Target: ≥ 15% session-to-lead conversion.
2. Engagement & Deflection (20%)
Measures autonomous resolution without human intervention or ticket escalation. Target: ≥ 85% deflection rate.
3. Accuracy & Grounding (25%)
Measures RAG precision, anti-hallucination guardrails, and valid tool executions. Target: ≥ 98% execution consistency.
4. Reliability & Latency (15%)
Measures p95 first-token latency, uptime, and fallback resilience. Target: p95 < 1,500ms for chat, < 500ms for voice.
5. Customer Satisfaction & CSAT (15%)
Measures verified post-chat 1–5 star ratings and positive sentiment balance. Target: ≥ 4.5/5.0 average CSAT score.
📈 Letter Grade Scale
| Grade | Composite Score | Operational Status | Interpretation |
|---|---|---|---|
| A+ | 93 – 100% | Exceptional | World-class conversion and autonomous deflection. Scale ready. |
| A | 90 – 92.9% | Excellent | Strong lead generation, low hallucination risk, high customer satisfaction. |
| A- | 87 – 89.9% | Very Good | Solid business metrics with minor optimization opportunities in edge cases. |
| B+ | 83 – 86.9% | Good | Reliable everyday performance; review unhandled queries to boost deflection. |
| B | 80 – 82.9% | Competent | Functional customer support bot; needs proactive campaign tuning. |
| C | 65 – 74.9% | Needs Attention | Elevated human escalation rate or noticeable latency bottlenecks. |
| D | 55 – 64.9% | Poor | High error rates or lack of relevant knowledge documents. |
| F | < 55% | Failing | Critical misconfiguration; bot is causing customer friction. |
🛠️ Three Ways to Run an Audit
1. Terminal CLI (Instant Evaluation)
Chatty includes a standalone auditor script that prints a rich terminal scorecard:
# Run with synthetic benchmark sample:
python scripts/audit_bot.py --sample
# Audit your live bot instance:
python scripts/audit_bot.py --bot-id <YOUR_BOT_UUID> --days 30
# Export markdown report to file:
python scripts/audit_bot.py --bot-id <YOUR_BOT_UUID> --format markdown --output audit-report.md2. Dashboard REST API
Integrate scorecard reporting into internal dashboards or CI/CD pipelines:
GET /api/admin/analytics/scorecard?bot_id=YOUR_BOT_ID&days=30
Authorization: Bearer <SUPABASE_JWT_OR_API_KEY>3. Model Context Protocol (MCP) Inside Claude or Cursor
Ask your AI coding assistant:
"Run an audit on my bot
8f9024b1-e25c-4122-8d77-a82a6fce921band explain the scorecard."
The MCP client invokes get_bot_performance_scorecard or reads chatty://bots/{bot_id}/scorecard automatically.