eval-chat-traces
Installation
SKILL.md
Eval Chat Traces
Systematically review AI chat traces to find failure patterns using error analysis methodology (inspired by Hamel Husain's evals framework).
Input
Optional: time range (default: last 7 days), sample size (default: 30 traces).
Process
Phase 1: Pull Traces
- Query CloudWatch Logs Insights on
/aws/lambda/extralife-ai-engagement-agentforChatQuerylogs. - Query for
ChatFeedbacklogs from/ecs/ExtraLifeWebAdmin. - Prioritize: negative feedback traces first, then outliers (response_time_ms > P90, empty tools_invoked, error responses), then random sample to fill remaining.
- For Athena deep dives:
SELECT * FROM extralife_chat_logs.chat_logs WHERE ...
Phase 2: Present Traces for Review
Present traces in a structured table for efficient binary judgment: