The Green Dashboard Trap: Why Agent Evals Pass but Users Still Fail

Hack Session

About the session

Your agent scores 90% on its offline benchmark—but what happens when the router selects the wrong tool, an API times out, context is missing, or the user’s language differs from your test set? In this live hack session, we will take an apparently production-ready tool-using agent and systematically break it using realistic environment perturbations. We will move beyond final-answer scores to diagnose routing, trajectory, state, recovery, latency, and cost failures. Finally, we will turn failed production-like traces into targeted regression datasets and evaluate an improved policy without blindly experimenting on users. Attendees will leave with a practical blueprint for closing the loop between offline evals, production telemetry, and safer agent improvement.
Audience takeaway
A reusable method for finding out why an agent fails and not merely whether its final answer received a good score.

Speaker

Download Brochure