An eval harness you trust
How to build offline and online evals that actually catch regressions — golden sets, LLM-as-judge done carefully, and the metrics that correlate with real user pain.
$ Free live webinar · Thursday, September 24, 2026 · 11:00 AM PT · Online
A free 90-minute live session for engineers shipping AI agents past the demo stage. We go deep on the three things that decide whether an agent survives real traffic: evaluation you can trust, guardrails that hold, and cost that stays under control. Live demo, live Q&A, recording sent to everyone who registers.

Agenda
All times Pacific. We start on the hour, keep the talks tight, and leave a real block for your questions. The whole thing wraps in 90 minutes.
| Time | Segment | Presenter | Format |
|---|---|---|---|
| 11:00 | Welcome & the state of agents in production | Alex Rivera | Intro |
| 11:08 | Why agents fail: five failure modes we see in the wild | Alex Rivera | Talk |
| 11:25 | Building an evaluation harness you can actually trust | Priya Nair | Talk |
| 11:45 | Live demo: adding guardrails & tool sandboxing to a real agent | Daniel Kim | Demo |
| 12:05 | Controlling cost & latency at scale — routing, caching, budgets | Priya Nair | Talk |
| 12:18 | Live Q&A: bring your hardest production questions | All presenters | Q&A |
| 12:28 | Wrap-up, resources & recording details | Alex Rivera | Close |
What you'll take away
No slideware theater. Every segment maps to a decision you have to make when an agent leaves the demo and starts handling real users.
How to build offline and online evals that actually catch regressions — golden sets, LLM-as-judge done carefully, and the metrics that correlate with real user pain.
Input and output validation, tool sandboxing, and permission layers that stop an agent from doing damage — shown live on a working agent, not on slides.
Model routing, prompt caching, and per-request budgets that cut spend 40-70% without gutting quality. Real numbers from real production traffic.
The five ways agents break in production — runaway loops, silent tool errors, context bloat, eval drift, and cost blowups — and how each team we work with fixes them.
Everyone who registers gets the full recording, the slide deck, and a starter repo with the eval and guardrail patterns from the demo — sent within 24 hours.
A real Q&A block where the presenters take your hardest production questions on camera. Bring the problem you're stuck on right now.
Your presenters
Three engineers who ship agents to production every week — not analysts talking about it from the sidelines.

Host · Co-founder, Runtime
Has shipped LLM products since the GPT-3 beta and now leads applied AI at Runtime. Your host for the session — keeping the pace tight and the demos honest.

Presenter · Staff AI Engineer
Builds the evaluation and cost-control tooling behind agents serving millions of requests a day. She will walk through the eval harness and the routing tricks that keep bills sane.

Presenter · Principal Engineer, Platform
Runs the sandboxing and guardrail layer that lets agents call real tools without breaking things. He drives the live demo, wiring guardrails into a working agent on screen.
The essentials
It's online, it's free, and the replay is yours whether or not you can make the live time.
Thursday, September 24, 2026, 11:00 AM Pacific (2:00 PM Eastern / 6:00 PM UTC). Runs 90 minutes, including live Q&A. Can't make it? Check the event page or contact the organizer about replay access.
Online, live on Zoom. Registration is stored for organizer review; this template does not email join links, calendar invites, or reminders automatically.
Engineers, ML/AI leads, and technical founders who already have an agent or LLM feature running — or about to ship one — and want it to hold up under real traffic. Free to attend.
Register once to store your request. Check this page or contact the organizer for join and replay details; attendee email is not automatic.