steering agents at runtime
a paper discussion: two hours on why agents fail and what you can do about it without touching the model. real failed runs, on screen.
we ran coding agents on t-bench 2.1 and read every failed run. most failures were not the model lacking skill: the agent was missing one fact about its environment. short runtime policies that catch those moments took success from 37% to 69%.

what it's about
two hours of nerding out about why agents fail and what you can do about it without touching the model.
we ran coding agents on t-bench 2.1 and read every failed run. most failures were not the model lacking skill. the agent was missing one fact about its environment: a server that dies when the shell closes, a damaged database opened before it was copied, an output that was right except for one byte. we wrote short runtime policies that catch those moments and tell the agent what it is missing. on tasks where a policy recognized the failure, success went from 37% to 69%.
that is the paper: fire: failure-informed runtime engineering for reliable language-model agents (arxiv:2609.26048). this is a discussion, not a demo day.
- 01real failed runs, up close
actual agent traces on screen, walked through step by step. the failures are more specific and more fixable than most people expect.
- 02the policies themselves
not a diagram of them. the actual code, what triggers each one, and why the exact wording of the instruction mattered more than when it fired.
- 03the parts we are unsure about
our text matching recognized only 2 of 18 reworded tasks. one policy set raised cost by almost 48%. one comparison did not clear significance. we want your take on all three.
- 04your own agent's failures
bring a failure mode you keep hitting. at the tables we work through whether a runtime policy could catch it, and what it would need to see.
- 05a panel that is asked to disagree
researchers who have read the paper in advance review it in front of the room, and then it is open floor.
- 06everything to rerun it
the policies, matching rules and task list go out at the event, so you can try it on your own agents the following week.
how the day runs
all times ist · subject to change- 15 minchai, find a tablearrive
arrive, get a chai, and get put on a table.
- 20 minwalkthrough of the method, traces and resultstalk
where each failed run went wrong, the policies written for those moments, and how success went from 37% to 69%.
- 30 mintable discussionsbuild
bring a failure mode you keep hitting. the table works out whether a runtime policy could catch it, and what it would need to see.
- 30 minpanel review and open floortalk
researchers who read the paper in advance review it in front of the room, then anyone can jump in.
- 25 minreplication kit, hang around and keep talkingsocial
the policies, matching rules and task list, so you can rerun it on your own agents.
who's on
getting there
questions
something else? ask on discord →do i need to have read the paper?
no. the 3:15 walkthrough covers the method, the traces and the results from the start.
if you want it beforehand: fire: failure-informed runtime engineering for reliable language-model agents, arxiv:2609.26048.
what should i bring?
a failure mode you keep hitting. the table discussions at 3:35 work through whether a runtime policy could catch it, and what it would need to see.
a laptop helps if you want to pull up a trace of your own.
is this a talk or a discussion?
a discussion, not a demo day. real failed runs on screen, the actual policy code, and the three results we are still unsure about.
the panel is asked to disagree with us, and the floor opens after it at 4:05 pm.
do i get the policies?
yes. the replication kit goes out at the event: policies, matching rules and the task list, so you can rerun it on your own agents the following week.
how does registration work?
register on luma. every registration is subject to host approval.
the exact address in hsr layout is shared once you're in.
MORE EVENTS
all events →- OCT06how we steer hermesengineering session · san francisco
- SEP27jev buildathonbuildathon · bengaluru · past
- SEP20jev colearnco-learning · bengaluru · past