← all events / steering-agents-at-runtimeshare x · linkedin ·
● registrations openpaper discussionin-person

steering agents at runtime

a paper discussion: two hours on why agents fail and what you can do about it without touching the model. real failed runs, on screen.

we ran coding agents on t-bench 2.1 and read every failed run. most failures were not the model lacking skill: the agent was missing one fact about its environment. short runtime policies that catch those moments took success from 37% to 69%.

steering agents at runtime poster

what it's about

two hours of nerding out about why agents fail and what you can do about it without touching the model.

we ran coding agents on t-bench 2.1 and read every failed run. most failures were not the model lacking skill. the agent was missing one fact about its environment: a server that dies when the shell closes, a damaged database opened before it was copied, an output that was right except for one byte. we wrote short runtime policies that catch those moments and tell the agent what it is missing. on tasks where a policy recognized the failure, success went from 37% to 69%.

that is the paper: fire: failure-informed runtime engineering for reliable language-model agents (arxiv:2609.26048). this is a discussion, not a demo day.

  1. 01real failed runs, up close

    actual agent traces on screen, walked through step by step. the failures are more specific and more fixable than most people expect.

  2. 02the policies themselves

    not a diagram of them. the actual code, what triggers each one, and why the exact wording of the instruction mattered more than when it fired.

  3. 03the parts we are unsure about

    our text matching recognized only 2 of 18 reworded tasks. one policy set raised cost by almost 48%. one comparison did not clear significance. we want your take on all three.

  4. 04your own agent's failures

    bring a failure mode you keep hitting. at the tables we work through whether a runtime policy could catch it, and what it would need to see.

  5. 05a panel that is asked to disagree

    researchers who have read the paper in advance review it in front of the room, and then it is open floor.

  6. 06everything to rerun it

    the policies, matching rules and task list go out at the event, so you can try it on your own agents the following week.

good foragent buildersresearchersanyone debugging the same failure twice

how the day runs

all times ist · subject to change
  1. 15 min
    chai, find a table

    arrive, get a chai, and get put on a table.

    arrive
  2. 20 min
    walkthrough of the method, traces and results

    where each failed run went wrong, the policies written for those moments, and how success went from 37% to 69%.

    talk
  3. 30 min
    table discussions

    bring a failure mode you keep hitting. the table works out whether a runtime policy could catch it, and what it would need to see.

    build
  4. 30 min
    panel review and open floor

    researchers who read the paper in advance review it in front of the room, then anyone can jump in.

    talk
  5. 25 min
    replication kit, hang around and keep talking

    the policies, matching rules and task list, so you can rerun it on your own agents.

    social

who's on

  • HOSTNivedit Jaincto & co-founder @ failproof aixlinkedin
  • HOSTNikita Agarwalceo & co-founder @ failproof aixlinkedin

getting there

hsr layout
1st sector, hsr layout, bengaluru, karnataka 560102
open in maps ↗

questions

something else? ask on discord →
do i need to have read the paper?

no. the 3:15 walkthrough covers the method, the traces and the results from the start.

if you want it beforehand: fire: failure-informed runtime engineering for reliable language-model agents, arxiv:2609.26048.

what should i bring?

a failure mode you keep hitting. the table discussions at 3:35 work through whether a runtime policy could catch it, and what it would need to see.

a laptop helps if you want to pull up a trace of your own.

is this a talk or a discussion?

a discussion, not a demo day. real failed runs on screen, the actual policy code, and the three results we are still unsure about.

the panel is asked to disagree with us, and the floor opens after it at 4:05 pm.

do i get the policies?

yes. the replication kit goes out at the event: policies, matching rules and the task list, so you can rerun it on your own agents the following week.

how does registration work?

register on luma. every registration is subject to host approval.

the exact address in hsr layout is shared once you're in.

MORE EVENTS

all events →

your room.
our agents. one night.

contact us