← all events / handling-10-million-agent-tracesshare x · linkedin ·
● registrations openengineering sessiononline

handling 10 million agent traces a day

an engineering session on storing and querying agent traces at 10 million a day: the write path, the schema, what it costs to keep, and the designs that fell over first.

an agent trace isn't a log line. one run can have dozens of nested tool calls, long model outputs, retries and branches, and you need the whole session back in one piece when something goes wrong. at a few thousand runs a day anything works. at 10 million, most of the obvious choices break.

handling 10 million agent traces a day poster

what it's about

an agent trace isn't a log line. one run can have dozens of nested tool calls, long model outputs, retries and branches, and you need to pull the whole session back in one piece when something goes wrong. at a few thousand runs a day, anything works. at 10 million traces a day, most of the obvious choices break.

sidd, founding engineer at failproof ai, walks through how we store and query agent traces at that volume, and what we changed to get there. this is an engineering session, not a product demo.

  1. 01what makes agent traces different

    why traces are wider, deeper and more uneven than regular observability data, and what that does to a storage design.

  2. 02the write path

    how traces come in, how we batch them, and how we keep ingestion from falling behind when traffic spikes.

  3. 03schema and layout

    how we lay out spans so that "show me this whole session" and "find every failed tool call this week" are both fast.

  4. 04keeping cost sane

    what we keep hot, what we move to cheaper storage, and where the money actually goes at this volume.

  5. 05what broke along the way

    the designs we tried first, where they fell over, and what we'd do differently.

  6. 06your questions, live

    the last 20 minutes are open. bring your own tracing setup and the query that's too slow.

good forengineers building agent observabilityanyone running agents in productionanyone on high-volume data pipelines

how the day runs

all times pdt · subject to change
  1. 15 min
    what agent traces look like at scale

    why one run is wider, deeper and more uneven than a log line, and what that does to a storage design.

    talk
  2. 25 min
    walkthrough of the storage design, with real numbers

    the write path and batching, how spans are laid out so whole-session and needle queries are both fast, and what we keep hot versus what moves to cheaper storage.

    talk
  3. 20 min
    open q&a

    bring your own tracing setup and the query that's too slow.

    social

who's on

  • SPEAKERSiddartha A Yfounding engineer @ failproof ai

    built the trace store, and walks through it from the inside: the write path, the span layout the queries are shaped around, and the designs that fell over on the way to 10 million a day.

    xlinkedin
  • HOSTNivedit Jaincto & co-founder @ failproof aixlinkedin
  • HOSTNikita Agarwalceo & co-founder @ failproof aixlinkedin
  • HOSTSahar Morbond ai

questions

something else? ask on discord →
do i need to use failproof to get anything out of this?

no. the session is about storing and querying agent traces at volume, and failproof is the system we hit the problem on.

the write path, the span layout and the hot-versus-cold split are the same questions on any trace store.

what will i actually see?

the storage design and real numbers, not a slide of a reference architecture.

how traces come in and get batched, how spans are laid out so "show me this whole session" and "find every failed tool call this week" are both fast, and where the money goes.

can i bring my own problem?

yes. the last 20 minutes are open, so bring your tracing setup and the query that is too slow.

the q&a starts at 6:40 pm pt.

how does registration work?

register on luma. it is free, and you are in straight away — no host approval to wait on.

it is online, so luma sends the joining link with your registration.

what is failproof ai?

failproof is your agent's oversight layer, steering it towards success. it traces your agents, finds failure patterns over time and sets up rules to prevent them from happening again.

the open-source cli is free: npm i -g failproofai.

MORE EVENTS

all events →

your room.
our agents. one night.

contact us