handling 10 million agent traces a day
an engineering session on storing and querying agent traces at 10 million a day: the write path, the schema, what it costs to keep, and the designs that fell over first.
an agent trace isn't a log line. one run can have dozens of nested tool calls, long model outputs, retries and branches, and you need the whole session back in one piece when something goes wrong. at a few thousand runs a day anything works. at 10 million, most of the obvious choices break.

what it's about
an agent trace isn't a log line. one run can have dozens of nested tool calls, long model outputs, retries and branches, and you need to pull the whole session back in one piece when something goes wrong. at a few thousand runs a day, anything works. at 10 million traces a day, most of the obvious choices break.
sidd, founding engineer at failproof ai, walks through how we store and query agent traces at that volume, and what we changed to get there. this is an engineering session, not a product demo.
- 01what makes agent traces different
why traces are wider, deeper and more uneven than regular observability data, and what that does to a storage design.
- 02the write path
how traces come in, how we batch them, and how we keep ingestion from falling behind when traffic spikes.
- 03schema and layout
how we lay out spans so that "show me this whole session" and "find every failed tool call this week" are both fast.
- 04keeping cost sane
what we keep hot, what we move to cheaper storage, and where the money actually goes at this volume.
- 05what broke along the way
the designs we tried first, where they fell over, and what we'd do differently.
- 06your questions, live
the last 20 minutes are open. bring your own tracing setup and the query that's too slow.
how the day runs
all times pdt · subject to change- 15 minwhat agent traces look like at scaletalk
why one run is wider, deeper and more uneven than a log line, and what that does to a storage design.
- 25 minwalkthrough of the storage design, with real numberstalk
the write path and batching, how spans are laid out so whole-session and needle queries are both fast, and what we keep hot versus what moves to cheaper storage.
- 20 minopen q&asocial
bring your own tracing setup and the query that's too slow.
who's on
- HOSTSahar Morbond ai
questions
something else? ask on discord →do i need to use failproof to get anything out of this?
no. the session is about storing and querying agent traces at volume, and failproof is the system we hit the problem on.
the write path, the span layout and the hot-versus-cold split are the same questions on any trace store.
what will i actually see?
the storage design and real numbers, not a slide of a reference architecture.
how traces come in and get batched, how spans are laid out so "show me this whole session" and "find every failed tool call this week" are both fast, and where the money goes.
can i bring my own problem?
yes. the last 20 minutes are open, so bring your tracing setup and the query that is too slow.
the q&a starts at 6:40 pm pt.
how does registration work?
register on luma. it is free, and you are in straight away — no host approval to wait on.
it is online, so luma sends the joining link with your registration.
what is failproof ai?
failproof is your agent's oversight layer, steering it towards success. it traces your agents, finds failure patterns over time and sets up rules to prevent them from happening again.
the open-source cli is free: npm i -g failproofai.
MORE EVENTS
all events →- OCT06how we steer hermesengineering session · san francisco
- OCT24steering agents at runtimepaper discussion · bengaluru
- OCT27steering agents with policiesengineering session · san francisco