jev buildathon
ship an agent that does the work. let jev evals and policies judge every move it makes.
bring an agent that does real work: support, payments, ops, coding or sales. write jev evals and policies for it, then demo the run where a verdict caught a bad action and steered it right. solo or in a team, three and a half hours, judged live.

what it's about
jev from typesafe ai returns typed verdicts (choices, scores, yes/no) instead of text. that makes it a sharp, cheap judge for an agent while it runs, not after. score every action, check it against a policy, and decide on the spot whether the agent proceeds, stops, or gets told what to do instead.
anyone can build a safe agent that summarizes docs. we want the agent you'd be nervous to ship, made trustworthy by jev.
- 01build an agent that does work
support, payments, ops, coding, sales. anything where a wrong action costs something.
- 02write jev evals and policies
the evals score every action it takes. the policies decide what happens next.
- 03demo it
show a run where a jev verdict caught a bad action and steered the agent to the right one.
how the winner is picked
to winthe highest-impact, highest-risk use case that runs reliably.
- stakes
how much damage could this agent do if it went wrong?
- reliability
does it hold up across runs, not just the one you rehearsed?
- the save
a clear moment where a jev verdict changed what the agent did.
ambitious and shaky loses to ambitious and solid.
how the day runs
all times ist · subject to change- 20 minkickofftalk
jev evals and policies walkthrough
- 1 h 55 minbuildbuild
ship the agent, then write its jev evals and policies. failproof runs them on live tool calls.
- 35 mindemostalk
show a run where a jev verdict caught a bad action and steered the agent right.
- 40 minwinner + wrapsocial
the highest-impact, highest-risk use case that runs reliably takes it.
who's on
getting there
questions
something else? ask on discord →what should i bring?
your laptop, your charger, and an agent idea, or a half-built one you want to finish. come solo or as a team. the build block runs just under two hours, so arrive with the idea picked and your setup ready.
do i need jev access before i arrive?
yes, ideally. get jev access through openrouter or the early access waitlist at typesafe.ai, and sort it out before sunday: build time is tight. no access yet? come anyway, we'll pair you up with someone who has it.
will failproof be set up?
yes. failproof will be set up at the venue, so your jev policies run on live agent tool calls: when a verdict says stop or instruct, you see the agent change course as it runs, not in a report afterwards.
i missed jev colearn. can i still come?
yes. came to jev colearn? this is the sequel. didn't? the 3:00 pm kickoff walks through jev evals and policies from the start, so you'll catch up in the first 15 minutes.
how does registration work?
register on luma. every registration is subject to host approval, so apply early; once you're approved, luma sends the confirmation and the day's details.
what is failproof ai?
failproof is your agent's oversight layer, steering it towards success. it traces your agents, finds failure patterns over time and sets up rules to prevent them from happening again. jev evals and policies just launched on failproof, for better reliability at the least cost and latency. the open-source cli is free: npm i -g failproofai.
MORE EVENTS
all events →- SEP20jev colearnco-learning · bengaluru · past
- APR30fresh context founder breakfastbreakfast · san francisco · past
- APR18hooked on claudemeetup · san francisco · past