Blogs
Evals do three jobs, not one
Ask most people what an eval is for and you’ll get some version of “testing”. You …
Anatomy of an evaluator: what happens when your prompt meets your traces
Last time I pulled apart the eval prompt - the role, criteria, rubric and examples you write to tell …
Anatomy of an eval prompt: what to actually put in it
When people decide to use an LLM as a judge, the prompt they reach for first is almost always some …
Spans, traces and sessions: the three zoom levels of an AI app
The moment you start tracing an AI app, three words turn up everywhere: span, trace, session. People …
Evals aren't a step at the end. They run the whole way through
There’s a version of building an AI app that goes like this. You build the thing, you get it …
Four ways to run an eval, from a cheap unit test to a full-blown agent
Someone asked me last week how you actually run an eval on an AI app. I gave the honest answer, …
I built a little Claude that dances when it needs me
I’ve got into a bad habit lately. I set Claude Code off on some task - refactor this, write …
How to build an eval you can actually trust
Here’s how most people build an eval. They open a file, write an LLM judge prompt that says …
Two labs started dreaming, and they built two different architectures
Originally published on the Arize AI blog: Two labs started dreaming, and they built two different …
Twenty-seven years ago I nearly put six people out of work with an Excel spreadsheet
Twenty-seven years ago I nearly put six people out of work with an Excel spreadsheet.
Every AI …
Memory is still a missing primitive: Cataloguing what the field is actually shipping
Originally published on the Arize AI blog: Memory is still a missing primitive: Cataloguing what the …
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both …











