Observability
SLMs for custom tasks: when small models beat frontier ones, and how to be sure
On the first of September, Tobi Lütke, the CEO of Shopify, posted something on X that interested me: …
A Skill Is Just an Agent. So Measure Your Changes.
TL;DR: A skill is just an agent, a prompt running in a harness, so you can test a change to it …
The EU AI Act wants a record. Your traces are already most of one.
I read a piece by Angie Jones from the Agentic AI Foundation this week, The EU AI Act and the new …
Evals Are a Revenue Strategy, Not a Safety Net
Here’s the moment every AI team knows. You have a demo that works. Not always, but most runs, …
Evals belong in your CI/CD pipeline
The other day I wrote that evals are just testing - the same old loop of “decide what good …
AI evals are just testing (with a much weirder answer key)
Years ago, before “AI” meant chatbots and before anyone said the word “eval” …
The Phoenix Project still holds up, even if you replaced all the code with agents
I reread The Phoenix Project last month. I do this every year or two — it’s one of those books …
Train, dev, test: the split that makes an LLM judge trustworthy
Say you’ve built an LLM judge and you want to know if it’s any good. The obvious move is …
Evals do three jobs, not one
Ask most people what an eval is for and you’ll get some version of “testing”. You …
Anatomy of an evaluator: what happens when your prompt meets your traces
Last time I pulled apart the eval prompt - the role, criteria, rubric and examples you write to tell …
Anatomy of an eval prompt: what to actually put in it
When people decide to use an LLM as a judge, the prompt they reach for first is almost always some …
Spans, traces and sessions: the three zoom levels of an AI app
The moment you start tracing an AI app, three words turn up everywhere: span, trace, session. People …











