Observability
Claude's hillclimb loop, with your traces underneath
On September 28, Anthropic’s Lance Martin published Automating eval design and hillclimbing …
NVIDIA proposes an AI agent kill switch in silicon after a year of sandbox escapes
TL;DR: NVIDIA’s new Open Agent Safety Platform assumes an agent can’t police itself, so …
SLMs for custom tasks: when small models beat frontier ones, and how to be sure
On the first of September, Tobi Lütke, the CEO of Shopify, posted something on X that interested me: …
A Skill Is Just an Agent. So Measure Your Changes.
TL;DR: A skill is just an agent, a prompt running in a harness, so you can test a change to it …
The EU AI Act wants a record. Your traces are already most of one.
I read a piece by Angie Jones from the Agentic AI Foundation this week, The EU AI Act and the new …
Evals Are a Revenue Strategy, Not a Safety Net
Here’s the moment every AI team knows. You have a demo that works. Not always, but most runs, …
Evals belong in your CI/CD pipeline
The other day I wrote that evals are just testing - the same old loop of “decide what good …
AI evals are just testing (with a much weirder answer key)
Years ago, before “AI” meant chatbots and before anyone said the word “eval” …
The Phoenix Project still holds up, even if you replaced all the code with agents
I reread The Phoenix Project last month. I do this every year or two — it’s one of those books …
Train, dev, test: the split that makes an LLM judge trustworthy
Say you’ve built an LLM judge and you want to know if it’s any good. The obvious move is …
Evals do three jobs, not one
Ask most people what an eval is for and you’ll get some version of “testing”. You …
Anatomy of an evaluator: what happens when your prompt meets your traces
Last time I pulled apart the eval prompt - the role, criteria, rubric and examples you write to tell …











