Evals

Arize's next top decision model
Arize's next top decision model

Eight decision models. Three challenges. One bracket. Only one can be Arize’s next top …

Anthropic says it fixed Claude's writing. I ran the evals to check.
Anthropic says it fixed Claude's writing. I ran the evals to check.

TL;DR: Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard …

Claude's hillclimb loop, with your traces underneath
NVIDIA proposes an AI agent kill switch in silicon after a year of sandbox escapes
NVIDIA proposes an AI agent kill switch in silicon after a year of sandbox escapes

TL;DR: NVIDIA’s new Open Agent Safety Platform assumes an agent can’t police itself, so …

Real-time LLM guardrails with Jev: comparing latency and cost
Real-time LLM guardrails with Jev: comparing latency and cost

When everything looks like a nail: jev, intent detection, and everything old is new again
When everything looks like a nail: jev, intent detection, and everything old is new again

I’ve been watching a small model launch bounce around my feed this week, and it sent me down a …

SLMs for custom tasks: when small models beat frontier ones, and how to be sure
SLMs for custom tasks: when small models beat frontier ones, and how to be sure

On the first of September, Tobi Lütke, the CEO of Shopify, posted something on X that interested me: …

A Skill Is Just an Agent. So Measure Your Changes.
A Skill Is Just an Agent. So Measure Your Changes.

TL;DR: A skill is just an agent, a prompt running in a harness, so you can test a change to it …

The two-person team: the domain expert prompts, the engineer connects
The two-person team: the domain expert prompts, the engineer connects

At RenderATL I had the same conversation twice, from two different sides.

The first was at the …

Your eval criteria are already written, just scattered across three systems
Your eval criteria are already written, just scattered across three systems

In an earlier post I built a self-improving agent by mining a context graph out of data the team …

Evals Are a Revenue Strategy, Not a Safety Net
Evals Are a Revenue Strategy, Not a Safety Net

Here’s the moment every AI team knows. You have a demo that works. Not always, but most runs, …

Evals belong in your CI/CD pipeline
Evals belong in your CI/CD pipeline

The other day I wrote that evals are just testing - the same old loop of “decide what good …