Llm

Anthropic says it fixed Claude's writing. I ran the evals to check.
Anthropic says it fixed Claude's writing. I ran the evals to check.

TL;DR: Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard …

Real-time LLM guardrails with Jev: comparing latency and cost
Real-time LLM guardrails with Jev: comparing latency and cost

When everything looks like a nail: jev, intent detection, and everything old is new again
When everything looks like a nail: jev, intent detection, and everything old is new again

I’ve been watching a small model launch bounce around my feed this week, and it sent me down a …

SLMs for custom tasks: when small models beat frontier ones, and how to be sure
SLMs for custom tasks: when small models beat frontier ones, and how to be sure

On the first of September, Tobi Lütke, the CEO of Shopify, posted something on X that interested me: …

Evals Are a Revenue Strategy, Not a Safety Net
Evals Are a Revenue Strategy, Not a Safety Net

Here’s the moment every AI team knows. You have a demo that works. Not always, but most runs, …

Evals belong in your CI/CD pipeline
Evals belong in your CI/CD pipeline

The other day I wrote that evals are just testing - the same old loop of “decide what good …

AI evals are just testing (with a much weirder answer key)
AI evals are just testing (with a much weirder answer key)

Years ago, before “AI” meant chatbots and before anyone said the word “eval” …

Train, dev, test: the split that makes an LLM judge trustworthy
Train, dev, test: the split that makes an LLM judge trustworthy

Say you’ve built an LLM judge and you want to know if it’s any good. The obvious move is …

Evals do three jobs, not one
Evals do three jobs, not one

Ask most people what an eval is for and you’ll get some version of “testing”. You …

Anatomy of an evaluator: what happens when your prompt meets your traces
Anatomy of an evaluator: what happens when your prompt meets your traces

Last time I pulled apart the eval prompt - the role, criteria, rubric and examples you write to tell …

Anatomy of an eval prompt: what to actually put in it
Anatomy of an eval prompt: what to actually put in it

When people decide to use an LLM as a judge, the prompt they reach for first is almost always some …