Llm
Anthropic says it fixed Claude's writing. I ran the evals to check.
TL;DR: Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard …
Real-time LLM guardrails with Jev: comparing latency and cost
I just bought a 2024 Chevy Tahoe for $1. pic.twitter.com/aq4wDitvQW
— Chris Bakke …
When everything looks like a nail: jev, intent detection, and everything old is new again
I’ve been watching a small model launch bounce around my feed this week, and it sent me down a …
SLMs for custom tasks: when small models beat frontier ones, and how to be sure
On the first of September, Tobi Lütke, the CEO of Shopify, posted something on X that interested me: …
Evals Are a Revenue Strategy, Not a Safety Net
Here’s the moment every AI team knows. You have a demo that works. Not always, but most runs, …
Evals belong in your CI/CD pipeline
The other day I wrote that evals are just testing - the same old loop of “decide what good …
Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures
Originally published on the Arize AI blog: Long-horizon agent benchmarks are fragmenting: a field …
AI evals are just testing (with a much weirder answer key)
Years ago, before “AI” meant chatbots and before anyone said the word “eval” …
Train, dev, test: the split that makes an LLM judge trustworthy
Say you’ve built an LLM judge and you want to know if it’s any good. The obvious move is …
Evals do three jobs, not one
Ask most people what an eval is for and you’ll get some version of “testing”. You …
Anatomy of an evaluator: what happens when your prompt meets your traces
Last time I pulled apart the eval prompt - the role, criteria, rubric and examples you write to tell …
Anatomy of an eval prompt: what to actually put in it
When people decide to use an LLM as a judge, the prompt they reach for first is almost always some …











