Home
  • Blog
[ New ] Escalating Evals: 72% Fewer LLM Judge Calls

Hi, I'm Sami!

AI ENGINEER & BUILDER
GitHub
INPUT PORT: 01

>

eval_set: refund_qa · 240 examples

metric: rubric · pass-rate

AGENT ENGINE SANDBOX

OPTIMIZING PROMPT...

  .   .   ✶   .   .
    . ✶ ✶ ✶ .
  . ✶ ✶ ✶ ✶ .
    . ✶ ✶ ✶ .
  .   .   ✶   .   .

CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL

v3 -

v5 -

v7 -

Latest Writing

Thoughts on software development, AI, and building systems

View all
Escalating Evals: 72% Fewer LLM Judge Calls
Oct 5, 2026

Escalating Evals: 72% Fewer LLM Judge Calls

The cheapest LLM-as-judge is not doing the judge at all! While the Jev output is not on-par with the LLM-as-judge, using it with escalation gives us great cost saving with small amount of difference in verdicts

Evals Apo Jev
TDD in the age of agents
Oct 3, 2026

TDD in the age of agents

The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer

TDD Apo
Agent-as-Judge: Harder Evals, More Computation
Sep 30, 2026

Agent-as-Judge: Harder Evals, More Computation

What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.

Evals Apo Jev
View all posts