Home
  • Blog
[ New ] Agent-as-Judge: Harder Evals, More Computation

Hi, I'm Sami!

AI ENGINEER & BUILDER
GitHub
INPUT PORT: 01

>

eval_set: refund_qa · 240 examples

metric: rubric · pass-rate

AGENT ENGINE SANDBOX

OPTIMIZING PROMPT...

  .   .   ✶   .   .
    . ✶ ✶ ✶ .
  . ✶ ✶ ✶ ✶ .
    . ✶ ✶ ✶ .
  .   .   ✶   .   .

CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL

v3 -

v5 -

v7 -

Latest Writing

Thoughts on software development, AI, and building systems

View all
Agent-as-Judge: Harder Evals, More Computation
Sep 30, 2026

Agent-as-Judge: Harder Evals, More Computation

What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.

Evals Apo Jev
Jev-as-Judge: A Confidence Signal for Evals
Sep 20, 2026

Jev-as-Judge: A Confidence Signal for Evals

System One models like Jev can be used in evals to give confidence signal for LLM-as-judge results, which can be easily used to debug eval setup problems.

Evals Apo Jev
Sep 11, 2026

Principles of Loop Engineering

What needs to be true, to automate work to the agents, while knowing the output is what we wanted.

Loop Engineering Apo
View all posts