INPUT PORT: 01
>
eval_set: refund_qa · 240 examples
metric: rubric · pass-rate
AGENT ENGINE SANDBOX
OPTIMIZING PROMPT...
. . ✶ . . . ✶ ✶ ✶ . . ✶ ✶ ✶ ✶ . . ✶ ✶ ✶ . . . ✶ . .
CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL
v3 -
v5 -
v7 -
Latest Writing
Thoughts on software development, AI, and building systems
Agent-as-Judge: Harder Evals, More Computation
What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.
Evals Apo Jev
Jev-as-Judge: A Confidence Signal for Evals
System One models like Jev can be used in evals to give confidence signal for LLM-as-judge results, which can be easily used to debug eval setup problems.
Evals Apo Jev
Principles of Loop Engineering
What needs to be true, to automate work to the agents, while knowing the output is what we wanted.
Loop Engineering Apo