INPUT PORT: 01
>
eval_set: refund_qa · 240 examples
metric: rubric · pass-rate
AGENT ENGINE SANDBOX
OPTIMIZING PROMPT...
. . ✶ . . . ✶ ✶ ✶ . . ✶ ✶ ✶ ✶ . . ✶ ✶ ✶ . . . ✶ . .
CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL
v3 -
v5 -
v7 -
Latest Writing
Thoughts on software development, AI, and building systems
Jev vs OpenAI Decisions: 2,094 Real Eval Checks
Does it matter what System One model you should be using? On same tasks, how well do Jev and OpenAI's Decision API agree, and which one turns out to be stronger
Evals Apo Jev
Escalating Evals: 72% Fewer LLM Judge Calls
The cheapest LLM-as-judge is not doing the judge at all! While the Jev output is not on-par with the LLM-as-judge, using it with escalation gives us great cost saving with small amount of difference in verdicts
Evals Apo Jev
TDD in the age of agents
The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer
TDD Apo