>
eval_set: refund_qa · 240 examples
metric: rubric · pass-rate
OPTIMIZING PROMPT...
. . ✶ . . . ✶ ✶ ✶ . . ✶ ✶ ✶ ✶ . . ✶ ✶ ✶ . . . ✶ . .
v3 -
v5 -
v7 -
Latest Writing
Thoughts on software development, AI, and building systems
TDD in the age of agents
The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer
Agent-as-Judge: Harder Evals, More Computation
What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.
Jev-as-Judge: A Confidence Signal for Evals
System One models like Jev can be used in evals to give confidence signal for LLM-as-judge results, which can be easily used to debug eval setup problems.