Home
  • Blog
[ New ] Jev-as-Judge: A Confidence Signal for Evals

Hi, I'm Sami!

AI ENGINEER & BUILDER
GitHub
INPUT PORT: 01

>

eval_set: refund_qa · 240 examples

metric: rubric · pass-rate

AGENT ENGINE SANDBOX

OPTIMIZING PROMPT...

  .   .   ✶   .   .
    . ✶ ✶ ✶ .
  . ✶ ✶ ✶ ✶ .
    . ✶ ✶ ✶ .
  .   .   ✶   .   .

CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL

v3 -

v5 -

v7 -

Latest Writing

Thoughts on software development, AI, and building systems

View all
Jev-as-Judge: A Confidence Signal for Evals
Sep 20, 2026

Jev-as-Judge: A Confidence Signal for Evals

System One models like Jev can be used in evals to give confidence signal for LLM-as-judge results, which can be easily used to debug eval setup problems.

Evals Apo Jev
Sep 11, 2026

Principles of Loop Engineering

What needs to be true, to automate work to the agents, while knowing the output is what we wanted.

Loop Engineering Apo
Closing the Loop on Harness Engineering
Aug 26, 2026

Closing the Loop on Harness Engineering

My exploration of harness engineering, and the primitives that allow the agents to do work on loop even when designing harnesses

Harness Engineering AI Agents Code Generation
View all posts