Home
  • Blog
[ New ] Jev vs OpenAI Decisions: 2,094 Real Eval Checks

Hi, I'm Sami!

AI ENGINEER & BUILDER
GitHub
INPUT PORT: 01

>

eval_set: refund_qa · 240 examples

metric: rubric · pass-rate

AGENT ENGINE SANDBOX

OPTIMIZING PROMPT...

  .   .   ✶   .   .
    . ✶ ✶ ✶ .
  . ✶ ✶ ✶ ✶ .
    . ✶ ✶ ✶ .
  .   .   ✶   .   .

CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL

v3 -

v5 -

v7 -

Latest Writing

Thoughts on software development, AI, and building systems

View all
Jev vs OpenAI Decisions: 2,094 Real Eval Checks
Oct 7, 2026

Jev vs OpenAI Decisions: 2,094 Real Eval Checks

Does it matter what System One model you should be using? On same tasks, how well do Jev and OpenAI's Decision API agree, and which one turns out to be stronger

Evals Apo Jev
Escalating Evals: 72% Fewer LLM Judge Calls
Oct 5, 2026

Escalating Evals: 72% Fewer LLM Judge Calls

The cheapest LLM-as-judge is not doing the judge at all! While the Jev output is not on-par with the LLM-as-judge, using it with escalation gives us great cost saving with small amount of difference in verdicts

Evals Apo Jev
TDD in the age of agents
Oct 3, 2026

TDD in the age of agents

The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer

TDD Apo
View all posts