← Back to news
Coding Evaluations Need Human Judgment

Photo: Jon Sullivan / Wikimedia Commons

29/08/2026

Coding Evaluations Need Human Judgment

When a benchmark becomes a product decision

Evaluation is not a side quest in AI-assisted development. It is the mechanism that decides what teams trust, what they ship, and what they later call "good enough." That is why OpenAI’s recent audit of SWE-Bench Pro matters beyond one benchmark. The team reported that roughly 30% of the tasks were broken, and the problems were not exotic. They included overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. In other words, the benchmark itself could tell the wrong story about model capability. If your yardstick is warped, your deployment decision will be warped too.

That lesson should feel familiar to any software team using coding agents. We already know that agents can draft code, summarize diffs, and speed up boring work. What is easier to forget is that the same system that writes code can also be used to evaluate code, triage failures, and suggest whether a task is solved. Once that happens, the quality of the evaluation becomes part of the product. You are no longer just asking, "Can the model fix the bug?" You are also asking, "Did we define the bug correctly, and did we define success in a way that matches reality?"

OpenAI’s audit is useful because it shows a mature way to handle that question. The review was not a single automated pass. It combined automated screening, repeated investigator-agent checks, and independent human review by experienced software engineers. The humans were more likely than the investigator agents to mark tasks as broken, and they sometimes found multiple problems in the same task. That is not a failure of automation. It is a reminder that the final call on whether a benchmark is trustworthy is itself a judgment call, and judgment still belongs to people.

What the audit found

The important detail is not only the 30% estimate. It is the shape of the failures. OpenAI grouped them into four families: overly strict tests that enforce a specific implementation without requiring it in the prompt; underspecified prompts that omit requirements the hidden tests still enforce; low-coverage tests that let incomplete fixes slip through; and misleading prompts that steer the model toward the wrong or contradictory behavior.

That mix is revealing because it exposes a point many product teams underestimate: a benchmark is not a neutral mirror. It is compressed policy. When the policy is wrong, it rewards the wrong behavior with great confidence. A passing score can then mean that the model learned to game the test, not that it understood the task.

The audit also shows why agents are useful without being sovereign. The models were excellent at sorting suspicious cases, reading traces, inspecting the repository, and comparing failure modes. But the final validation step still depended on experienced humans. That makes sense: identifying a failure pattern and deciding whether that pattern truly invalidates an evaluation are different jobs. The first is an extraction problem. The second is a governance problem.

A model can surface a defect; it cannot own the standard.

Why AI is good at triage, not arbitration

An agent is excellent at anything that looks like mechanical triage. It can scan a repository, find related functions, read an error trace, compare several patches, and flag inconsistencies. In a validation pipeline, that matters a lot. A huge amount of human time is wasted rediscovering context, filtering obvious cases, and doing the first pass on oversized batches.

But arbitration requires something different: deciding what "correct" means in a given context. Should the test check a user-visible behavior, or an internal implementation constraint? Does the business rule live in the repository, in a product note, or in a security discussion? Should the team accept a change that passes tests but violates an implicit compatibility expectation? A model can propose a plausible answer to each of those questions. It cannot, by itself, decide which answer actually reflects the team’s intent.

That is where the most expensive trap lives: the fluency of a summary can make us confuse speed with truth. A clean output feels like the problem is solved. In reality, it may just move responsibility onto an unverified assumption. If a team’s evaluations are already fragile, a convincing agent can make the fragility less visible instead of fixing it.

The hidden trap: fluent wrongness

Evaluation failures often arrive in very polite forms. Overly strict tests reward compliance with a hidden implementation instead of the expected behavior. Underspecified prompts let the model choose assumptions nobody intended. Low-coverage tests allow an incomplete fix to pass. Misleading prompts invert the lesson entirely. In all of these cases, the machine can announce success that will not survive real product use.

The practical danger for a team is simple: it starts optimizing the wrong thing. If the benchmark becomes the implicit truth, people begin to look for patches that pass the tests, to produce cleaner-looking diffs, or to chase pass rates that mean nothing outside the sandbox. A system can learn to satisfy the evaluator without becoming more useful to users.

That is why "the eval passed" is not enough evidence in a serious workflow. A pass only means the model satisfied the evaluator’s rules. It does not mean the rules matched the real intent, or even that the task was fair. Human review exists to ask the uncomfortable questions: Are the hidden tests legitimate? Is the prompt vague? Does the benchmark reward one style of fix over another? Those are system questions, not just output questions.

How human-in-command evals work in practice

If we want humans in command, we have to make that command lighter, not optional. In practice, that means using models for the mechanical work and reserving people for interpretation. A useful flow looks like this:

  1. Write the specification before you write the benchmark.
  2. Ask an agent to screen for suspicious tasks, traces, and tests.
  3. Have humans sample both successes and failures, not just the obvious outliers.
  4. Classify issues by severity and by type.
  5. Escalate anything ambiguous, high-stakes, or security-sensitive to a person with authority to decide.

That pattern is not slow bureaucracy. It is how you avoid turning an eval into a lottery. It also scales better than total manual review because the machine does the first pass across the long tail of obvious cases. Human time is then reserved for the small set of examples where judgment is actually valuable.

Anthropic’s framework for trustworthy agents points in the same direction. Their emphasis on human control, transparency, and careful permissioning is the product version of the same engineering principle: autonomy is useful only when the path back to human authority remains clear. GitHub’s recent guidance on reliable AI workflows adds another layer to that idea by insisting on validation checkpoints and explicit human gates. The vendors are converging because the underlying problem is the same: speed without supervision is brittle.

A minimal governance pattern

  • One owner per benchmark. Someone has to be accountable for what "correct" means.
  • One audit trail per decision. Keep the reason a task was accepted or rejected.
  • One adversarial sample per cycle. Regularly test whether the benchmark can be gamed.
  • One escalation path for uncertainty. If no one can explain a result, it should not quietly pass.

What a healthier evaluation culture looks like

In practice, the best teams treat evals as living specifications. They version them, review them, and change them when the product changes. A benchmark that was honest six months ago may become misleading after a new API, a new permission model, or a new user journey. That means the right question is not whether an eval is stable forever. It is whether the team can explain its drift and update it deliberately. AI can help by diffing test changes, summarizing regressions, and surfacing places where the benchmark has started to encode yesterday's assumptions.

Another useful habit is to separate three things that are often blended together in conversation: model quality, test quality, and product quality. A model can improve while the benchmark gets worse. A benchmark can become more selective while the product still becomes less reliable. And a product can ship faster while user trust declines. When those distinctions are visible, teams stop chasing one dashboard number and start asking whether the whole system still serves the users. That is the essence of human command: the metric is useful, but it is not the boss.

What this means for AI-assisted development teams

For teams shipping software with coding agents, the practical implication is simple: use AI to increase throughput, but do not let AI define the standards. Let the agent draft tests, summarize diffs, surface missing edge cases, and even flag suspicious evaluation tasks. But the benchmark rubric, the release criteria, and the decision to accept risk should stay with humans who understand the product and its users.

That matters even more when the team feels pressure to move faster. The temptation is to treat a green eval or a polished agent summary as proof of readiness. Yet the fastest way to accumulate technical debt is to let the test harness replace product judgment. If the benchmark says something is done, but no one can explain why it is correct, the work is not done. It is only automated.

The right question is not "Can the model complete the task?" It is "Would I accept this result if a junior engineer produced it and I had to defend the merge in front of the team, the security reviewer, and the user?" That question keeps the human where they belong: in command of the standard, not just the tool.

Keep the human in the loop, because the loop defines the work

AI is very good at producing drafts, hypotheses, and first passes. It is also good at helping us audit our own systems. But when the subject is evaluation quality, the job is not to maximize automation at any cost. The job is to make sure the numbers mean something. A benchmark that lies, even politely, is worse than no benchmark at all.

The best teams also keep a small, explicit review loop around their benchmarks. They re-run edge cases after each major change, compare today’s failures with last month’s, and ask whether the benchmark is drifting toward convenience instead of truth. That sounds modest, but it is how you keep the evaluation from becoming a decorative scorecard. A benchmark is only useful when it stays close to the work it claims to represent. If it drifts too far, it becomes a ritual: something the team reports, but no longer learns from.

Benchmarks should behave like contracts, not trophies. A contract tells everyone what the system owes, and it can be revised when the product changes. A trophy is just something to display. If a team confuses the two, it will optimize for visible success instead of reliable behavior. Human review is what keeps the contract readable. That is why the healthiest teams review the rules themselves, not just the score, whenever the product, the threat model, or the user journey changes.

Sources