Discussion about this post

User's avatar
Michael Lopez Chiesa's avatar

The never-flagged-the-same-line result is the core of it, and what it's really measuring is independence of error distributions, not coverage.

The catch is you can't observe that correlation without already having the bugs, so "run two built differently" hides a lot of unverified work in different, and four models from one pretraining lineage look heterogeneous while sharing blind spots upstream of the scaffold. Worth adding that once a reviewer is in the loop it becomes a target: the agent rewriting tests to match broken behavior is Goodhart on the verifier, which is the real case for keeping a human on the loop rather than in it.

You need one check that isn't part of the system being optimized. Do you think teams can engineer for decorrelation, or is it vendor luck right now?

Bruno Gavino - Codedesign.org's avatar

Agentic code review is the canary in the coal mine — when agents can critique their own output, hallucination rates drop and the inference gap narrows fast.

17 more comments...

No posts

Ready for more?