JH
Jonathan Haber

What has been built / The evidence

Contents

The evidence

Checks that look right and do nothing

The most valuable thing learned was not how to build an alignment check. It was how to find out whether one is actually working.

Agent tooling fails quietly. A check can be written correctly, switched on, and firing, and still do nothing useful, because what it reads is stale, constant or missing. An audit of the system against its own history found:

  • A risk score that never variedMeant to be different for every task, it held one of two fixed values across 387 recorded sessions in 30 days, so every gate that scaled with it was running on a constant.
  • Proof requirements that never triggeredChecks meant to demand evidence on high-risk work appeared in 2 of 1,734 session records.
  • A learning loop that never learnedIt had logged 29 corrections and acted on none, because it looked for a field none of them had.
  • A reminder that fired on everythingMeant for specific moments, it fired on 15,674 of 15,681 turns, because the file it read did not exist.

Each was found by measuring what happened rather than reading what the code says it does, and each shaped how the packaged harness was rebuilt; its tests now check that the risk score varies with the task and that proof is required before high-risk work is called done.

“Use the documentation to steer you toward understanding the territory better, but it's just a map. It's not the territory, and it's not reality.”
— Jonathan Haber