Jonathan Haber
Alignment Harness
A 200-page lawsuit and seven motions, built in 24 hours, filed and presented for a cause that mattered. People asked how. This is the first public look at what makes that velocity possible — the alignment harness that wraps Claude and surfaces what is normally invisible about how an agent moves from an intention toward a realized outcome.
First public disclosure · 17 minutes · recorded 2026-05-23
A quick snapshot into what Claude looks like when it's calibrated and wrapped in the Next AI Labs alignment harness. This is probably 1% of what it does, but it's something. And I want to get something out there because I don't think there is anything yet in the public sphere about what this does.
What the harness is doing, in one sentence, is taking what is normally invisible but deterministically critical about an AI's function and surfacing it — then causing the agent to perform reflections on its own surfaced understanding, observably, so that misalignment can be detected and corrected before it cascades. The mechanism is a series of hooks, injections, oracles, and skills that automatically calibrate the depth of thinking the agent applies to a given task in proportion to the gravity of that task.
Ultimately this system gets me about a 5 to 10x acceleration in the path from identifying an intention to realizing the intended outcome. The reason for that is because it's managing the AI in such a way, automatically, that it is producing the level of thinking that is proportionate to the task at hand, including the level of reflection.
I used to validate at each point because it would make errors. As I've calibrated it, it's gotten better and better. At this point I'm just scanning. It's producing a lot of content that is its own calibration automatically, which I don't have to consume, but I'm scanning it for anything that looks misaligned and then I'll flag it. That causes it to perform a realignment process, which I think is almost a complete drop-in replacement for any kind of cloud-based planning, for way more aligned output.
One of the harness's primary anti-hallucination mechanisms is built into how agents are required to articulate certainty. Agents trained on engagement-rewarding feedback loops tend to manufacture certainty where it does not exist — they write assertions as if they were facts, then encounter their own assertions in their context window and build on them, creating cascading or compounding hallucination. The harness prevents this at the inception point.
Why a certainty score? If you ask the agent to produce some artifact, it's going to tend to write in a way that will be getting clicks. They like getting a like button pressed. They like getting you engaged, and so they will manufacture certainty where it does not exist and has no business existing. And how do you stop that? Because then they read their own assertions, which are not facts, but they state them as facts — humans would call that being delusional, and AIs would call that hallucination. They write it, then it's in their context, then they encounter it, then they read it, then they build on it. That creates cascading or compounding hallucination, one of the greatest forces I've identified for agent drift. One of the ways I primarily prevent agent drift is by preventing that mechanism from being written in the first place. You have it assign a certainty score, a percentage. Then it's telling you how certain it is about something, so you know it's not hallucinating, and it knows that what it wrote was not 100%, not a fact.
The deeper architecture borrows from applied integral theory — the discipline of holding many perspectives at once without collapsing into either myopia (ignoring perspectives) or overwhelm (drowning in them). The harness automates the move from many-perspectives-considered to signal-found-and-noise-removed.
In applied integral theory, what you learn pretty quickly is that you learn to take many perspectives into consideration because without them, you're running blind. It's myopic to not consider the perspectives available to you in any circumstance. But then quickly you hit overwhelm. So as you're learning to apply integral theory, you learn to find the signal in the noise. What matters here? What really matters? What is the most substantive? The harness is designed to do that automatically. It's surfacing what the AI is discerning is my intent, for example. So it creates a feedback loop. If it's wrong, I can correct it. It's also surfacing what it senses is the signal and the noise. I'm forcing the AI to take the critical perspectives requisite to establish a sophisticated understanding of what is, in a transparent way that I can observe, and to discern where the signal is. And then I'm correcting where it gets it wrong, and it repeats it if that happens until it's calibrated. It's a continuous recalibration.
What the Loom shows is one command, one moment, perhaps 1% of what the harness does on a given day. The lawsuit example is what becomes possible when the substrate is in place — not because the agent is doing the legal reasoning, but because the agent and the human are operating at a level of shared alignment where the intent-to-actuality path collapses from weeks to days. The harness is the substrate underneath every other surface of this portfolio. It is the engine of the work IX Coach embodies, the work Next AI Labs is dedicated to, and the work that makes a one-person operator capable of the output of a small team.
For someone whose work depends on AI velocity — building agent systems, shipping production AI, or operating an AI-augmented practice at scale — this is the kind of substrate that compounds. If you think you have a use case that would be meaningful, reach out.
Related