What you'd have
How the AI agents that build all this are kept aligned
Most of this was built with AI agents. What made that workable is a system that checks, at each point in a session where work tends to drift, whether the agent is still working on what the person meant.
Institutional memory for agents
A search every agent is directed to use before touching code, across past sessions, research, product intent, decisions and history, so a new session starts knowing what was already settled.
What it gives you
Agents stop having to assume what has already been decided: a new session draws on the decisions, corrections and research that came before it, instead of someone re-explaining them.
What makes this version different
- Part of how every session is required to work, not an optional add-on.
- Has a tested fallback for when its primary connection fails, a failure most one-off setups don’t plan for.
- Research artifacts held
- 2,505
- Commits
- 187
- Since
- Oct 2025
Measured 26 September 2026
Checks at every stage of a session
The agent states what it understood before acting, sizes how careful to be by what could go wrong for a real person, and must show proof before calling work done.
What it gives you
Fewer confident mistakes, and effort scaled to risk: a typo moves fast, a billing change gets every check.
What makes this version different
- Enforced at the level of the tools the agent uses, not only written as guidance it can ignore.
- Audited against its own history, which found checks that looked right and did nothing, and rebuilt them (see What measuring revealed).
- Checks wired in
- 31
- Commits
- 91
- Since
- Mar 2026
Measured 26 September 2026
A loop that learns from corrections
When the same correction to an agent recurs across separate sessions, it becomes a standing instruction, with its effect tracked.
What it gives you
A correction that keeps recurring becomes standing guidance instead of being made again. Measuring this loop is also how we found it had quietly stopped acting (see What measuring revealed).
What makes this version different
- Designed to act only on corrections seen across at least three sessions, never on one anecdote, with enforcement that escalates gradually.
- Active patterns
- 26
Measured 26 September 2026
A playbook of 457 procedures
Written procedures covering how the business builds, tests, ships, markets, measures and writes, each callable by an agent at the right moment.
What it gives you
Much of how the company builds, tests, ships and markets, written down in a form agents can use rather than living only in one person’s head.
What makes this version different
- Procedures route to one another, so the set works as a system rather than a pile of documents.
- Built and revised continuously over eight months.
- Procedures
- 457
- Lines of instruction
- ~149,000
- Since
- Feb 2026
Measured 26 September 2026
The packaged alignment harness
The alignment system above, packaged to install in one command into a Claude Code setup, with a switch and a strength setting for every piece. Its public release follows first-install testing.
What it gives you
The system in a form other people can adopt.
What makes this version different
- 191 pieces read one by one against the intent of the whole before being packaged.
- Tested in real sessions as a first-time user, and in automated runs alongside an existing setup to check that nothing of theirs changes.
- Pieces reviewed
- 191
- Automated checks passing
- 48 of 48
Measured 26 September 2026
A record of how agents actually behave
A continuous record of agent activity and full session histories.
What it gives you
Evidence to study what fails and what works when AI agents do real production work.
What makes this version different
- It exists because the work happened; a fresh system starts with none of it.
- Recorded agent events
- 270,446
- Session histories
- 27,823
Measured 26 September 2026