Thanks for this. Of your six principles, #4 is the one that gets harder as models improve, not easier. A more capable generator produces outputs that are wrong in more plausible, harder-to-catch ways, so the cost of independent verification scales up precisely as the value of everything else you list scales down. That asymmetry has a measurable face: across independent reviews of 100-plus AI models, roughly 45% of generated code carries a common security flaw, and models catch far fewer of those flaws reviewing than they introduce writing. Most teams building harnesses today are compounding the generation side; the few pulling durable value are the ones who built independent verification capacity to keep pace with it.
Fully agreed. Not to mention, as reasoning gets better, the bar for quality keeps going up. And we start having AI attempt to do harder and harder things. The people who can close that loop quickly and have their AI systems consistently perform without creating slop debt will quickly outpace others
Thank you for this article - it teaches and crystallizes something for me: as model capabilities improve, the strategic value shifts from the model itself to the harness around it.
Then, enterprises eventually face a second-order problem: not how to build a harness, but how to govern thousands of them.
Once agents are embedded across workflows, teams, partners, and vendors, harness engineering becomes a control-plane problem: identity, policy, observability, auditability, and the ability to manage authority at scale.
The more capable the reasoning becomes, the more valuable the governance layer becomes.
This is highly complementary to a multi-player agentic AI article I’m working on.
Thanks for this. Of your six principles, #4 is the one that gets harder as models improve, not easier. A more capable generator produces outputs that are wrong in more plausible, harder-to-catch ways, so the cost of independent verification scales up precisely as the value of everything else you list scales down. That asymmetry has a measurable face: across independent reviews of 100-plus AI models, roughly 45% of generated code carries a common security flaw, and models catch far fewer of those flaws reviewing than they introduce writing. Most teams building harnesses today are compounding the generation side; the few pulling durable value are the ones who built independent verification capacity to keep pace with it.
Fully agreed. Not to mention, as reasoning gets better, the bar for quality keeps going up. And we start having AI attempt to do harder and harder things. The people who can close that loop quickly and have their AI systems consistently perform without creating slop debt will quickly outpace others
Thank you for this article - it teaches and crystallizes something for me: as model capabilities improve, the strategic value shifts from the model itself to the harness around it.
Then, enterprises eventually face a second-order problem: not how to build a harness, but how to govern thousands of them.
Once agents are embedded across workflows, teams, partners, and vendors, harness engineering becomes a control-plane problem: identity, policy, observability, auditability, and the ability to manage authority at scale.
The more capable the reasoning becomes, the more valuable the governance layer becomes.
This is highly complementary to a multi-player agentic AI article I’m working on.