Discussion about this post

User's avatar
State of Play's avatar

Thanks for this. Of your six principles, #4 is the one that gets harder as models improve, not easier. A more capable generator produces outputs that are wrong in more plausible, harder-to-catch ways, so the cost of independent verification scales up precisely as the value of everything else you list scales down. That asymmetry has a measurable face: across independent reviews of 100-plus AI models, roughly 45% of generated code carries a common security flaw, and models catch far fewer of those flaws reviewing than they introduce writing. Most teams building harnesses today are compounding the generation side; the few pulling durable value are the ones who built independent verification capacity to keep pace with it.

Vlad Stojanovski's avatar

Thank you for this article - it teaches and crystallizes something for me: as model capabilities improve, the strategic value shifts from the model itself to the harness around it.

Then, enterprises eventually face a second-order problem: not how to build a harness, but how to govern thousands of them.

Once agents are embedded across workflows, teams, partners, and vendors, harness engineering becomes a control-plane problem: identity, policy, observability, auditability, and the ability to manage authority at scale.

The more capable the reasoning becomes, the more valuable the governance layer becomes.

This is highly complementary to a multi-player agentic AI article I’m working on.

1 more comment...

No posts

Ready for more?