Great piece - this is the framing I've been building against for the last 6 months.
One thing I'd add, because it changes the ending: nearly every example here (Granola, Cursor, Sierra) wraps a rented model and builds the moat around it -- memory, context, workflow. I've took it a step further and moved the model out of the critical path. Apps are produced by a deterministic compiler from a signed contract; the LLM only makes bounded decisions at the edges.
That does two useful things to your three-question test. It passes "would it survive a model swap?" structurally rather than as a promise -- we've run the whole pipeline (normally Opus 4.8 + Sonnet 5) end-to-end on NVIDIA Nemotron 3 Ultra. And it turns "we have data" from theater into a real flywheel: because every build is compiled and validated, every build is execution-verified training data -- the compounding kind, not a rotting pile -- and we can start moving away from proprietary frontier models.
When the model gets cheap and revocable, owning the compiler and an ownable model -- air-gapped where the customer needs it -- stops being a wrapper problem and becomes the moat.
Honestly not sure Granola passes your own second test. Two-year-old meeting notes look a lot like data that accumulates without compounding, and your rot point would seem to apply to them too. Maybe the real lock is that leaving feels like walking away from your own memory, like you say, but that's a switching cost living in the user's head, not sure it's a loop.
a moat is not just a bigger pile of notes, old context has to actively improve current and future work, in a controlled (by the user) and measurable way.
...and context is more than meetings, it is the screens, apps, chats, docs and decisions around them - shared and agreed.
Moats built on model performance are evaporating because the tech is becoming a commodity. Real defensibility today comes from proprietary data access or deep integration into specific workflows. My piece on governance aligns with this because authority to act is the only moat that lasts once the models equalize. If your agent can execute tasks but lacks the internal controls to verify intent or scope, you are just building a liability. The market is shifting from who has the best math to who owns the process. Stop obsessing over the architecture and start securing the decision chain.
Great post, overall - but I would disagree on two issues:
1) we use and appreciate Open Router - but do they really have a moat? Routing is, by definition, a commodity business: if a new router came along offering access to the same set of models we use, with a similarly functional (or improved) routing spec operating at a lower token margin, we would switch In a heartbeat. Similarly, once a model consumer reaches a large enough size, it would seem to make logical sense to bring the routing function in-house, maintaining token accounts with the direct suppliers: thus, demonstrating that open Router has a very limited upside as well.
2) will the frontier labs really be OK? Their product is increasingly a niche super-premium offering, and it's not clear to me that there's much of a market for that at all. Every challenge we have faced, we have been able to solve with 2nd or 3rd tier models, and we vastly prefer small and open models for all the reasons in this post, and many more. I would honestly like to understand who really needs the SOTA frontier models, and what on earth they are needed for. Even more, every new release brings a new round of concern of unwanted excess model behavior - more expense, less speed, and more guardrail concerns: who really needs these products?!
This is exactly the right way to think about the next phase of AI competition.
As open-weight models continue to narrow the capability gap, the durable advantage increasingly moves outward from the model itself and into the context layer: proprietary data, accumulated institutional memory, embedded workflows, permissions, orchestration, trust, and the ability to turn intelligence into action.
There is also a national-security implication here. If models become more abundant and interchangeable, then the most valuable target may no longer be the model weights themselves. It may be the operational context surrounding them: internal communications, supplier relationships, engineering decisions, meeting notes, decision authorities, and the workflows through which an organization actually functions.
That suggests a different understanding of sovereign AI. Sovereignty may depend less on owning the world’s best model than on controlling the system in which the model operates, and being able to replace the model without surrendering the data, memory, workflows, or trust architecture around it.
The model may answer the question. The surrounding system determines which question is asked, what information shapes the answer, and what happens next.
Agree the moat question matters more than the model leaderboard now. The valuation angle we’d add: distribution and switching-cost moats convert to durable FCF, but a lot of current AI multiples are pricing that durability as if it’s already proven. The test is whether pricing power survives once the model layer commoditizes. Which moat type do you think holds up best there?
The arguments for and against we have heard before: SaaS (proprietary models accessed via external API) and On-Prem (local, open source models).
The decision points are also the same:
1) Operational Complexity: are you willing to maintain, upgrade, and manage the On-Prem service yourself? Are you willing to invest in the expertise indefinitely?
2) Feature Velocity: Are the open-source keeping up with you feature needs?
Great piece - this is the framing I've been building against for the last 6 months.
One thing I'd add, because it changes the ending: nearly every example here (Granola, Cursor, Sierra) wraps a rented model and builds the moat around it -- memory, context, workflow. I've took it a step further and moved the model out of the critical path. Apps are produced by a deterministic compiler from a signed contract; the LLM only makes bounded decisions at the edges.
That does two useful things to your three-question test. It passes "would it survive a model swap?" structurally rather than as a promise -- we've run the whole pipeline (normally Opus 4.8 + Sonnet 5) end-to-end on NVIDIA Nemotron 3 Ultra. And it turns "we have data" from theater into a real flywheel: because every build is compiled and validated, every build is execution-verified training data -- the compounding kind, not a rotting pile -- and we can start moving away from proprietary frontier models.
When the model gets cheap and revocable, owning the compiler and an ownable model -- air-gapped where the customer needs it -- stops being a wrapper problem and becomes the moat.
Honestly not sure Granola passes your own second test. Two-year-old meeting notes look a lot like data that accumulates without compounding, and your rot point would seem to apply to them too. Maybe the real lock is that leaving feels like walking away from your own memory, like you say, but that's a switching cost living in the user's head, not sure it's a loop.
a moat is not just a bigger pile of notes, old context has to actively improve current and future work, in a controlled (by the user) and measurable way.
...and context is more than meetings, it is the screens, apps, chats, docs and decisions around them - shared and agreed.
an archive just accumulates
Moats built on model performance are evaporating because the tech is becoming a commodity. Real defensibility today comes from proprietary data access or deep integration into specific workflows. My piece on governance aligns with this because authority to act is the only moat that lasts once the models equalize. If your agent can execute tasks but lacks the internal controls to verify intent or scope, you are just building a liability. The market is shifting from who has the best math to who owns the process. Stop obsessing over the architecture and start securing the decision chain.
https://cyrilsimonnet.substack.com/p/everyone-is-shopping-for-agentic
Great post, overall - but I would disagree on two issues:
1) we use and appreciate Open Router - but do they really have a moat? Routing is, by definition, a commodity business: if a new router came along offering access to the same set of models we use, with a similarly functional (or improved) routing spec operating at a lower token margin, we would switch In a heartbeat. Similarly, once a model consumer reaches a large enough size, it would seem to make logical sense to bring the routing function in-house, maintaining token accounts with the direct suppliers: thus, demonstrating that open Router has a very limited upside as well.
2) will the frontier labs really be OK? Their product is increasingly a niche super-premium offering, and it's not clear to me that there's much of a market for that at all. Every challenge we have faced, we have been able to solve with 2nd or 3rd tier models, and we vastly prefer small and open models for all the reasons in this post, and many more. I would honestly like to understand who really needs the SOTA frontier models, and what on earth they are needed for. Even more, every new release brings a new round of concern of unwanted excess model behavior - more expense, less speed, and more guardrail concerns: who really needs these products?!
This is exactly the right way to think about the next phase of AI competition.
As open-weight models continue to narrow the capability gap, the durable advantage increasingly moves outward from the model itself and into the context layer: proprietary data, accumulated institutional memory, embedded workflows, permissions, orchestration, trust, and the ability to turn intelligence into action.
There is also a national-security implication here. If models become more abundant and interchangeable, then the most valuable target may no longer be the model weights themselves. It may be the operational context surrounding them: internal communications, supplier relationships, engineering decisions, meeting notes, decision authorities, and the workflows through which an organization actually functions.
That suggests a different understanding of sovereign AI. Sovereignty may depend less on owning the world’s best model than on controlling the system in which the model operates, and being able to replace the model without surrendering the data, memory, workflows, or trust architecture around it.
The model may answer the question. The surrounding system determines which question is asked, what information shapes the answer, and what happens next.
We expanded on that strategic dimension in today’s edition of Building Our Future: https://buildingourfuture.substack.com/p/the-model-is-no-longer-the-crown
Agree the moat question matters more than the model leaderboard now. The valuation angle we’d add: distribution and switching-cost moats convert to durable FCF, but a lot of current AI multiples are pricing that durability as if it’s already proven. The test is whether pricing power survives once the model layer commoditizes. Which moat type do you think holds up best there?
The arguments for and against we have heard before: SaaS (proprietary models accessed via external API) and On-Prem (local, open source models).
The decision points are also the same:
1) Operational Complexity: are you willing to maintain, upgrade, and manage the On-Prem service yourself? Are you willing to invest in the expertise indefinitely?
2) Feature Velocity: Are the open-source keeping up with you feature needs?
3) Cost: pure opez vs. opex and capex