The short-term investment signal from this lecture is semiconductor infrastructure. The long-term signal is energy. Jensen's 1,000x compute growth over ten years is well understood by markets. His observation that AI energy demand is probably 1,000x current levels, and that the market is now large enough to fund the solution without subsidies, is less fully priced.
What most allocators are signalling in the June consensus is constructive on utilities and energy, partly on AI power demand as a structural tailwind. But the CMA frameworks underlying those overweights are still anchored to incremental demand growth, not to a compute pattern shift from on-demand to continuous that structurally changes the energy profile of the entire digital economy.
The tokens-per-watt improvement curve is the number worth tracking for long-term expected return models. The teams compounding it fastest will determine whether the energy constraint becomes the dominant infrastructure investment thesis of the decade.
Excellent synthesis. The section on agent compute patterns is the one that doesn’t get nearly enough attention, and it maps directly to what’s happening at the edge right now.
Jensen’s framing — that agents require low-latency, single-threaded, local memory operations — is essentially a description of what embedded and IoT engineers have been designing for for decades, except now with intelligence baked in. The irony is that cloud-centric AI architectures are having to re-learn design principles that edge hardware has always demanded: minimal latency, low power budgets, local context, and tolerance for intermittent connectivity.
The memory bandwidth bottleneck point (Section 3) is even more acute at the edge. On a device with 4-8GB of LPDDR5 shared across CPU, NPU, and radio subsystems, memory bandwidth is the single hardest constraint. This is why architectures like Liquid AI’s LFM2 (which replaces KV-cache-heavy attention layers with gated convolutions) are genuinely important — they’re not incremental improvements, they’re re-architecting inference for the bandwidth reality of real hardware.
On the energy point — Eelco’s read is right that the market hasn’t fully priced the demand shift. I’d add one layer: at the embedded/IoT scale, the energy constraint isn’t about grid infrastructure, it’s about battery life measured in months, not megawatts. The NPU efficiency curves from Qualcomm, Apple, and MediaTek matter enormously here — the difference between 0.5W and 2W sustained inference isn’t a footnote for a battery-powered connected device, it’s the product viability line.
The co-design insight (Section 1) also applies in a very literal way when you’re designing edge hardware: silicon, firmware, radio stack (eSIM/cellular modem), and inference runtime have to be designed as a single system, or you’re leaving significant efficiency on the table. Fascinating era to be building in.
Point four is the one operators should sit with. Co-design is an organizational claim before it is a technical one. The reason a million-x is hard to copy is not the silicon, it is that almost no company will let one owner optimize across layers that separate budgets and separate VPs control. Stanford's fragmented grants are the same structure that produces stalled pilots inside large enterprises: every function optimizes its own layer, and nobody owns the whole. The accountability line lands for the same reason. Until someone holds the P&L across the layers, the stack stays a sum of parts, and the compounding never arrives.
The short-term investment signal from this lecture is semiconductor infrastructure. The long-term signal is energy. Jensen's 1,000x compute growth over ten years is well understood by markets. His observation that AI energy demand is probably 1,000x current levels, and that the market is now large enough to fund the solution without subsidies, is less fully priced.
What most allocators are signalling in the June consensus is constructive on utilities and energy, partly on AI power demand as a structural tailwind. But the CMA frameworks underlying those overweights are still anchored to incremental demand growth, not to a compute pattern shift from on-demand to continuous that structurally changes the energy profile of the entire digital economy.
The tokens-per-watt improvement curve is the number worth tracking for long-term expected return models. The teams compounding it fastest will determine whether the energy constraint becomes the dominant infrastructure investment thesis of the decade.
The teams that win are the ones who design an entire stack to work together from the beginning.
Excellent synthesis. The section on agent compute patterns is the one that doesn’t get nearly enough attention, and it maps directly to what’s happening at the edge right now.
Jensen’s framing — that agents require low-latency, single-threaded, local memory operations — is essentially a description of what embedded and IoT engineers have been designing for for decades, except now with intelligence baked in. The irony is that cloud-centric AI architectures are having to re-learn design principles that edge hardware has always demanded: minimal latency, low power budgets, local context, and tolerance for intermittent connectivity.
The memory bandwidth bottleneck point (Section 3) is even more acute at the edge. On a device with 4-8GB of LPDDR5 shared across CPU, NPU, and radio subsystems, memory bandwidth is the single hardest constraint. This is why architectures like Liquid AI’s LFM2 (which replaces KV-cache-heavy attention layers with gated convolutions) are genuinely important — they’re not incremental improvements, they’re re-architecting inference for the bandwidth reality of real hardware.
On the energy point — Eelco’s read is right that the market hasn’t fully priced the demand shift. I’d add one layer: at the embedded/IoT scale, the energy constraint isn’t about grid infrastructure, it’s about battery life measured in months, not megawatts. The NPU efficiency curves from Qualcomm, Apple, and MediaTek matter enormously here — the difference between 0.5W and 2W sustained inference isn’t a footnote for a battery-powered connected device, it’s the product viability line.
The co-design insight (Section 1) also applies in a very literal way when you’re designing edge hardware: silicon, firmware, radio stack (eSIM/cellular modem), and inference runtime have to be designed as a single system, or you’re leaving significant efficiency on the table. Fascinating era to be building in.
Point four is the one operators should sit with. Co-design is an organizational claim before it is a technical one. The reason a million-x is hard to copy is not the silicon, it is that almost no company will let one owner optimize across layers that separate budgets and separate VPs control. Stanford's fragmented grants are the same structure that produces stalled pilots inside large enterprises: every function optimizes its own layer, and nobody owns the whole. The accountability line lands for the same reason. Until someone holds the P&L across the layers, the stack stays a sum of parts, and the compounding never arrives.
Good article
Ai like 👌 link 🔗