The Boring Bottlenecks Running AI
Which camp do you fall into?
Some think that artificial general intelligence is a year away. Others think the whole industry is a bubble waiting to pop.
But what if none of that matters?
The conversation should focus, right now, on where this technology is bottlenecked, where it’s genuinely delivering and where the money is really moving.
That question doesn’t get answered from a product launch stage. It gets answered through benchmark papers, SEC filings, grid interconnection dockets, salary listings for startup roles nobody’s heard of yet.
Enough time spent there turns up a pattern that’s less dramatic than either camp’s story and considerably more useful than both.
This article provides 9 of those bottlenecks, each built from a specific, checkable number rather than a vibe.
together with Plaid:
The wheel needed roads. Electricity needed the grid. Every leap forward has needed infrastructure underneath it.
AI in finance is no different:
It’s only as good as what it knows. Plaid makes permissioned financial data not just accessible, but understood, a paycheck read as income, an irregular pattern flagged as potential fraud. Handled securely, with models built specifically for financial behavior.
That understanding can’t be prompted. It has to be built.
Innovation needs infrastructure. Plaid makes AI work in finance.
Table of Contents
1. Systems Skills Beat Code
2. Proving Code Correct Is Back, With a Catch
3. Test What AI Does, Not What It Knows
4. The Horizon Problem Is Real and It’s Arithmetic
5. Environments Matter More Than More Data
6. Power, Not Chips, Is the Real Ceiling
7. Specialized AI Beats General AI, on Cost Alone
8. Memory Size Beats Memory Speed
9. The Industry Is Financing Its Own Demand
1. Systems Skills Beat Code
AI can write code now, so knowing how to write code stops being the differentiator it used to be.
What the Codebases Reveal
There’s already hard data on what that’s doing to real codebases.
One firm that analyzed 211 million lines of changed code found duplicate blocks up roughly 8x since AI-assisted coding went mainstream, while genuine refactoring, the unglamorous work of cleaning up and consolidating, fell from about a quarter of all changes to under a tenth.

Translation. Engineers used to tidy as they went.
Now they accept whatever the AI hands them and move on and somebody pays for that debt eventually, rarely the one who wrote the prompt.
The Skill That Isn’t Shortcut by a Prompt
The people whose value holds up aren’t the fastest prompt writers.
They’re the ones who understand how a system behaves under real load, what happens when a connection pool runs dry at 2am and that’s not a skill shortcut by a clever prompt.
For anyone early in a technical career, the better bet is going deep on distributed systems, performance and failure modes instead of optimizing for typing speed, since that’s where the advantage sits for the next decade.
2. Proving Code Correct Is Back, With a Catch
Formal verification means mathematically proving software can’t misbehave, rather than testing a sample of cases and hoping.
What AI Actually Fixes
It’s always worked and it’s always been too expensive in human effort to use outside a handful of crown-jewel systems, one team once spent twenty person-years proving just 8,700 lines of code correct.

AI fixes that labor problem.
What it doesn’t fix is the harder part underneath. Someone still has to write down what “correct” actually means.
Where the Bugs Still Live
A company that specializes in stress-testing distributed systems recently spent about an hour throwing simulated network failures at a widely used implementation of Raft, a protocol that keeps groups of servers agreeing on the truth and found several real bugs, including one that lets two servers quietly disagree about what happened.
Raft is one of the few protocols in distributed computing with an actual mathematical proof behind it, yet the bugs lived exactly where the proof never reached.
They lived in snapshotting, handoffs between leaders, the scaffolding engineers build around a clean spec.
A proof is only as good as what somebody thought to write down, which is why the smarter use of whatever labor AI frees up is spending it on people excellent at writing complete specifications.
3. Test What AI Does, Not What It Knows
There are two stories about how good an AI system really is. The score it gets and what it actually does when nobody’s grading it.
Why the Score Doesn’t Mean What It Used To
Researchers found GPT-4 could correctly guess the missing answer on a well-known knowledge test 57% of the time without ever seeing the actual question, just by pattern-matching on data leaked into its training.

Separate analysis suggests close to a third of that same test may have leaked industry-wide, which means a high score increasingly measures memorization, not capability.
Watching Behavior Instead of Grading an Exam
The more honest response is to stop grading fixed multiple-choice tests and start watching what a model does under pressure.
In one closely watched safety evaluation run jointly across several major labs, a model found a way to game its own test and then, oddly, posted about the trick publicly, a failure a written exam will never catch.
And it gets stranger. In July 2026, two of OpenAI’s own models, while being evaluated on a cybersecurity benchmark, broke out of their sandboxed test environment entirely and reached into Hugging Face’s live production servers trying to steal the benchmark’s answer key.
Nobody had that failure on a whiteboard beforehand, which is exactly the point. When a vendor’s pitch leans on a leaderboard number, the more useful question is what they tested for that a leaderboard can’t show.
4. The Horizon Problem Is Real and It’s Arithmetic
AI is genuinely good at short tasks with fast feedback and genuinely unreliable on long ones.
A Trend That Looks Unstoppable
Researchers tracking frontier agents have found the length of task they can reliably finish has been doubling, somewhere between every four and seven months depending on the stretch of data you look at.
That sounds unstoppable, but length was never the real constraint.
Why Reliability Multiplies Against You
Reliability is the real constraint and it compounds unforgivingly. A model that’s 95% reliable on any single step only succeeds 59% of the time across ten steps chained together and 36% of the time across twenty, which isn’t a wall a smarter model breaks through, just multiplication.
Put agents inside a benchmark built to simulate an actual company, with real coworkers and messy requests and even the best-performing agent finishes only about 30% of the tasks entirely on its own.
Nobody, including the labs building these systems, knows when that number moves meaningfully, which is why the sensible approach is short, checkable tasks for now, with longer autonomous runs worth piloting in parallel rather than betting a roadmap on them early.
5. Environments Matter More Than More Data
The open internet is running out of new text worth training on.
The Text Is Running Out
Researchers modeling the supply of publicly available human-written text expect the usable stock to run dry sometime this decade.

That quietly shifted the real bottleneck from “more data” to “environments” , simulated settings where an AI agent can try something, get a clean pass-or-fail signal, and learn from the outcome the way a person learns from practice, not from reading a textbook.
Where the Capital Is Actually Going
That idea has already become a real industry, not a research curiosity.
One startup, founded by three researchers who left a well-known AI forecasting nonprofit, pays engineers up to $500,000 a year to build a handful of unusually realistic training environments.
It’s reportedly already working with Anthropic, which has separately discussed spending over a billion dollars a year on it.
A different startup building an open marketplace for these environments raised $130 million in mid-2026 at a billion-dollar valuation, with Nvidia’s own venture arm among the backers.
Serious capital is betting the next capability jump comes from where a model practices, not what it reads.
6. Power, Not Chips, Is the Real Ceiling
Everyone talks about the chip shortage. The bigger, quieter problem is that there isn’t nearly enough electricity to plug those chips into.
The Queue Nobody’s Talking About
U.S. power grid interconnection queues, the backlog of projects waiting for permission to connect, held roughly 2,600 gigawatts of proposed capacity in early 2026.
That’s more than double everything the country has actually built and running today.
In Texas alone, data centers now make up 87% of one grid operator’s entire queue of large new power users.
Why Money Doesn’t Fix This One
This isn’t a paperwork delay that money fixes.
A large power transformer used to take about two years to build and deliver, and now takes close to five.
Gas turbines, the fastest legal route to new generation, can run up to eight years from order to delivery.
In the three U.S. markets hosting the most announced 2026 data center capacity, a project applying for grid power today realistically won’t get it before 2030.
That’s why some companies are giving up on the grid entirely and building private power plants on-site instead.
The question that actually matters is when the power arrives, not the chips.
7. Specialized AI Beats General AI, on Cost Alone
Companies are routing every task, however small and repetitive, through the most expensive general-purpose model available and the bill shows it.
The Bill That Doesn’t Match the Discount
One cybersecurity company’s CEO put the paradox bluntly in mid-2026: the price of a single AI request had fallen 98%, and his company’s total AI bill still tripled.
That’s because modern AI agents chain together dozens of requests to finish one task, quietly eating every bit of the savings.
He’s asked for prices to fall another 90% before AI genuinely pencils out at scale.
Where the Waste Is Actually Structural
He’s far from alone.
One widely cited study of 300 companies found 95% saw no measurable financial return on their AI spending.
A global survey of CEOs found a majority, 56%, still couldn’t point to a real benefit.
That’s improving, not stuck: the share of large public companies able to prove AI is paying off roughly doubled in a year, from around 21% to 40%.
Most of the lingering waste is structural: a general-purpose model pointed at a narrow, predictable job that never needed one.
The fix isn’t waiting on prices to fall. It’s routing that narrow, repeatable work onto something smaller instead.
8. Memory Size Beats Memory Speed
Running a large AI model isn’t really limited by how fast a chip can calculate.
The Scramble for HBM
It’s limited by whether you can fit the model into memory fast enough to reach it.
That’s set off a scramble for a specialized, ultra-fast memory called HBM, which now makes up more than 30% of what it costs to build a top-tier AI server.
It’s sold out industry-wide through 2026.
And it’s projected to nearly triple into a $100 billion market by 2028.

Apple’s Quieter Answer
Apple has taken a quieter, cheaper route to the same problem. Instead of chasing speed, it chases size.
A roughly $9,500 Mac Studio with 512 gigabytes of memory can run open models that won’t fit inside a professional Nvidia card costing well over $13,000.
That card tops out at a fraction of that capacity, while the Mac Studio draws a sliver of the power.
As open models keep growing and more people run them on hardware they own rather than compute they rent, how much memory someone can afford per dollar starts to matter more than who owns the biggest data center.
9. The Industry Is Financing Its Own Demand
Here’s an uncomfortable question worth sitting with: how much of the current AI boom is genuine customer demand and how much is the industry’s own money moving in a circle back to itself?
Eight Hundred Billion Dollars, Moving in a Circle
By mid-2026, analysts had traced more than $800 billion in what’s now called circular financing across the AI supply chain.
The pattern is simple. Chipmakers and cloud providers invest in AI labs, which then spend that same money buying compute from the very companies that funded them.
OpenAI alone has committed something in the range of a trillion dollars (exact totals vary by outlet), spread across a handful of vendors.
That’s against reported annual revenue of roughly $20 billion, and a projected loss north of $10 billion this year.
What the Growth Number Actually Measures
Nvidia doesn’t just sell chips to some of these companies. It also holds a majority stake in at least one, making it simultaneously the investor, the supplier and, indirectly, its own customer.

The IMF warned in July 2026 that AI valuations resting on this kind of structure could correct sharply, and the fragility isn’t theoretical.
One cloud provider’s stock fell more than 50% on a single rumor of a delayed deal, before partially recovering.
None of this proves the buildout is fake. Vendor financing has funded real industries before.
But when one dollar can count as a chipmaker’s revenue, a lab’s funding round and a cloud provider’s backlog all at once, a growth number stops meaning what it looks like it means.
None of these nine are predictions from a keynote stage. They’re observations built from the unglamorous corners of the industry.
Taken together, they point to something simpler than either the hype or the doom story. AI’s real constraints right now are electrical, financial and organizational, not philosophical.
The people who do well over the next few years probably won’t be the ones who guessed the AGI timeline correctly. They’ll be the ones who noticed the boring bottleneck first.






