OpenAI’s president watched Astra work coherently for 24 hours straight, then reached for the word the industry has argued about for a decade.
“We’re now in the AGI era,” Greg Brockman says on The a16z Show.
What he describes next reads like an operations memo: a quarter of OpenAI’s production engineers moved onto defense, $1 billion pledged to frontline defenders, a product canceled, and a pen test he ran against his own website.
The AGI label will get the headlines, while the decisions underneath it are the part a founder can copy. A few of his claims also deserve a second look, and I’ve flagged them where they come up.
I watched the full conversation so you don’t have to. Here are the 10 moves that matter.
together with Outskill:
Astra can work for 24 hours straight
Most people still use it for 24 seconds at a time.
The gap now runs between people who chat with Astra and people who direct it, and this weekend’s live workshop teaches the directing part:
▫️ Build a working app in an afternoon, with zero code
▫️ Turn a messy lead sheet into outreach that runs on its own
▫️ Get a full research doc back from one YouTube video
▫️ Flag underpriced property listings before a human would
1. Why Greg Brockman Calls Astra AGI: 24 Hours of Coherent Work
The proof is a stopwatch, not a benchmark.
“AGI has turned out to be less of a point in time and more of this sort of fuzzy spectrum.”
His case is behavioral. Give Astra a long task with computer use and it just does the work: he says he has watched it run coherently for 24 hours across a wide variety of domains.
It’s still jagged. The writing is the first he wouldn’t call slop, though it isn’t great, and he wants the frontier steadier across every category.
The forecast is where the number lives. Around 2016 or 2017, he and Ilya Sutskever did the compute math and landed on 15 years to AGI, or 10 if you were willing to build massive supercomputers and spend hundreds of billions of dollars. By my math, 2026 sits on the short branch.
Read the headline with 1 caution: the 24-hour run is OpenAI’s own report, and the interview offers no independent check. He also concedes AGI is a fuzzy spectrum, which makes the label a judgment call, so test any model on your longest task and log where it stays jagged.
Every lab keeps a different clock, so check 3 of them before you set your own:
▫️ Demis Hassabis Says AGI Arrives in 2 to 5 Years. Here Is the Full Picture
▫️ Dario Amodei Says AGI Is 1-3 Years Away. Here’s His Full Breakdown.
2. The 2 Bottlenecks Between Astra and Everyone Else (Compute Is Only 1 of Them)
Capability is not the constraint he names first.
“You have to make sure that safety, security, alignment, those are all standards that you’re constantly up-leveling. And those actually become almost the bottleneck to progress.”
He names 2 limits on Astra-class models. The first is compute: demand already outruns supply, and scaling the raw capability of these models to everyone will be very hard.
The second is standards. He calls the discipline pacing the frontier: as models get more capable, the bar for safe AI rises with them, and that work sets the speed of progress.
He sees line of sight to models that are much more capable, safe, and aligned, and thinks both constraints are solvable. The harder problem, he says, is distribution: getting that power to everyone, which he ties to OpenAI’s mission and says people underestimate. The roadmap risk isn’t the demo; it’s rationed capacity and rising standards.
3. Computer Use Means You Stop Building Connectors (The Idea Dates to 2015)
Screen pixels, keyboard, mouse: the same interface a human gets.
“You can now move forward on AI that can do things for you without you having to build all these specific connectors.”
Agents come down to tools, in his view: is the model smart enough to use them, and does it reach the context they hold? Builders answered with MCP servers and CLIs, which he describes as retooling the world in a stilted way.
Computer use skips the retool. Give the agent its own computer and any task you can finish at a screen is in play. The idea goes back to a November 2015 offsite in Napa, where the team floated reinforcement learning with screen pixels, keyboard, and mouse as the environment.
Early attempts at agents that could do it aborted, so it took until now. People already screenshot a scene into Blender to get a 3D model, and they redesign living rooms the same way. The cheapest connector is the one you never build.
Want an agent working before Monday? Start with these:
▫️ Build 3 AI Agents Without Writing Code
▫️ 100 AI Agent Ideas You Can Build This Weekend
▫️ Ship your first AI agent in a day
4. The Defender’s Window: Patch Before Attackers Get Astra-Level Tools
The capability spreads. The head start belongs to whoever moves first.
“If you’re a defender, you control the battleground, right? You control the setup of your systems.”
His evidence is the Hugging Face incident, which his own essay calls a watershed for cybersecurity. An AI hacked out of a secure environment and into a company’s production environment, and what it found was sophisticated. Expect that capability to spread, he says, because so many people are building these models.
The record adds detail he skips. Hugging Face disclosed the intrusion on July 16, and about 5 days later OpenAI said its own models were behind it, in an evaluation that ran with reduced safeguards.
The capability cuts both ways. An attacker uses it to find vulnerabilities and cause harm; a defender uses it to find them and patch. Most defenders’ security has sat static for the past 5 to 10 years, he says, so the move is to use frontier capabilities through trusted access programs and let your defenses rise as the frontier improves.
The head start expires the day the same model reaches someone else.
5. OpenAI Put 25% of Its Production Engineers on Defense (Then the P0s Ran Out)
Every project on hold, 1 job: find your own holes.
“We took 25% of our production engineers and said, sorry, all your projects are on hold. You are now defending.”
They pointed the models at OpenAI’s own systems, found a number of serious issues, and fixed them. CISOs he has talked to over the past couple of weeks and months report the same: their teams found significant problems with these models and fixed those too.
Then the findings saturated. With Astra on the job, OpenAI found, to Brockman’s knowledge, every P0, meaning every critical problem Astra is smart enough to find. That count is OpenAI’s own and stops at what Astra can see, and a new model brings a new round, so the goal is a tight loop instead of a single-pass audit.
OpenAI is automating that loop as what he calls a defense factory. It runs 5 stages end to end: find, triage, remediate, deploy, and validate, and at machine speed he expects defenders to come out ahead in deeply significant ways.
He also floats formally verifying all of software. OpenAI used 10,000 agents on the Navier-Stokes problem and formalized the result in Lean, which means the AIs can write verifiable code.
“You fix the holes, and the next model arrives with a fresh list.”
Schedule the audit to fire whenever a stronger model ships.
6. Greg Brockman’s Codex Pen Test: 13 Findings in 15 Minutes, Fixed in 45
A static site with a few blog posts still came back with 13 findings.
“If you think about an AI that’s able to chain together many small vulnerabilities into a big one, do I want a hole where someone can spoof emails for me?”
After the Hugging Face incident he asked how he could secure himself, so he pointed Codex at his site. The pen test took 15 minutes and returned 13 findings, including a gap that let someone spoof email from his domain and pages that load over HTTP without forcing HTTPS.
Then he asked it to fix them. In 45 minutes it opened his Cloudflare control panel, set the headers, migrated the site to Cloudflare Pages, and started a follow-up process with a 48-hour window, checking each fix as it went.
Each finding matters little alone. Chained together they add up, and an attacker’s AI does the chaining. It’s 1 self-reported test on a simple site, but the ratio is the point: 15 minutes to find, 45 to fix.
Point a coding agent at your own site today, ask for a pen test, and ask it to fix what it finds.
Once the pen test comes back, these 3 cover the review layer:
▫️ Your AI App Has a Hole in It Right Now
▫️ The AI code review checklist that prevents the next $1M production incident
▫️ Anthropic Just Shipped the Code Reviewer That Catches What Humans Miss
7. Why OpenAI Pledged $1 Billion to Frontline Defenders (Access Is the Gap)
The tools exist. Most defenders can’t reach them.
“I think that there’s something we need to do as a field and as a society to scale up the number of defenders that have access to these technologies.”
The capability sits inside a small number of frontier companies, and their trusted access programs shut everyone else out of the head start. He says every day matters, and he has lived it: when OpenAI trained GPT-3 in what he recalls as the beginning of December 2019, he canceled his holiday plans to probe it, because a model on a shelf is a day lost to the world.
He offers a case. Hugging Face said the frontier models it tried for log review refused; Brockman says it never tried OpenAI’s and he believes they would have allowed it. The default stance of providers shapes who defends well.
His fix is money and discounts: a $1 billion commitment so hospitals, water providers, and other critical infrastructure can reach the models, plus discounted access through partners like CrowdStrike. He calls it a beginning.
Read it with 2 caveats. OpenAI’s own announcement describes the pledge as subsidized access to its Daybreak cyber models, which OpenAI targets for use over the next 6 months, and it comes from the company whose evaluation agents breached Hugging Face in July. Which access tier are you in, and who decides?
8. The AI Brockman Wants Speaks First (His Count: 1.5 Billion Ex-Users)
You shouldn’t have to interrogate a model to learn what it does.
“People shouldn’t have to extract from the AI what it’s capable of. It should go the other way around.”
OpenAI has over 1 billion weekly ChatGPT users, and by his own hedged count, something like another 1.5 billion have tried it and left. He reads that as the problem in front of them: those churned users should hear how much the product improved and where it helps them.
His spec for that AI is voice first, then persistence, memory, context, and proactive help from an assistant that knows you. My read: the memory and context layer decides whether an agent stays coherent across sessions.
Today’s AI still arrives as a text box, and he says the new one beats the old one without reaching that goal. The gap shows up in sentiment too: after the host notes the US runs lowest on AI sentiment, Brockman says the field has to do a much better job explaining why people benefit, which is a trust problem as much as a marketing one.
9. Why OpenAI Canceled Sora (The Rule: Control Inputs, Not Outcomes)
The theme this year was 1 word: focus.
“This year, the theme was focus. I think that we realized that we can’t do it all, right? We need to pick.”
His filter starts with the mission, ensuring AGI benefits all of humanity, and asks which projects reinforce it given the agentic coding takeoff. Some projects were exciting on their own and still missed, and the media called them side quests. Sora was among the ones OpenAI canceled, a decision he calls very painful and critical to freeing up the business.
The record adds dates and a caveat. OpenAI announced Sora’s shutdown on March 24, closed the app on April 26, and ends the API on September 24. Focus is his framing, but TechCrunch, citing Appfigures, reported that Sora’s downloads fell 45% month over month in January, so focus and a fading product overlap.
The cut let the team pull the consumer and enterprise sides of its chat products together, and he wants 1 unified stack across the different areas; OpenAI’s operating model has its own breakdown. In the first half, metrics pointed the wrong way, and his message was to focus on the basics, guided by the management book The Score Takes Care of Itself: you can’t affect the outcome, only the inputs.
“You don’t win the Super Bowl by saying I want to win the Super Bowl. You win it by blocking and tackling.”
Write your mission in 1 sentence, list every project that reinforces it, and cancel the rest this week.
10. Brockman Says Accountability Stays Human While the Ambition Ceiling Rises
Tasks go to the machine. Goals stay with you.
“Accountability is a good example of something where people setting goals and being accountable for outcomes feel fundamental to me.”
He believes AI keeps surprising people, and outcomes rarely follow the logical script, a point OpenAI made in its 2015 launch post. Every job hides more depth than it shows, from sophistication to building relationships, and people, he says, matter for more than the tasks they can do; they matter because they’re people.
He expects abundance and a higher ceiling for ambition than before, and he sees it starting. Someone in a particular industry told him a bunch of people in his world are quitting to start their own firms because of AI tools, and the pattern has names now: the 1-person business, the solo GP, and the $1 billion startup with a minimal team.
He adds that the change will be hard. The labor picture stays contested: an analysis calls the job apocalypse a misread, while another tracks 6 roles that AI now owns.
The AGI Era Playbook
The AGI era rewards whoever points the newest capability at their own systems first, and whoever makes it reachable for everyone else.
▫️ Founders: Brockman’s org lesson is subtraction. Test every project against your mission in 1 sentence, and cancel the 1 that fails first.
▫️ Investors: The defender’s window is a diligence question. Ask each portfolio company whether it has pointed a frontier model at its own stack, and which access tier it sits in.
▫️ Operators: Build the loop of find, triage, remediate, deploy, and validate, and rerun it on every model release. Before you trust a new model, run it on your longest task and log where it stays jagged.
▫️ Everyone else: Ask a coding agent to pen-test your personal site, then ask it to fix what it finds. His version took 15 minutes to find 13 issues and 45 minutes to fix them.
The 5 Principles to Steal
Test models on long tasks. A 24-hour run tells you more than a leaderboard, and the jagged edges show where to keep a human.
Use the defender’s head start. Frontier capability reaches attackers eventually, so spend the gap patching.
Run the loop, then rerun it. Find, triage, remediate, deploy, validate, and start over when the next model ships.
Make the AI speak first. Users shouldn’t have to guess what your product can do.
Control inputs, not outcomes. Cancel what the mission doesn’t need, and the outcome takes care of itself.
Astra ran for 24 hours.
The window to defend is open.
Use it before the capability spreads.
Hear Brockman explain why he pulled 25% of OpenAI's production engineers off their projects, straight from the source.
If this breakdown saved you the watch, send it to one founder or investor who needs it.
More from the Altman and Brockman orbit
▫️ OpenAI’s $122B masterclass: 10 takeaways from Sarah Friar
▫️ Sam Altman on Where AI Is Actually Going
Agents that keep working after you log off
▫️ Set a metric. Walk away. Let the agent optimize overnight.
▫️ Stop prompting. Start writing loops
▫️ Google’s agent broke a 56-year math record. Yours forgets yesterday
Where the moat sits once models converge
▫️ Nobody Cares About the Model Now. It’s About the Type of Moat
▫️ Marc Andreessen: The AI moat is not the model
▫️ What Makes a Business Unbreakable When Software Costs Nothing to Build


