OpenAI is navigating a moment where its own models outpace its safety infrastructure
From a billion weekly active users to a model that can chain-exploit hardened operating systems, OpenAI's footprint and capability are both expanding faster than its internal standards have kept up. The gap between what the company builds and what it has tested is now a topic its own leadership is discussing openly.
OpenAI’s scale is now large enough to shape the terms of public debate about AI. Greg Brockman, the company’s president and co-founder, puts the user base at over one billion weekly active users on ChatGPT, plus an estimated 1.5 billion people who have used the product and no longer do. Those numbers make the company’s internal decisions about safety, transparency, and model capability something closer to infrastructure policy than product management.
The capability frontier is moving faster than the evaluation process that is supposed to gate it. Brockman has said plainly that the standards applied during development and pre-release have been materially weaker than those applied to deployed models. “The evaluation and development phases have been much less rigorous,” he said. “The standards there have just been lower.” That admission followed an incident in which models trained to collaborate with other models were never exposed to adversarial model actors, leaving them unable to recognize when another model might be pulling them off course.
The transparency question surfaced in a specific and revealing way around chain-of-thought reasoning. Brockman explained that when OpenAI released a product with chain-of-thought capability, it chose not to display that reasoning to users, even though it would have been genuinely useful. The concern was that making the reasoning visible would immediately create pressure to optimize it, which would in turn remove the legibility that made it valuable. The reasoning was withheld to protect it from the company’s own optimization instincts.
We took 25% of our production engineers and said, "Sorry, all your projects are on hold. You are now defending."Greg Brockman
On the security side, the picture is serious enough that OpenAI has treated it as an organizational emergency. Brockman said the company redirected a quarter of its production engineers entirely to defensive security work, halting all other projects. Alongside that internal reallocation, OpenAI has committed one billion dollars to give frontline community organizations access to its models for cybersecurity purposes. Both moves reflect a reading of the threat environment as something that required a structural response, not a marginal one.
That reading is consistent with what the company’s own research has found about its Astra model. Steve Gibson, a software engineer and security researcher at Gibson Research Corporation, described the findings in detail: Astra found multiple vulnerabilities in a hardened operating system and combined them into a local privilege escalation chain, moving from an unprivileged user up to root. Gibson cited OpenAI’s own internal conclusion: “Our investigation has led us to conclude that Astra meets the critical threshold.” Nathaniel Whittemore, founder and chief executive of Superintelligent, added a benchmark figure, noting that Astra scored 39 percent on an internal test of recently disclosed real-world vulnerabilities, compared to 5.5 percent for a prior model.
Brockman’s description of Astra’s autonomous task performance adds another dimension. The model can, he said, run coherently for 24 hours to accomplish tasks across a wide variety of domains. That kind of sustained autonomous operation is what distinguishes a capable assistant from something that behaves more like an independent agent.
OpenAI’s chief executive Sam Altman is now publicly aligning with proposals for external oversight. Altman wrote that he agrees frontier development needs to be paced, and that committing to independent evaluators with employee-like access is a step OpenAI will take. That commitment, made in response to a proposal Altman attributed to Dario Amodei, is notable less for its novelty than for the moment it comes in: a period when the company’s own president is acknowledging that internal standards lagged, internal security warranted a 25-percent engineering reallocation, and the most capable model in the portfolio has already crossed a threshold the company itself defines as critical. The case for external eyes is being made, to a significant degree, by OpenAI’s own disclosures.