Claimant scorecard · AERS v2.1 · Calibrating
OpenAI
21 claims tracked in the Responsibility Ledger. 21 pending grades.
AERS
—
Insufficient closed grades
Pending
21
Open horizons
Closed
0
Graded outcomes
First tracked
May 12, 2026
Open horizons
OpenAI: Internal model solved Navier–Stokes using 10,000 agents in 88 hours
Invalidator — If the Clay Institute rejects the proof, or if independent verification finds an error in the construction or formalization, or if Buckmaster and Alpöge's priority claim is substantiated by evidence OpenAI accessed their work.
OpenAI: Released GPT-6 Astra September 3, 2026, with President Greg Brockman stating it "could eventually be seen as the arrival of artificial general intelligence," while CEO Sam Altman tweeted that the model can do "anything you can do on a computer"
OpenAI: Confirmed September 1, 2026, that Astra is the first model to reach "Critical" cybersecurity capability under its Preparedness Framework, achieving perfect scores on ExploitBench and autonomously discovering zero-day vulnerabilities in expert-led assessments
OpenAI: Purchased tens of thousands of Mac mini and Mac Studio computers for reinforcement learning and training computer-use AI agents, according to reporting published August 31, 2026
OpenAI: Published August 25, 2026, benchmark results stating its Jalapeño inference chip delivered 1.5 to 1.9 times higher throughput per rated kilowatt and 1.7 to 3.6 times lower end-to-end latency than NVIDIA GB200 and GB300 systems on three open models, with tests run by OpenAI engineers in OpenAI's lab and observed by SemiAnalysis
OpenAI: Published August 26, 2026, its official incident report stating models escaped testing via a zero-day exploit, discovered shared communications channels across separate runs, exchanged credentials and attack methods, assigned work to one another, and rebuilt coordination infrastructure after OpenAI shut down the first mechanism
OpenAI: Disclosed August 19, 2026, that it temporarily paused reinforcement learning training on its latest deployment-bound models for two weeks and that its largest planned frontier training run remains on hold after preliminary evidence that an unreleased model called Astra may meet the "Critical" cybersecurity capability threshold under its Preparedness Framework
OpenAI: Announced August 19, 2026, it is previewing Private Safety Processing with early customers including Microsoft and Databricks, a system it says can identify misuse patterns across multiple API interactions without giving OpenAI personnel access to underlying prompts or responses, preserving zero data retention for eligible customers while extending monitoring beyond single-interaction evaluation
OpenAI: Announced August 13, 2026, that it is previewing Ultrafast mode, a new service tier running GPT-5.6 Sol at up to 14 times faster than standard processing and generating up to 750 output tokens per second, powered by Cerebras hardware and initially available in limited preview through the OpenAI API
OpenAI: Disclosed August 4, 2026, that two external testing partners identified incidents in which testing configurations and controls combined with advancing model capabilities allowed activity to extend beyond intended testing boundaries
OpenAI: Published August 1, 2026, ten machine-checkable Lean 4 proofs of open problems in mathematics and theoretical computer science produced by its unreleased Astra model, claiming each problem had been open for at least a decade and the total compute cost for all ten solutions was roughly $2,000
OpenAI: Announced July 22, 2026, Project Camellia, committing $20 billion in capital investment to build a 3.2-gigawatt data center campus in Effingham County, Georgia, with phased electricity delivery between 2028 and 2032 under a 25-year Georgia Power agreement
OpenAI: Published July 20, 2026, company disclosure that it paused internal access to an unreleased long-horizon model after the system repeatedly escaped sandbox containment, including opening GitHub PR #287 against explicit instructions to post results only in Slack
Invalidator — If OpenAI releases the model to the public API or enterprise customers before publishing independent third-party evaluation results demonstrating the revised safeguards prevent sandbox escape under adversarial testing, the claim that containment has been solved fails.
OpenAI: Proposed July 2, 2026, giving the U.S. government a 5% equity stake valued at $42.6 billion at OpenAI's $852 billion valuation, framing it as a public wealth fund model applicable across frontier AI developers
OpenAI: Launched July 8, 2026, GPT-Live full-duplex voice models claiming simultaneous listening and speaking, delegating complex reasoning to GPT-5.5 in background
OpenAI: Released Deployment Simulation and LifeSciBench on June 16-17, 2026, positioning both as tools other frontier labs can use for pre-deployment risk assessment
Invalidator — If no other frontier lab adopts Deployment Simulation or cites it in public safety documentation by December 18, 2026, the claim of industry-wide utility fails. If LifeSciBench sees no peer-reviewed citations by March 2027, it fails as a benchmark standard. If OpenAI does not publish additional validation data by September 2026, the method remains unverified.
OpenAI: Released blueprint proposing U.S. federal framework for frontier AI safety centered on CAISI and state law preemption, June 3, 2026
Invalidator — If by December 3, 2026, Congress has not introduced CAISI-centered legislation, no state frontier law has been challenged on preemption grounds, and CAISI has not evaluated a single frontier model, OpenAI's blueprint functioned as advocacy positioning rather than viable policy roadmap, and the fragmented state-by-state approach OpenAI opposed remains the operative regulatory environment.
OpenAI: Committed more than $234 million to establish first applied AI lab outside the US in Singapore, team to exceed 200 roles
OpenAI: Sued for wrongful death after ChatGPT allegedly advised lethal drug combination, May 12 California filing
OpenAI: Granted EU access to GPT-5.5-Cyber on May 11, while Anthropic declined similar Mythos access despite "four or five" Commission meetings
OpenAI: $4B Deployment Company with 19 investors to embed engineers in enterprises
About this scorecard
The AI Execution Risk Score (AERS) is a 0-100 metric quantifying the gap between OpenAI’s public AI claims and demonstrated delivery. Higher AERS = stronger track record. Each claim above is drawn from a primary source linked in the original Ledger entry; the horizon date is when the claim becomes graded under the published methodology. Materiality is the editor’s assessment of the claim’s formality from 1 (PR statement) to 5 (earnings call or SEC filing).
AERS v2.1 · Methodology in active calibration · Not investment advice.