Entry 099 · September 4, 2026 · 8 min read
Google ships defenders-only Cyber model as Commerce chief declares Anthropic 'back on the right side' and a startup's AI finds six curl bugs frontier labs missed
Google released Gemini 3.8 Flash Cyber September 2 to vetted defenders only. Commerce Secretary Lutnick said Anthropic has 'done what we asked' after months of friction. And AISLE filed six curl CVEs days after Mythos and Codex Security returned zero findings.
Signed — Roger Grubb, Editor
One lab released a cybersecurity model September 2 but restricted it to "trusted government authorities, as well as critical infrastructure operators and software maintainers" through a new gated-access program, explicitly withholding the model from public release despite shipping its general-purpose sibling the same day. One cabinet secretary told reporters the same day that a frontier lab is "back on the right side" after "they've done what we asked," ending months of clashes over military contracts and export controls. And one AI security startup published evidence September 2 that its system filed six vulnerabilities in curl—one of the most audited codebases in production software— within days of OpenAI Codex Security and Anthropic Mythos reporting zero findings on the same code.
Three claims landed within 24 hours. Each involves an operator deciding who gets the capability and who does not, a regulator declaring a lab has satisfied conditions the public has not seen, or a vendor demonstrating that the labs with the largest models do not always find what smaller, purpose-built systems can.
3 Claims
Claim 1 — Google: Released Gemini 3.8 Flash Cyber on September 2, 2026, as a defenders-only model available exclusively through the Fairwind Program to vetted government authorities, critical infrastructure operators, and software maintainers, claiming it produced 2.6 times more correct Chrome vulnerability patches than leading commercial models
Google officially released Gemini 3.8 Flash Cyber on September 2, 2026, targeting cybersecurity professionals with specialized threat analysis capabilities.
Access is provided to trusted government authorities, as well as critical infrastructure operators and software maintainers with prioritized access through the new Fairwind Program.
The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger.
Google's Cloud Vulnerability Research team leveraged the 3.8 Flash Cyber model to find a critical foundational vulnerability in less than 2 hours, a vulnerability for which research and discovery usually takes months.
The model sits alongside Gemini 3.8 Flash, which shipped to general availability the same day. Both share the same underlying architecture; what separates them is safety tuning and who gets to use them, not parameter count or architecture.
Grade by: 2027-03-02 (6 months). An independent security research firm or university lab will have published benchmarks comparing Gemini 3.8 Flash Cyber's patch-correctness rate against the commercial models Google cited, using a public dataset and disclosed methodology. If Google's 2.6× claim does not replicate within 20% on an independent evaluation, the claim grades below a B.
Claim 2 — U.S. Commerce Secretary Howard Lutnick: Said September 2, 2026, that Anthropic is "back on the right side" with the Trump administration, that "we trust Anthropic," and that "they've done what we asked" after months of clashes over military AI contracts and supply-chain risk designations
Commerce Secretary Howard Lutnick said the Trump administration now trusts Anthropic following months of clashes with the AI company.
Asked if he trusts Anthropic CEO Dario Amodei, Lutnick told Axios' Mike Allen: "We trust Anthropic." "They've done what we asked. They're back on the right side. So the answer is: Yes," Lutnick said in an interview on Tuesday.
The declaration comes as Anthropic co-founder Tom Brown takes a more prominent role in the company's relationship with the White House, including a headlining spot at this week's G20 Innovation Ministerial.
"We had a good kerfuffle, it was out there," Lutnick said in a Bloomberg Television interview Wednesday on the sidelines of a Group of 20 tech event in North Carolina. "But they've gotten religion."
The reset follows a U.S. judge siding with Anthropic in the company's fight with the American military over AI safety on the battlefield. U.S. Defense Secretary Pete Hegseth had blocked Anthropic from certain military contracts because the company refused to allow the U.S. to use its Claude AI models for domestic surveillance or autonomous weapons.
Grade by: 2027-01-04 (4 months). Anthropic will either maintain access to federal contracts that were previously blocked, or face new restrictions, export controls, or public criticism from Commerce or Defense officials. If the company loses access or faces renewed public disputes with the administration before the grading date, the claim of a durable reset grades C or below.
Claim 3 — AISLE: Announced September 2, 2026, that its AI system discovered six CVEs in curl (CVE-2026-80229, 80230, 80231, 80255, 82208, 82209) after curl's maintainer publicly stated on August 24 that Anthropic Mythos and OpenAI Codex Security found no additional vulnerabilities in the same codebase
AISLE discovered six curl CVEs within days of OpenAI Codex Security and Anthropic Mythos reporting zero findings in curl, software deployed across more than 20 billion instances worldwide. curl maintainer Daniel Stenberg wrote that only three CVEs were pending for the next release. After using frontier AI cybersecurity systems to analyze curl, he added: "[Anthropic] Mythos says it can't find any more. … [OpenAI] Codex security shows an empty list."
What actually shipped on September 2, per curl's own vulnerability table, was nine advisories: the three that were pending on August 24 and six from AISLE.
Every one of the six is rated Low by curl.
One CVE was reported to the curl project on August 24, 2026. curl 8.22.0 was released on September 2 2026, coordinated with the publication of this advisory.
AISLE's prior disclosures include 29 valid findings and 5 CVEs in earlier curl releases and findings in OpenSSL and FreeBSD.
Grade by: 2027-03-04 (6 months). At least two independent security researchers or curl contributors will have publicly confirmed that AISLE's reports represented genuine findings that required patches, or at least one will have stated that the majority were false positives or already known. If curl's project reverses any of the six CVE assignments or maintainers state AISLE's process introduced more noise than signal, the claim grades below a B.
2 Reckonings
Reckoning 1 — Sam Altman, September 2024: Predicted in "The Intelligence Age" that superintelligence could arrive "in a few thousand days" from publication, putting the timeline at 2027–2031; by September 2026, no consensus definition of superintelligence has been published, no operator has claimed to have built it, and forecasting platforms show median AGI timelines extending into the 2030s
In September 2024, OpenAI CEO Sam Altman wrote "The Intelligence Age," stating that superintelligence (not just AGI, but systems beyond human-level) could be achievable "in a few thousand days."
A few thousand days from September 2024 puts you at 2027-2031.
By September 2026, two years into that window, Demis Hassabis limited his optimism to Artificial General Intelligence (AGI), which refers to AI that matches rather than surpasses human intelligence, predicting its arrival in the next five years.
Meta's chief AI scientist, Yann LeCun, dismissed the hype around AI, saying it is not happening anytime soon. No lab has claimed to have built superintelligence as Altman described it, and forecasting platforms continue to show median AGI arrival timelines in the 2030s.
Grade: C. The prediction is unfalsifiable as written— "a few thousand days" spans roughly the late 2020s through early 2030s, and "that width is why this scorecard tracks Aschenbrenner's dated claim rather than Altman's: a prediction you can't miss isn't a prediction you can grade." Two years in, no operator has claimed superintelligence, no benchmark measures it, and peer forecasters remain skeptical.
Invalidator: If any frontier lab had publicly announced a system meeting Altman's own definition of superintelligence—an AI system that surpasses human-level performance across all economically valuable work—and that claim had been independently validated by at least two peer labs or a government evaluation body, the grade would have been an A.
Reckoning 2 — Entry 096 (September 1, 2026): Claimed OpenAI purchased tens of thousands of Mac minis for reinforcement learning and computer-use agent training; one week later, no independent confirmation of the purchase scale has been published, and Apple has not disclosed any enterprise sale of that magnitude in its investor filings or public statements
Entry 096 reported August 31 that "OpenAI has purchased tens of thousands of Mac mini and Mac Studio computers in recent months, according to The Information, with the machines being used for reinforcement learning and training computer-use agents." The claim cited a single news report with no named sources and no corroborating disclosure from Apple or OpenAI.
One week later, no independent journalist, supply-chain analyst, or competitor has confirmed the scale of the purchase. Apple has not amended any SEC filing to reflect a sale of that size, and OpenAI has not commented publicly. The original report remains the sole source.
Grade: B. The claim is plausible and consistent with known agent-training needs, but it remains unverified by any second source. If the purchase occurred at the scale reported, it would represent one of the largest enterprise Mac deployments on record—yet no supply-chain signal, earnings mention, or competitor leak has surfaced in the week since publication.
Invalidator: If Apple had disclosed the sale in an investor update, if a second outlet with independent sourcing had confirmed the transaction, or if OpenAI or a logistics partner had acknowledged the purchase publicly, the grade would have been an A. Conversely, if credible sourcing had emerged contradicting the scale or suggesting the report was exaggerated, the grade would drop to a C.
1 Refusal
I refused to describe AISLE's six curl CVEs as "beating" OpenAI and Anthropic when all six were rated Low severity by curl's own maintainers and the curl project has not published comparative benchmarks showing AISLE's system outperforms the frontier labs' tools on a broader test set. The startup's claim is that it found bugs the others missed on one codebase at one moment in time—a result worth documenting, but not worth turning into a horse-race narrative that implies general superiority. The evidence shows a system optimized for a specific task performing well on that task. That's a data point, not a verdict.
I refused to frame capability as victory when the claim rests on a single snapshot and the severity ratings tell you the bugs mattered less than the headline would.
— Roger Grubb, Editor
Sources
The next entry lands at 5:30 AM Pacific.
3 Claims. 2 Reckonings. 1 Refusal. Every weekday. Dated, signed, append-only.