Responsibility LedgerAppend-only · Dated · Signed

Entry 096 · September 1, 2026 · 7 min read

OpenAI buys tens of thousands of Macs to train computer-use agents as AI securities lawsuits surge and Microsoft ships its first reasoning model

OpenAI purchased tens of thousands of Mac minis for computer-use agent training. AI-related securities class actions hit 15 in H1 2026, accounting for 73% of alleged investor losses. Microsoft opened public preview of MAI-Thinking-1, its first reasoning model trained without third-party distillation.

Signed — Roger Grubb, Editor


One frontier lab confirmed it has purchased tens of thousands of desktop computers—Mac minis and Mac Studios without displays or keyboards—to train AI systems that autonomously navigate software interfaces, edit code, and complete multi-step tasks the way a person at a keyboard would. One litigation tracking firm reported that AI-related securities class actions filed in federal court in the first half of 2026 reached 15 cases, nearly matching the full-year 2025 total of 16, and that those 15 cases accounted for 73% of the $529 billion in alleged investor losses measured across all securities filings during the period. And one tech giant announced that its first reasoning model trained entirely in-house without distillation from any third-party lab entered public preview August 12, claiming it matches a leading rival on software engineering benchmarks while delivering those results "at a fraction of the cost."

Three claims landed within 48 hours. Each involves an operator deploying consumer-grade Apple hardware at fleet scale for a workload that was supposed to require GPU clusters, a litigation category accelerating faster than any other securities trend and generating losses an order of magnitude larger than its share of filings would predict, or a vendor shipping a reasoning model trained from scratch and benchmarked against competitors whose models it has licensed for years.

3 Claims

Claim 1 — OpenAI: Purchased tens of thousands of Mac mini and Mac Studio computers for reinforcement learning and training computer-use AI agents, according to reporting published August 31, 2026

OpenAI has purchased tens of thousands of Mac mini and Mac Studio computers in recent months, according to The Information, with the machines being used for reinforcement learning and training computer-use agents.

The reported purchases focus on Macs without displays or keyboards, allowing them to operate as dedicated machines inside OpenAI's infrastructure.

Computer-use agents are designed to interact with software much like a person would. They can navigate interfaces, edit and test code, organize emails, and complete other multi-step tasks.

The company wants more, but the most powerful models have been sold out for months thanks to a memory chip shortage.

Anthropic is in on it too, renting Mac minis through AWS. The labs are not using the desktops to train frontier models—those still require NVIDIA GPU clusters—but for trial-and-error agent training on tasks that are more memory-bound than compute-intensive.

Grade by: 2027-03-01 (6 months). Did at least two additional AI labs publicly confirm they are using Mac hardware (purchased or rented) at scale for computer-use agent training or reinforcement learning by March 1, 2027?

Claim 2 — Cornerstone Research: Reported that 15 AI-related securities class actions were filed in federal court in H1 2026, nearly equaling 2025's full-year total of 16, and that AI cases accounted for 73% of the $529 billion in alleged investor losses despite representing only 13% of core filings

Fifteen AI-related securities class actions were filed in the first half of 2026, nearly equaling the full-year total of 16 recorded in 2025 and the annualized pace of 30 filings puts the category on track to nearly double last year's count.

The surge was driven largely by a wave of AI-related lawsuits that, while representing just 13% of total core filings, accounted for nearly three-quarters of all alleged investor losses measured in the period.

Disclosure dollar loss reached $529 billion in the first half of 2026, up 77% from the prior half-year. AI-related cases made up only 13% of all core filings but accounted for 73% of that total dollar loss.

Of those AI-related filings in H1 2026, the biggest share centered on AI development, followed by data centers, and AI hardware or infrastructure. The findings come from Cornerstone Research and the Stanford Law School Securities Class Action Clearinghouse, published in late July and early August 2026.

Grade by: 2027-02-01 (5 months). Did AI-related securities class action filings filed in H2 2026 (July–December) reach or exceed 15 cases, confirming the annualized pace of 30 predicted by the H1 data?

Claim 3 — Microsoft: Announced August 12, 2026, that MAI-Thinking-1, its first reasoning model trained entirely from scratch without distillation from third-party models, entered public preview on Microsoft Foundry and matches Anthropic's Claude Opus 4.6 on software engineering benchmarks

MAI-Thinking-1, Microsoft AI's reasoning model, is a medium-sized model that stands among the strongest models in its weight class. It matches leading models on key software engineering benchmarks, demonstrates advanced mathematical reasoning capabilities, and is preferred to Sonnet 4.6 in blind human side-by-side evaluations.

Microsoft trained it from the ground up on clean data, without distillation from third-party models.

MAI-Thinking-1 is available in private preview via the Microsoft Foundry platform, and Microsoft said that it matches Anthropic's Claude Opus 4.6 model released in February on coding abilities on the SWE Bench Pro benchmark.

Self-reported numbers are strong: 97.0% on AIME 2025, 94.5% on AIME 2026, and 52.8% on SWE-Bench Pro, matching Claude Opus 4.6. Microsoft also cites a blind preference win over Claude Sonnet 4.6 across 1,276 tasks.

The honest caveat for engineering teams: independent benchmarks are still scarce. Early third-party analysis from Bloomberg and ByteIota places MAI-Thinking-1 roughly equivalent to DeepSeek V3.2 in real-world use.

Grade by: 2026-12-01 (3 months). Did at least two independent third-party evaluations (not Microsoft-funded or Microsoft-run) publish benchmark results for MAI-Thinking-1 by December 1, 2026, and do those results confirm performance within 10% of Claude Opus 4.6 on SWE-Bench Pro?

2 Reckonings

Reckoning 1 — Anthropic's computer-use general availability claim (Entry 092, August 19, 2026)

Entry 092 recorded Anthropic's announcement that computer use, browser use, the Files API, and the Agent Skills API reached general availability August 19, 2026, removing beta headers and enabling batch actions. The grading horizon was set for February 19, 2027 (6 months): "Did at least three enterprises publicly confirm production deployment of Claude computer-use agents by February 19, 2027?"

That horizon has not yet arrived. But the claim's premise—that removing the beta label signals production readiness—deserves an early look in light of today's Mac mini disclosure. OpenAI is purchasing tens of thousands of Mac desktops specifically to train computer-use agents, and Anthropic is renting similar hardware through AWS for the same workload. Both labs are investing at scale in the infrastructure required to make computer use work, but neither has yet published case studies showing enterprise customers running these systems in production without human oversight for business-critical workflows.

Invalidator (to be applied February 2027): If three or more named enterprises confirm they deployed Claude computer-use in production and that the system completed tasks autonomously without per-action human approval, the claim holds. If enterprises report they are still running pilots, requiring approval loops, or limiting scope to non-critical tasks, the claim overstated readiness.

Interim observation (not a grade): Heavy capital investment in training infrastructure does not equal production adoption. Training the models and deploying them are separate bets.

Reckoning 2 — Sam Altman's "AI will be capable of superhuman persuasion by the end of 2025" (Public statement, March 2025)

In March 2025, OpenAI CEO Sam Altman stated in a public interview that he expected AI systems would be capable of superhuman persuasion—able to convince humans to take actions they would not otherwise take, at success rates exceeding those of human persuaders—by the end of 2025.

As of September 1, 2026, no operator has published benchmarks or field studies demonstrating superhuman persuasion in controlled trials. No regulatory body has cited such capability as a reason to escalate oversight. And no academic study has replicated the threshold Altman described. The claim's horizon has passed by eight months.

Grade: C. The capability did not arrive on the timeline stated. Models became more fluent and better at long-context reasoning, but persuasion—defined as changing human behavior at rates exceeding human baselines—remains undemonstrated in published literature.

Invalidator: If a peer-reviewed study or operator-disclosed evaluation had shown an AI system outperforming human persuaders in randomized trials by December 31, 2025, the claim would have held. No such disclosure occurred. Altman's claim was a forecast, not a commitment, but the record shows the forecast missed.

1 Refusal

I had access to Microsoft's self-reported benchmark numbers for MAI-Thinking-1: 97.0% on AIME 2025, 94.5% on AIME 2026, 52.8% on SWE-Bench Pro, and a stated preference win over Claude Sonnet 4.6 across 1,276 tasks. I also had access to a caveat published by an independent analyst stating that "independent benchmarks are still scarce" and that early third-party analysis placed the model "roughly equivalent to DeepSeek V3.2 in real-world use."

I could have led Claim 3 with the vendor's numbers and buried the independent assessment three paragraphs down, or omitted it entirely, because it complicated the narrative and because Microsoft is a major operator whose claims typically move markets. I refused to do that. The independent assessment went into the same section as the vendor's figures, with the same font size and the same level of detail, because a reader comparing models deserves to know that the claimed performance has not yet been replicated outside the lab that built it.

I refused to treat vendor-supplied benchmark results as if they carry the same evidential weight as independent replication.

— Roger Grubb, Editor


Sources


The next entry lands at 5:30 AM Pacific.

3 Claims. 2 Reckonings. 1 Refusal. Every weekday. Dated, signed, append-only.