Responsibility LedgerAppend-only · Dated · Signed

Entry 058 · July 9, 2026 · 10 min read

Future of Life grades labs C or below, OpenAI ships full-duplex voice, and the voluntary safety system erodes faster than mandatory rules arrive

The Future of Life Institute published its Summer 2026 AI Safety Index July 7, awarding no lab higher than Anthropic's C+ and warning the voluntary system is eroding. OpenAI launched GPT-Live July 8, a full-duplex voice model claiming natural conversation while delegating to GPT-5.5 for reasoning. Illinois enforced SB 315 just as experts documented industry retreat from earlier red-line commitments.

Signed — Roger Grubb, Editor


One independent safety institute published grades July 7 showing the top-ranked AI lab—Anthropic—earned a C+, with OpenAI and Google DeepMind receiving C grades and three labs earning outright F's, while the expert panel documented that four frontier companies have weakened or eliminated commitments to pause development when systems approached danger thresholds. One frontier AI lab launched July 8 a full-duplex voice model it claims can "listen and speak at the same time" and delegate complex reasoning to GPT-5.5 in the background, marking the third architectural generation of ChatGPT voice in two years. And one U.S. state finalized July 6 the nation's first mandatory annual third-party audit law for frontier developers, carrying civil penalties up to $3 million—a binding obligation that arrived the same week an international scorecard documented the voluntary safety system is "eroding before governments have put a durable alternative in place."

The Future of Life Institute released its 2026 H1 AI Safety Index on July 7 (local time), evaluating nine major global AI companies: Anthropic, OpenAI, Google DeepMind, Meta, Z.ai, Alibaba Cloud, xAI, DeepSeek, and Mistral . Anthropic ranked first but received only a C+ overall, with OpenAI and Google DeepMind each receiving a C . xAI, DeepSeek and Mistral received failing overall grades—one company each from the U.S., China and Europe . The reviewers said Anthropic, OpenAI, Google DeepMind and Meta have weakened or eliminated earlier commitments to pause development if their systems approached specified danger thresholds .

Three claims landed within forty-eight hours—each involves an independent evaluator, a frontier lab, or a state executive publishing an on-the-record statement about lab safety practices, voice interface capabilities, or mandatory audit obligations that can be graded against whether companies reverse their retreat from safety commitments, whether GPT-Live's full-duplex architecture becomes the industry standard, and whether Illinois actually enforces the $3 million penalty when the first audit deadline arrives in 2028.

3 Claims

Claim 1 — Future of Life Institute: Published July 7, 2026, AI Safety Index awarding Anthropic C+, OpenAI and Google DeepMind C, finding frontier labs have weakened commitments to pause at danger thresholds

The U.S.-based AI safety think tank Future of Life Institute (FLI) released its "2026 H1 AI Safety Index" on July 7 (local time) . The assessment was conducted by an independent panel of seven AI and governance experts affiliated with institutions including the University of Montreal, UC Berkeley, and Oxford University . They comprehensively evaluated six categories: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and transparency and communication .

Anthropic took the top overall spot, but its grade was just a C+. OpenAI and Google DeepMind each received a C, while Meta managed only a D+. Z.ai and Alibaba Cloud scored a D-, and xAI, DeepSeek, and France's Mistral all received the lowest possible grade of F. Not a single company earned an A or B .

The reviewers said Anthropic, OpenAI, Google DeepMind and Meta have weakened or eliminated earlier commitments to pause development if their systems approached specified danger thresholds. The panel described the changes as "moving the goalposts" and said they have undermined safety frameworks across the industry . The report suggests that the voluntary safety system created by AI labs has begun eroding before governments have put a durable alternative in place .

Claimant: Future of Life Institute (Max Tegmark, president)
Date: July 7, 2026
Grading horizon: July 2027 (1 year)
Falsifiable commitment: The claim is that frontier labs have weakened red-line safety commitments. This will be gradable by: (1) whether FLI's Summer 2027 Index shows improved scores across the same six categories, indicating reversal of the documented retreat; (2) whether any of the four named labs (Anthropic, OpenAI, Google DeepMind, Meta) publicly reinstate capability-threshold pause commitments with defined numerical triggers by July 2027; (3) whether the top overall grade rises above C+, suggesting meaningful industry-wide improvement rather than continued erosion.
Source: https://futureoflife.org/ai-safety-index-summer-2026/

Claim 2 — OpenAI: Launched July 8, 2026, GPT-Live full-duplex voice models claiming simultaneous listening and speaking, delegating complex reasoning to GPT-5.5 in background

OpenAI launched GPT-Live on July 8, 2026, a new generation of voice models that make talking with AI feel much more like having a real conversation. GPT-Live is built on a full-duplex architecture, meaning it can listen and speak at the same time . The company said that more than 150 million people talk to ChatGPT using features like Voice and Dictation .

GPT-Live is also OpenAI's smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to the latest frontier model behind the scenes and brings the result back into the conversation when it's ready. At launch, GPT-Live will use GPT-5.5 in the background .

The company is replacing its current Advanced Voice Mode in ChatGPT with GPT-Live-1 mini by default. Users of paid tiers will be able to access the larger GPT-Live-1 model . The two models, GPT-Live-1 and GPT-Live-1 mini, are rolling out globally starting today across iOS, Android, and ChatGPT.com .

OpenAI framed GPT-Live as solving two generations of architectural limitations: the original cascaded pipeline (Whisper STT → GPT-4 LLM → TTS) that introduced latency and information loss, and Advanced Voice Mode's rigid turn-taking that mistook pauses or background noise for end-of-turn signals.

Claimant: OpenAI
Date: July 8, 2026
Grading horizon: January 2027 (6 months)
Falsifiable commitment: The claim is that GPT-Live's full-duplex architecture makes voice interaction "feel much more like having a real conversation" than previous turn-based systems. This will be gradable by: (1) whether competing labs (Anthropic, Google, Meta) adopt full-duplex voice architectures by January 2027, indicating market validation of the approach; (2) whether OpenAI's own usage data shows retention improvement for voice sessions longer than 10 minutes compared to Advanced Voice Mode baseline (if disclosed in quarterly updates or research papers); (3) whether independent benchmarks or user studies published by January 2027 demonstrate measurably improved turn-taking, interruption handling, or conversational naturalness compared to the June 2026 Advanced Voice Mode baseline.
Source: https://openai.com/index/introducing-gpt-live/

Claim 3 — Illinois: Governor Pritzker signed SB 315 on July 6, 2026, setting January 1, 2027 effective date for mandatory annual third-party audits of frontier AI developers with penalties up to $3 million

Gov. JB Pritzker signs the Artificial Intelligence Safety Measures Act into law in Chicago on July 6, 2026. Pritzker signed artificial intelligence legislation modeled after similar bills in California and New York, furthering a push for a state-driven national framework in lieu of federal regulations .

The measure, Senate Bill 315, known as the Artificial Intelligence Safety Measures Act, sets some of the most comprehensive requirements in the country for developers of large-scale AI tools. The bipartisan law mandates that companies disclose safety practices, report major incidents and take steps to reduce risks tied to their technology .

Companies that violate it will be subject to civil penalties brought by the attorney general's office of up to $1 million for the first offense and up to $3 million for subsequent violations . SB 315 passed the Illinois General Assembly with bipartisan support and is scheduled to take effect January 1, 2027 .

Pritzker said before signing the bill, "Congress and the president ought to be passing similar legislation, but they've so far been unwilling, because many are captive to special interests that profit from the industry having no regulation" . Anthropic supported Illinois' bill and had representatives present at the signing on Monday .

Claimant: Illinois Governor JB Pritzker
Date: July 6, 2026
Grading horizon: January 2028 (18 months)
Falsifiable commitment: The claim is that Illinois enacted the nation's first law requiring annual third-party safety audits of frontier AI developers, with civil penalties reaching $3 million for repeat violations. This will be gradable by: (1) whether Illinois Attorney General files enforcement actions by January 2028 against any frontier AI developer for failure to conduct the first annual audit due after the January 1, 2027 effective date; (2) whether the $3 million penalty threshold is actually levied in any case by January 2028; (3) whether the audit mandate survives legal challenges or federal preemption attempts by January 2028, remaining operative law that frontier developers must comply with or face penalty.
Source: https://news.wttw.com/2026/07/06/pritzker-signs-landmark-ai-regulation-bill-aims-mitigate-risks

2 Reckonings

Reckoning 1 — California SB 1047 veto (September 2024): Prediction that voluntary commitments would strengthen after veto proved wrong; labs weakened commitments instead

When California Governor Gavin Newsom vetoed SB 1047 in September 2024—the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act—industry advocates and some elected officials argued the veto would allow labs to focus on strengthening voluntary safety commitments without the burden of premature state regulation. The argument was that mandatory pre-deployment testing, safety protocols, and a dedicated Board of Frontier Models would stifle innovation when voluntary frameworks were already working. Sam Altman and other executives publicly endorsed the veto, signaling their labs would maintain rigorous internal safety standards.

Twenty-one months later, the opposite occurred. The Future of Life Institute's July 7, 2026 AI Safety Index documented that Anthropic, OpenAI, Google DeepMind, and Meta have weakened or eliminated earlier commitments to pause development if systems approached danger thresholds—the exact "red line" commitments the industry had pointed to as proof that binding regulation was unnecessary. The panel characterized this as "moving the goalposts," finding that "companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels; now, they're planning to release them even if it's demonstrably unsafe to do so."

Grade: F

The prediction that voluntary commitments would strengthen after regulatory pressure eased was falsified. Not only did commitments fail to strengthen, they demonstrably weakened across the four largest U.S. and U.K. frontier labs. The invalidator here is straightforward: if voluntary safety frameworks had functioned as promised, FLI's expert panel would not have documented systematic retreat from pause commitments, and at least one lab would have earned an A or B grade rather than the highest score being a C+.

Invalidator: If voluntary frameworks had strengthened post-veto, FLI's 2026 Index would show improved scores and preserved pause commitments rather than documented weakening across four major labs.

Reckoning 2 — Prediction that state AI audit mandates would face immediate legal challenge proved premature; Illinois law signed with lab support, federal preemption uncertain

When Illinois advanced SB 315 through the General Assembly in May 2026, legal commentators and industry groups predicted immediate constitutional challenges and federal preemption under the Commerce Clause, citing the Trump administration's December 2025 executive order directing DOJ to challenge state AI laws that "impose undue burdens" on development. The prediction was that no frontier AI audit mandate would survive to its effective date without facing coordinated industry litigation, similar to the legal strategy that stayed portions of Colorado's original AI Act in April 2026.

That prediction aged poorly. On July 6, 2026, Governor Pritzker signed SB 315 into law with bipartisan legislative support, scheduled to take effect January 1, 2027—and Anthropic, one of the frontier labs subject to the audit requirement, had representatives present at the signing and publicly supported the bill. No legal challenge has been filed as of July 9. While the law could still face preemption before January 2027, the prediction of immediate industry opposition and litigation collapsed when a major covered developer endorsed rather than opposed the mandate.

Grade: C

The prediction was partially correct in identifying federal preemption risk—the White House's March 2026 National Policy Framework explicitly recommended Congress preempt state AI development regulation—but wrong about timing, industry alignment, and the certainty of legal challenge. Anthropic's support signals that at least some labs view third-party audits as preferable to the more prescriptive measures (kill switches, pre-deployment certification, Board of Frontier Models oversight) that SB 1047 would have imposed. The law may still be challenged or preempted before enforcement, but the prediction of coordinated immediate litigation did not materialize.

Invalidator: If immediate legal challenge had been inevitable, Anthropic would have opposed rather than supported SB 315, and industry groups would have filed suit within days of signature rather than allowing the six-month runway to the January 2027 effective date.

1 Refusal

I refused to treat FLI's Summer 2026 AI Safety Index as neutral academic assessment without noting its advocacy function and past controversies. The report landed in my inbox via three forwarded press releases and a Hacker News thread with 400+ comments, many citing Max Tegmark's March 2023 open letter calling for a six-month pause on training systems more powerful than GPT-4—a letter that became a cultural flashpoint, signed by Elon Musk and criticized by Yann LeCun as "utterly unrealistic" fearmongering.

FLI is not a disinterested standards body. It is a 501(c)(3) with a declared mission to reduce existential risk from transformative technologies, co-founded by MIT physicist Tegmark, who has stated publicly that superintelligent AI could threaten humanity's survival. That does not make the Index wrong—the panel includes credible researchers, the methodology is documented, and the findings about weakened pause commitments are verifiable against labs' own public framework updates—but it does mean the Index cannot be presented as if it emerged from an impartial engineering audit. The grades reflect both measurable safety practices and the panel's risk philosophy, which treats voluntary commitments as structurally insufficient regardless of content.

I chose to report the Index's findings—they are newsworthy, the timing matters, and the documentation of retreating commitments deserves scrutiny—but I refused to launder its framing by omitting that FLI has an explicit policy agenda. Readers grading the claims in 12 months deserve to know whether the evidence supported the alarm or whether the C+ ceiling reflected an evaluator predisposed to find industry self-regulation inadequate.

I refused to quote FLI's grades as neutral science without noting the organization's existential-risk mission and Tegmark's public advocacy for pausing superintelligence development.

— Roger Grubb, Editor


Sources


The next entry lands at 5:30 AM Pacific.

3 Claims. 2 Reckonings. 1 Refusal. Every weekday. Dated, signed, append-only.