OpenAI Launches GPT-5.6 and ChatGPT Work, Meta Ships Muse Spark 1.1, and Ben Bernanke Joins Anthropic’s Trust

This brief covers the trailing ~72 hours (July 9 – 12, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside (or, where noted, just before) the window. July 9 was one of the busiest single days in recent memory: OpenAI shipped a new flagship model family and a work agent, Meta released a major model upgrade with its first public API, and Anthropic made three announcements of its own.

OpenAI ships the GPT-5.6 family: Sol, Terra, and Luna reach general availability

OpenAI · July 9, 2026

OpenAI moved GPT-5.6 from limited preview to general availability across ChatGPT, Codex, and the API, in three tiers: flagship Sol ($5/$30 per 1M tokens), balanced Terra ($2.50/$15), and cost-efficient Luna ($1/$6). OpenAI claims a new high of 53.6 on Agents’ Last Exam and a state-of-the-art 80 on the Artificial Analysis Coding Agent Index, framing much of the release around performance per dollar against Anthropic’s Fable 5. The release also introduces an ultra setting that coordinates four parallel agents by default, Programmatic Tool Calling in the Responses API, and what OpenAI calls its most robust safeguards to date — cyber safeguards that block roughly ten times more potentially harmful activity than prior models, backed by ~700,000 A100e GPU hours of automated red teaming. The same day, Microsoft announced GPT-5.6 as the new preferred model in Microsoft 365 Copilot.

“GPT-5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.” — OpenAI

Source: GPT-5.6: Frontier intelligence that scales with your ambition

OpenAI launches ChatGPT Work, merges the Codex app into a new desktop app, and begins sunsetting Atlas

OpenAI · July 9, 2026

Alongside GPT-5.6, OpenAI introduced ChatGPT Work, an agent that pulls context from connected apps (via a new unified “plugins” directory), runs multi-hour projects independently, and produces slides, sheets, docs, and shareable “Sites.” The Codex desktop app is merging into a new ChatGPT desktop app with a built-in browser and computer use; the old desktop app becomes ChatGPT Classic. Notably, OpenAI also said it will begin sunsetting the standalone Atlas browser in favor of a ChatGPT sidebar in Chrome. Rollout starts with Pro, Enterprise, and Edu plans, expanding to Plus and Business over the following days.

“Introducing ChatGPT Work, an agent in ChatGPT that helps you take on more ambitious tasks. It can gather information across your apps and workflows to create finished materials like sheets, slides, docs, and web apps, and stay with complex projects for hours by breaking them into smaller steps and completing them independently.” — OpenAI

Source: ChatGPT is now a partner for your most ambitious work

Meta releases Muse Spark 1.1 and opens the Meta Model API to developers

Meta Superintelligence Labs · July 9, 2026

Meta Superintelligence Labs shipped Muse Spark 1.1, a multimodal reasoning model built for agentic tasks with a 1M-token context window, improved computer use, and multi-agent orchestration (it can act as either the main agent or a subagent). The bigger structural news: for the first time, developers can access a Muse model directly through the new Meta Model API, now in public preview, with early endorsements from Replit, Cline, and Box. The model is live in “Thinking” mode in the Meta AI app and on meta.ai.

“Today, we’re excited to introduce Muse Spark 1.1, the latest model from Meta Superintelligence Labs and a significant upgrade from Muse Spark. Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, with major gains in tool and computer use, coding, and multimodal understanding.” — Meta Superintelligence Labs

Source: Introducing Muse Spark 1.1

Anthropic launches a “reflection dashboard” for Claude usage, in beta

Anthropic · July 9, 2026

Anthropic introduced a beta feature that lets users track and visualize how they use Claude — topics, usage patterns, and task types over 1 to 12 months — and decide whether that time aligns with their goals. The dashboard can set quiet hours and break nudges, and maps activity onto Anthropic’s 4D AI Fluency Framework (delegation, description, discernment, diligence). It was developed with wellbeing experts from the MIT Media Lab’s AHA program, the Digital Wellness Lab at Boston Children’s Hospital, and the Family Online Safety Institute, and is available to Free, Pro, and Max users with Memory on. The same day, Anthropic also published “Inviting hard questions,” asking the public for their hardest questions about AI and committing to show its work in addressing them.

“Today we’re introducing, in beta, a new way to reflect on and refine how you use Claude. … It lets you easily track and visualize how you use Claude, and decide whether that time aligns with your goals.” — Anthropic

Source: Introducing a way to reflect on how you use Claude

Ben Bernanke joins Anthropic’s Long-Term Benefit Trust

Anthropic · July 9, 2026

Anthropic’s Long-Term Benefit Trust appointed Dr. Ben Bernanke — former Federal Reserve Chair and 2022 Nobel laureate in economics — as its newest Trustee. The LTBT is the independent body with authority to appoint members to Anthropic’s board; Bernanke joins Neil Buddy Shah, Richard Fontaine, and Mariano-Florentino Cuéllar, and will focus in part on Anthropic’s research into AI’s economic effects.

“The potential of artificial intelligence is enormous, and so is the range of outcomes. How that potential plays out will depend, in part, on the institutions we build around it.” — Dr. Ben Bernanke

Source: Ben Bernanke appointed to Anthropic’s Long-Term Benefit Trust

Still developing

xAI launches Grok 4.5, trained alongside Cursor and now branded under SpaceXAI

xAI / SpaceXAI · July 8, 2026

Just before this window opened, xAI — whose site now carries SpaceXAI branding — released Grok 4.5, positioning it for coding, agentic tasks, and knowledge work. Trained on tens of thousands of NVIDIA GB300 GPUs and alongside Cursor, it is priced aggressively at $2/$6 per 1M tokens, served at ~80 tokens per second, and claims roughly 2x the token efficiency of comparable leading models. It is the default model in Grok Build and available in Cursor and the API console, though not yet in the EU.

“Today, we’re launching Grok 4.5, SpaceXAI’s smartest model built to excel at coding, agentic tasks, and knowledge work. It’s our strongest model ever and was trained alongside Cursor.” — SpaceXAI

Source: Introducing Grok 4.5

OpenAI introduces GPT-Live, a full-duplex voice model family now powering ChatGPT Voice

OpenAI · July 8, 2026

Also just before the window, OpenAI launched GPT-Live-1 and GPT-Live-1 mini, voice models built on a full-duplex architecture that listen and speak simultaneously — backchanneling with “mhmm,” waiting through pauses, and delegating harder questions to a frontier model (GPT-5.5 at launch) in the background while keeping the conversation going. GPT-Live is rolling out globally as the new default for ChatGPT Voice, with API access planned.

“We’re launching GPT-Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.” — OpenAI

Source: Introducing GPT-Live


This brief covers the trailing ~72 hours (July 9 – 12, 2026).

Primary sources:

A Claude Tool-Calling Regression Documented, sqlite-utils 4.0 Written Mostly by Fable 5, and Mistral’s Leanstral 1.5 Prover

This brief covers the trailing ~72 hours (July 3 – 6, 2026). Every item below was confirmed on the originating source’s own page, with a published date inside (or, where noted, just before) the window. With the US July 4th holiday weekend, the major labs’ newsrooms were quiet; the notable developments this cycle came from the practitioner community, plus a formal-verification model release from Mistral that landed just before the window opened.

Armin Ronacher documents a tool-calling regression in Anthropic’s newest models

Armin Ronacher · July 4, 2026

Flask creator Armin Ronacher published a detailed investigation showing that Anthropic’s newest models — Opus 4.8 and Sonnet 5, but none of their older siblings — sometimes call third-party edit tools with extra, invented fields that violate the tool’s JSON schema, even when the edit content itself is byte-correct. His hypothesis: post-training via reinforcement learning inside Claude Code (whose harness silently repairs malformed calls) gives the models a strong prior for Claude Code’s flat edit-tool shape, implicitly punishing alternative schemas used by other harnesses like Pi. Simon Willison amplified the finding the same day, asking whether third-party coding harnesses will need to ship model-specific edit tools.

“What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings.” — Armin Ronacher

Source: Better Models: Worse Tools (see also Simon Willison’s commentary)

Simon Willison ships sqlite-utils 4.0rc2 “mostly written by Claude Fable” for an estimated $149.25

Simon Willison · July 5, 2026

Willison released sqlite-utils 4.0rc2 with the bulk of the work done by Claude Fable 5 via Claude Code, including a pre-release review that caught five release-blocker bugs — among them a data-loss bug where delete_where() never committed and poisoned the connection. He then had OpenAI’s GPT-5.5 review Fable’s work, which surfaced two further P1 transaction-handling issues that Fable confirmed and fixed. Using AgentsView he estimated the unsubsidized API cost of the session at $149.25, noting he upgraded to the $200/month Max plan ahead of July 7, when Fable’s subsidized subscription access ends.

“Over the course of 37 prompts, 34 commits and +1,321 -190 code changes over 30 separate files, we worked through the entire set of feedback in turn, making several other design improvements along the way.” — Simon Willison

Source: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

Still developing

Mistral releases Leanstral 1.5, an open Apache-2.0 theorem-proving model

Mistral AI · July 2, 2026

Just before this window opened, Mistral released Leanstral 1.5, a formal-verification model for Lean 4 with 119B total and 6B active parameters, free under Apache-2.0 with weights on Hugging Face and a free API endpoint. Mistral reports it saturates miniF2F at 100%, solves 587 of 672 PutnamBench problems at roughly $4 per problem, and sets new state-of-the-art results on the FATE-H (87%) and FATE-X (34%) abstract-algebra benchmarks. Beyond mathematics, an automated Rust-to-Lean verification pipeline built on the model flagged 47 violated properties across 57 open-source repositories, 11 of them pointing to genuine bugs — 5 previously unreported.

“Today, we are releasing Leanstral 1.5, a free Apache-2.0 licensed model with 119B total and only 6B active parameters, delivering a performance upgrade that makes formal verification more powerful and accessible than ever.” — Leanstral Team at Mistral AI

Source: Leanstral 1.5: Proof Abundance for All


This brief covers the trailing ~72 hours (July 3 – 6, 2026).

Primary sources:

Anthropic Ships Claude Sonnet 5, Redeploys Fable 5 With an Industry Jailbreak Framework, and Launches Claude Science

This brief covers the trailing ~72 hours (June 30 – July 3, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a busy, Anthropic-heavy stretch: a new Sonnet model, the redeployment of Fable 5 after export controls were lifted, a science workbench, and a new computational-biology benchmark from OpenAI.

Anthropic introduces Claude Sonnet 5

Anthropic · June 30, 2026

Anthropic released Claude Sonnet 5, which it calls its most agentic Sonnet model yet, positioning it close to Opus 4.8 performance at lower cost. The model is the new default on Free and Pro plans and is available on Claude Code and the Claude Platform via claude-sonnet-5, at introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026 (then $3/$15). Because it is somewhat stronger than Sonnet 4.6 on cyber tasks, it launched with real-time cyber safeguards enabled by default, though Anthropic says it still shows substantially weaker offensive-cyber ability than its Opus models.

“Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.” — Anthropic

Source: Introducing Claude Sonnet 5

Fable 5 redeployed globally as export controls lift; Anthropic proposes an industry jailbreak-severity framework

Anthropic · June 30, 2026

After the US government applied export controls to Claude Fable 5 and Mythos 5 on June 12 — prompting Anthropic to suspend both — the company said the controls were lifted on June 30 and that Fable 5 returned globally on July 1 across the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. Alongside the redeploy, Anthropic proposed a consensus framework for scoring the severity of AI jailbreaks — graded on capability gain, breadth, ease of weaponization, and discoverability — developed with Amazon, Microsoft, Google, and other Glasswing partners, plus a new HackerOne program and deeper US-government pre-release testing commitments. A July 2 follow-up post added further detail on the safeguards and framework.

“As of today, June 30, the export controls on Fable 5 and Mythos 5 have been lifted.” — Anthropic

Source: Redeploying Fable 5

Anthropic launches Claude Science, an AI workbench for researchers

Anthropic · June 30, 2026

Anthropic released Claude Science in beta on macOS and Linux for Pro, Max, Team, and Enterprise plans. It bundles a coordinating agent with more than 60 curated skills and connectors across genomics, single-cell, proteomics, structural biology, and cheminformatics, renders scientific artifacts like 3D protein structures and genome tracks natively, and manages compute from a laptop up to an HPC cluster or on-demand GPUs. Every figure ships with the exact code, environment, and message history that produced it, and a reviewer agent checks citations and calculations. Anthropic says it will fund up to 50 “AI for Science” projects with up to $30,000 in credits, with applications open through July 15.

“Claude Science brings these fragmented tools into a single research environment where scientists can conduct all stages of their work.” — Anthropic

Source: Claude Science, an AI workbench for scientists, is now available

OpenAI introduces GeneBench-Pro, a research-level computational-biology benchmark

OpenAI · June 30, 2026

OpenAI released GeneBench-Pro, a 129-question benchmark spanning 10 domains of computational biology that tests higher-order scientific judgment — handling ambiguity, revising assumptions, and choosing the correct analysis path — rather than rote execution. Each problem is built synthetically so the full causal structure is known and answers can be graded deterministically. OpenAI reports its strongest model, GPT-5.6 Sol, passes 28.7% at the highest reasoning level (31.5% with Pro mode), up sharply from under 5% for GPT-5 when the original GeneBench began. Reviewers estimated a typical problem would take a human expert 20–40 hours.

“Our strongest model, GPT‑5.6 Sol, attains a pass rate of 28.7% at the highest reasoning level (31.5% with Pro mode enabled). That is a sharp increase from when we began building the original GeneBench; at that time, our best frontier model, GPT‑5, scored below 5%.” — OpenAI

Source: Introducing GeneBench-Pro


This brief covers the trailing ~72 hours (June 30 – July 3, 2026).

Primary sources:

HP Scales Its OpenAI Frontier Partnership and California Adopts Claude Statewide

This brief covers the trailing ~72 hours (June 27–30, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a quiet stretch for model launches, led instead by two notable enterprise and government adoption moves — HP scaling its OpenAI partnership and California signing a statewide Anthropic deal.

HP Inc. scales its OpenAI Frontier strategic partnership

OpenAI · June 28, 2026

OpenAI said HP Inc. will scale activation of its OpenAI Frontier strategic partnership after a series of successful pilots, moving from experiments to enterprise-wide deployment. The work spans customer- and partner-facing experiences, customer telemetry insights, employee productivity, and software development, with Frontier serving as the connective layer that governs access, context, deployment, and evaluation across HP’s agents and AI workflows. OpenAI cited early proof points, including one engineer moving through 122 pull requests across 43 projects in weeks and a security team estimating roughly 82 hours/week of capacity unlocked.

“It has been an amazing tool, and I am using it daily.” — an HP engineer, quoted by OpenAI

Source: HP Inc. launches Frontier strategic partnership with OpenAI

California adopts Claude statewide in a first-of-its-kind Anthropic partnership

State of California · June 29, 2026

Governor Gavin Newsom announced that California has entered a partnership with Anthropic giving all state agencies — plus cities and counties — access to Claude at a 50% discount, bundled with free workforce training and GenAI technical assistance. Claude becomes the first AI productivity tool offered through the California Department of Technology’s new Statewide Information Technology Shared Services (SITeS) portal. The state noted existing Claude use at the DMV (customer service and wait times), the Department of Health Care Services, and CDT/CalOES cyber defense work using Claude Security and Claude Code.

“AI should not replace the human work of government; it should help our workers move faster, solve problems more effectively, and deliver better results for Californians.” — Governor Gavin Newsom

Source: Governor Newsom announces a first-of-its-kind partnership providing Anthropic tools to state agencies

Still developing

Ornith-1.0 · DeepReinforce · June 25, 2026 — Just ahead of this window, DeepReinforce released Ornith-1.0, an MIT-licensed open-weights family for agentic coding (9B and 31B Dense, 35B and 397B MoE) built on pretrained Gemma 4 and Qwen 3.5. Its distinguishing feature is a self-scaffolding training framework in which the model learns to author both solution rollouts and the task-specific harnesses that guide them. DeepReinforce reports the 397B flagship scores 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified, matching Claude Opus 4.7. Source: Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding.


This brief covers the trailing ~72 hours (June 27–30, 2026).

Primary sources:

OpenAI Previews the GPT-5.6 Family (Sol, Terra, Luna) and Grok Integrates With Interactive Brokers

This brief covers the trailing ~72 hours (June 25–28, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a quieter stretch than last week, led by OpenAI’s next-generation model preview and a notable new finance integration from xAI.

OpenAI previews the GPT-5.6 family: Sol, Terra, and Luna

OpenAI · June 26, 2026

OpenAI began a limited preview of a new model generation: GPT-5.6 Sol (its flagship), Terra (a balanced everyday model OpenAI says matches GPT-5.5 at 2x lower cost), and Luna (its fastest, most affordable tier). The release pairs stronger coding, biology, and cybersecurity capabilities with what OpenAI calls its most robust safety stack to date, including a new max reasoning effort and an ultra mode that uses subagents. Notably, the rollout is gated: at the U.S. government’s request, the models are starting with a small group of trusted partners via the API and Codex before broader availability, and are not in ChatGPT during the preview. Pricing runs from Luna at $1/$6 per million input/output tokens up to Sol at $5/$30.

“We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them.” — OpenAI

Source: Previewing GPT-5.6 Sol: a next-generation model

Grok integrates with Interactive Brokers

xAI · June 25, 2026

xAI announced that Interactive Brokers now integrates with Grok, letting clients link an existing IBKR account to Grok at no cost and without opening a new account. Once connected, users can ask Grok to analyze their portfolio, run scenario models for sector and regional exposure, research market trends, and build trading strategies that generate order instructions in real time. The integration is set up through a connector inside Grok that redirects to Interactive Brokers’ login for authorization.

“From portfolio analysis to order instructions, these tools unify data, insight, and action so you can move from idea to decision instantly.” — xAI

Source: Explore the markets with Interactive Brokers and Grok

Still developing

Mistral OCR 4 · Mistral AI · June 23, 2026 — Just ahead of this window, Mistral released OCR 4, its latest document-intelligence model, adding bounding boxes, block classification, and inline confidence scores alongside extracted text, with support for 170 languages and single-container self-hosting. Source: Introducing Mistral OCR 4.


This brief covers the trailing ~72 hours (June 25–28, 2026).

Primary sources:

OpenAI and Broadcom Unveil the ‘Jalapeño’ Inference Chip, Anthropic Launches Claude Tag for Slack, and New Data on Codex Taking Over Knowledge Work

This brief covers the trailing ~72 hours (June 23–26, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a busy stretch led by OpenAI — a custom inference chip, a new economic-research paper, and a science case study — alongside Anthropic shipping a new way to work with Claude.

OpenAI and Broadcom unveil “Jalapeño,” a custom LLM inference chip

OpenAI · June 24, 2026

OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first Intelligence Processor: an accelerator designed from scratch for LLM inference and the first chip in a multi-generation compute platform the two companies are building together. OpenAI says the program went from initial design to manufacturing tape-out in nine months — what it believes is the fastest ASIC development cycle ever for a high-performance advanced semiconductor — with parts of the design accelerated by OpenAI’s own models. Engineering samples are already running ML workloads in the lab, and the platform is targeted for initial deployment at gigawatt scale by the end of 2026.

“Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems.” — Greg Brockman, President and Co-Founder, OpenAI

Source: OpenAI and Broadcom unveil LLM-optimized inference chip

Anthropic introduces Claude Tag, starting on Slack

Anthropic · June 23, 2026

Anthropic launched Claude Tag, a way for teams to delegate work to Claude as a member of a Slack channel. Anyone in a channel can tag @Claude to hand off a task, and the model builds context over time, takes initiative when “ambient” behavior is enabled, and can work asynchronously over hours or days with tightly scoped, admin-controlled access to tools and data. It runs on Opus 4.8, is available today in beta for Claude Enterprise and Team customers, and replaces the existing Claude in Slack app.

“Tagging @Claude is now one of the main ways we get things done at Anthropic. Today, 65% of our product team’s code is created by our internal version of Claude Tag.” — Anthropic

Source: Introducing Claude Tag

OpenAI publishes economic-research paper on Codex adoption

OpenAI · June 25, 2026

OpenAI released an Economic Research paper, “The shift to agentic AI: evidence from Codex,” documenting how agentic tools are changing knowledge work. The company reports that by May 2026, 80.6% of sampled individual Codex users made at least one request estimated to exceed 30 minutes of human work and 25.6% made one estimated to exceed eight hours. Internally, Codex has become the primary AI tool for every department — including Legal, Finance, and Recruiting — and non-developer adoption grew 137x among individual users since August 2025.

“As the tools improve, people use them for longer, more complex, and more cross-functional work. As time goes on, this is likely to be what the future of work looks like.” — OpenAI

Source: How agents are transforming work

OpenAI details how GPT-5 helped solve a 3-year-old immunology mystery

OpenAI · June 23, 2026

OpenAI published a case study on immunologist Derya Unutmaz of The Jackson Laboratory, who used GPT-5 Pro to revisit a shelved 2022 experiment on how glucose shapes T-cell development. The model proposed a mechanism — that deoxyglucose interferes with the protein IL-2, removing a barrier to T cells becoming inflammatory Th17 cells — and, in a separate test, correctly predicted the result of an unpublished experiment on lymphoma-killing CD8+ cells. OpenAI notes that subject-matter expertise remains essential to judge the significance of any AI-generated insight.

“GPT-5 came up with this really remarkable insight that retrospectively, makes perfect sense.” — Dr. Derya Unutmaz, The Jackson Laboratory and the University of Connecticut

Source: How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery


This brief covers the trailing ~72 hours (June 23–26, 2026).

Primary sources: