Moonshot’s Kimi K3 Becomes the Largest Open Model, Thinking Machines Debuts Inkling, and OpenAI Details GPT-Red Adversarial Safety Training

This brief covers the trailing ~72 hours (July 15–18, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a heavyweight stretch for open-weight models: Moonshot shipped the largest open model ever announced, and Mira Murati’s Thinking Machines Lab released its first model, while OpenAI published three posts spanning safety research, teen policy, and AI economics, and Google DeepMind and Isomorphic Labs laid out a joint biosecurity strategy.

Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model it calls the first open 3T-class model

Moonshot AI · July 16, 2026

Moonshot introduced Kimi K3, a 2.8T-parameter Mixture-of-Experts model (16 of 896 experts active) built on its Kimi Delta Attention and Attention Residuals architectures, with native vision and a 1-million-token context window. Moonshot says K3 trails only Claude Fable 5 and GPT-5.6 Sol overall while consistently outperforming other tested models, and it is priced at $3/$15 per million tokens — the most expensive Chinese-lab model to date. K3 is live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with full open weights promised by July 27, 2026.

“Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world’s first open 3T-class model.” — Moonshot AI

Source: Kimi K3: Open Frontier Intelligence

Thinking Machines Lab releases Inkling, its first open-weights model

Thinking Machines Lab · July 15, 2026

Mira Murati’s Thinking Machines Lab released Inkling, a 975B-parameter Mixture-of-Experts model (41B active) trained from scratch on 45 trillion tokens of text, images, audio, and video, with a 1M-token context window and controllable thinking effort. The lab positions Inkling not as a benchmark leader but as a broad, balanced base for customization via its Tinker fine-tuning platform, and it trained the model specifically for calibration, instruction following, and resistance to censorship. Full weights are on Hugging Face, and a lighter Inkling-Small (12B active) is previewed.

“Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.” — Thinking Machines Lab

Source: Inkling: Our open-weights model

OpenAI reveals GPT-Red, an internal red-teaming model trained via self-play to harden GPT-5.6

OpenAI · July 15, 2026

OpenAI detailed GPT-Red, an internal-only automated red-teaming model trained with self-play reinforcement learning at the compute scale of some of its largest post-training runs. GPT-Red found successful prompt-injection attacks in 84% of held-out scenarios versus 13% for human red-teamers, and even broke a live Andon Labs vending-machine agent in OpenAI’s office. Used adversarially during GPT-5.6’s training, it helped drive the model’s failure rate on GPT-Red’s direct prompt injections down to 0.05%, with a pre-print promised the following week.

“We believe automated red-teaming unlocks a crucial form of self-improvement for safety: using today’s models to directly help make future models safer.” — OpenAI

Source: GPT-Red: Unlocking Self-Improvement for Robustness

Google DeepMind and Isomorphic Labs publish a joint approach to bioresilience

Google DeepMind · July 16, 2026

Google DeepMind and Isomorphic Labs published a joint bioresilience framework organized around prevention, detection, and response: adapting SynthID watermarking to biological sequences for DNA-synthesis screening, using AlphaEvolve to cut the cost of metagenomic pathogen surveillance, and granting trusted researchers access to frontier systems to accelerate vaccine and countermeasure design. Isomorphic Labs has stood up a dedicated unit to rapidly deploy its Drug Design Engine during novel outbreaks, and the pair report more than 15 partnerships with governments and biosecurity organizations over the past 12 months.

“Our work is twofold – to prevent threat actors from misusing our models, and to ensure that governments, scientists, biosecurity experts and our teams can harness these technologies to build a more resilient world.” — Google DeepMind and Isomorphic Labs

Source: Our approach to bioresilience

OpenAI argues teens deserve access to safe AI, expands Study Mode parental controls

OpenAI · July 16, 2026

OpenAI published its case for teen access to AI paired with age-appropriate protections, citing that nearly 9 in 10 teens on ChatGPT use it weekly for learning or productivity. New measures include letting parents enable Study Mode by default from Parental Controls, education-focused starter prompts, more frequent break reminders for teens, and expanded parent notifications — now covering account deactivations for violent-threat policy violations, developed with violence-prevention firm Moonshot (unrelated to Moonshot AI). OpenAI also announced it has joined the Family Online Safety Institute.

“Keeping teens from using it until adulthood would be like asking a previous generation to avoid the internet or search engines until they turned 18, leaving them less prepared to use one of the defining technologies of their time.” — OpenAI

Source: Why teens deserve access to safe AI

OpenAI CFO Sarah Friar proposes a four-part “scorecard for the AI age”

OpenAI · July 17, 2026

In a follow-up to last week’s enterprise spend playbook, OpenAI CFO Sarah Friar proposed measuring AI ROI as “Useful Intelligence per Dollar” across four questions: how much useful work gets done, what a successful task costs, how dependable the results are, and whether each AI dollar buys more work at scale. The post frames GPT-5.6’s three tiers (Sol, Terra, Luna) as levers in that equation and claims GPT-5.6 Sol set a new state of the art on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model.

“The ultimate scorecard for the age of AI could be looked at as ‘Useful Intelligence per Dollar.’” — Sarah Friar, CFO, OpenAI

Source: A scorecard for the AI age


This brief covers the trailing ~72 hours (July 15–18, 2026).

Primary sources:

Anthropic Launches Claude for Teachers and a $10M Canadian Research Commitment, While OpenAI Publishes an Agentic-Era Spend Playbook

This brief covers the trailing ~72 hours (July 12–15, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. After one of the busiest weeks in recent memory (GPT-5.6, ChatGPT Work, Muse Spark 1.1, and Grok 4.5 all landed July 8–9), this window was notably quiet: the verified news is concentrated on a single day, July 14, led by two Anthropic announcements.

Anthropic launches Claude for Teachers, free for verified US K-12 educators

Anthropic · July 14, 2026

Anthropic introduced Claude for Teachers, giving verified US K-12 educators free access to premium Claude capabilities, a library of teaching skills grounded in learning science, and connections to evidence-based curricula mapped to academic standards in all 50 states via Learning Commons. The product connects to an ecosystem of K-12 tools (ASSISTments, Brisk Teaching, Canva Education, MagicSchool, and others), includes Claude Code and Cowork for tasks like analyzing class data and scheduling recurring work, and ships with K-12-specific privacy terms written to comply with FERPA — teacher data is not used for model training. Anthropic will pilot an evaluation in the Detroit Public Schools Community District, is open-sourcing the teaching skills, and is working with the American Federation of Teachers on privacy standards. Educators who sign up by June 30, 2027 get a full year of free access.

“We’re introducing Claude for Teachers, providing verified K-12 educators in the US free access to premium Claude capabilities, a library of teaching skills, and a direct connection to evidence-based curricula, mapped to academic standards in all 50 states.” — Anthropic

Source: Introducing Claude for Teachers

Anthropic commits $10 million CAD to Canadian AI research

Anthropic · July 14, 2026

Anthropic announced a $10 million CAD commitment to Canadian research institutions, with partnerships spanning the country’s three regional AI institutes — Amii (Edmonton), Mila (Montréal), and the Vector Institute (Toronto) — plus CHEO, CAMH, Université Laval, the University of Toronto, and the University of Saskatchewan. The funding targets beneficial and responsible AI applications from reinforcement learning and AI safety to children’s health and low-resource languages like Quebec French and Indigenous languages. Anthropic also published its first Canadian country brief from the Anthropic Economic Index, finding Canada ranks eighth worldwide in Claude.ai use and second in per-capita adoption among the top ten countries.

“Some of the foundations of modern AI came out of Toronto, Montréal, and Edmonton— and so, strikingly, did many of the researchers most committed to making it safe. I was formed by that culture, and I’m proud Anthropic can support the next chapter.” — Chris Olah, Co-Founder, Anthropic

Source: Anthropic commits $10 million to Canadian AI research

OpenAI publishes an enterprise playbook for managing AI spend in the agentic era

OpenAI · July 14, 2026

Following last week’s GPT-5.6 and ChatGPT Work launches, OpenAI published guidance for enterprise leaders on managing AI investments as agents take on longer-running work. The five-step framework centers on measuring “useful work per dollar” rather than token price, tracking cost per accepted outcome, and governing agentic workflows before they scale. The post also highlights updated usage analytics and spend controls in the ChatGPT Admin Console — adoption, credit usage, and spend broken down by user, product, and model — and notes that token prices fell 97% from GPT-4 to GPT-5.4, with GPT-5.6 completing coding-agent tasks with 54% fewer output tokens.

“But token price alone does not show whether AI is creating value. Leaders should look at useful work per dollar: tasks completed, time saved, decisions improved, and workflows ready to scale.” — OpenAI

Source: How to manage AI investments in the agentic era


This brief covers the trailing ~72 hours (July 12–15, 2026).

Primary sources:

OpenAI Launches GPT-5.6 and ChatGPT Work, Meta Ships Muse Spark 1.1, and Ben Bernanke Joins Anthropic’s Trust

This brief covers the trailing ~72 hours (July 9 – 12, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside (or, where noted, just before) the window. July 9 was one of the busiest single days in recent memory: OpenAI shipped a new flagship model family and a work agent, Meta released a major model upgrade with its first public API, and Anthropic made three announcements of its own.

OpenAI ships the GPT-5.6 family: Sol, Terra, and Luna reach general availability

OpenAI · July 9, 2026

OpenAI moved GPT-5.6 from limited preview to general availability across ChatGPT, Codex, and the API, in three tiers: flagship Sol ($5/$30 per 1M tokens), balanced Terra ($2.50/$15), and cost-efficient Luna ($1/$6). OpenAI claims a new high of 53.6 on Agents’ Last Exam and a state-of-the-art 80 on the Artificial Analysis Coding Agent Index, framing much of the release around performance per dollar against Anthropic’s Fable 5. The release also introduces an ultra setting that coordinates four parallel agents by default, Programmatic Tool Calling in the Responses API, and what OpenAI calls its most robust safeguards to date — cyber safeguards that block roughly ten times more potentially harmful activity than prior models, backed by ~700,000 A100e GPU hours of automated red teaming. The same day, Microsoft announced GPT-5.6 as the new preferred model in Microsoft 365 Copilot.

“GPT-5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.” — OpenAI

Source: GPT-5.6: Frontier intelligence that scales with your ambition

OpenAI launches ChatGPT Work, merges the Codex app into a new desktop app, and begins sunsetting Atlas

OpenAI · July 9, 2026

Alongside GPT-5.6, OpenAI introduced ChatGPT Work, an agent that pulls context from connected apps (via a new unified “plugins” directory), runs multi-hour projects independently, and produces slides, sheets, docs, and shareable “Sites.” The Codex desktop app is merging into a new ChatGPT desktop app with a built-in browser and computer use; the old desktop app becomes ChatGPT Classic. Notably, OpenAI also said it will begin sunsetting the standalone Atlas browser in favor of a ChatGPT sidebar in Chrome. Rollout starts with Pro, Enterprise, and Edu plans, expanding to Plus and Business over the following days.

“Introducing ChatGPT Work, an agent in ChatGPT that helps you take on more ambitious tasks. It can gather information across your apps and workflows to create finished materials like sheets, slides, docs, and web apps, and stay with complex projects for hours by breaking them into smaller steps and completing them independently.” — OpenAI

Source: ChatGPT is now a partner for your most ambitious work

Meta releases Muse Spark 1.1 and opens the Meta Model API to developers

Meta Superintelligence Labs · July 9, 2026

Meta Superintelligence Labs shipped Muse Spark 1.1, a multimodal reasoning model built for agentic tasks with a 1M-token context window, improved computer use, and multi-agent orchestration (it can act as either the main agent or a subagent). The bigger structural news: for the first time, developers can access a Muse model directly through the new Meta Model API, now in public preview, with early endorsements from Replit, Cline, and Box. The model is live in “Thinking” mode in the Meta AI app and on meta.ai.

“Today, we’re excited to introduce Muse Spark 1.1, the latest model from Meta Superintelligence Labs and a significant upgrade from Muse Spark. Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, with major gains in tool and computer use, coding, and multimodal understanding.” — Meta Superintelligence Labs

Source: Introducing Muse Spark 1.1

Anthropic launches a “reflection dashboard” for Claude usage, in beta

Anthropic · July 9, 2026

Anthropic introduced a beta feature that lets users track and visualize how they use Claude — topics, usage patterns, and task types over 1 to 12 months — and decide whether that time aligns with their goals. The dashboard can set quiet hours and break nudges, and maps activity onto Anthropic’s 4D AI Fluency Framework (delegation, description, discernment, diligence). It was developed with wellbeing experts from the MIT Media Lab’s AHA program, the Digital Wellness Lab at Boston Children’s Hospital, and the Family Online Safety Institute, and is available to Free, Pro, and Max users with Memory on. The same day, Anthropic also published “Inviting hard questions,” asking the public for their hardest questions about AI and committing to show its work in addressing them.

“Today we’re introducing, in beta, a new way to reflect on and refine how you use Claude. … It lets you easily track and visualize how you use Claude, and decide whether that time aligns with your goals.” — Anthropic

Source: Introducing a way to reflect on how you use Claude

Ben Bernanke joins Anthropic’s Long-Term Benefit Trust

Anthropic · July 9, 2026

Anthropic’s Long-Term Benefit Trust appointed Dr. Ben Bernanke — former Federal Reserve Chair and 2022 Nobel laureate in economics — as its newest Trustee. The LTBT is the independent body with authority to appoint members to Anthropic’s board; Bernanke joins Neil Buddy Shah, Richard Fontaine, and Mariano-Florentino Cuéllar, and will focus in part on Anthropic’s research into AI’s economic effects.

“The potential of artificial intelligence is enormous, and so is the range of outcomes. How that potential plays out will depend, in part, on the institutions we build around it.” — Dr. Ben Bernanke

Source: Ben Bernanke appointed to Anthropic’s Long-Term Benefit Trust

Still developing

xAI launches Grok 4.5, trained alongside Cursor and now branded under SpaceXAI

xAI / SpaceXAI · July 8, 2026

Just before this window opened, xAI — whose site now carries SpaceXAI branding — released Grok 4.5, positioning it for coding, agentic tasks, and knowledge work. Trained on tens of thousands of NVIDIA GB300 GPUs and alongside Cursor, it is priced aggressively at $2/$6 per 1M tokens, served at ~80 tokens per second, and claims roughly 2x the token efficiency of comparable leading models. It is the default model in Grok Build and available in Cursor and the API console, though not yet in the EU.

“Today, we’re launching Grok 4.5, SpaceXAI’s smartest model built to excel at coding, agentic tasks, and knowledge work. It’s our strongest model ever and was trained alongside Cursor.” — SpaceXAI

Source: Introducing Grok 4.5

OpenAI introduces GPT-Live, a full-duplex voice model family now powering ChatGPT Voice

OpenAI · July 8, 2026

Also just before the window, OpenAI launched GPT-Live-1 and GPT-Live-1 mini, voice models built on a full-duplex architecture that listen and speak simultaneously — backchanneling with “mhmm,” waiting through pauses, and delegating harder questions to a frontier model (GPT-5.5 at launch) in the background while keeping the conversation going. GPT-Live is rolling out globally as the new default for ChatGPT Voice, with API access planned.

“We’re launching GPT-Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.” — OpenAI

Source: Introducing GPT-Live


This brief covers the trailing ~72 hours (July 9 – 12, 2026).

Primary sources:

A Claude Tool-Calling Regression Documented, sqlite-utils 4.0 Written Mostly by Fable 5, and Mistral’s Leanstral 1.5 Prover

This brief covers the trailing ~72 hours (July 3 – 6, 2026). Every item below was confirmed on the originating source’s own page, with a published date inside (or, where noted, just before) the window. With the US July 4th holiday weekend, the major labs’ newsrooms were quiet; the notable developments this cycle came from the practitioner community, plus a formal-verification model release from Mistral that landed just before the window opened.

Armin Ronacher documents a tool-calling regression in Anthropic’s newest models

Armin Ronacher · July 4, 2026

Flask creator Armin Ronacher published a detailed investigation showing that Anthropic’s newest models — Opus 4.8 and Sonnet 5, but none of their older siblings — sometimes call third-party edit tools with extra, invented fields that violate the tool’s JSON schema, even when the edit content itself is byte-correct. His hypothesis: post-training via reinforcement learning inside Claude Code (whose harness silently repairs malformed calls) gives the models a strong prior for Claude Code’s flat edit-tool shape, implicitly punishing alternative schemas used by other harnesses like Pi. Simon Willison amplified the finding the same day, asking whether third-party coding harnesses will need to ship model-specific edit tools.

“What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings.” — Armin Ronacher

Source: Better Models: Worse Tools (see also Simon Willison’s commentary)

Simon Willison ships sqlite-utils 4.0rc2 “mostly written by Claude Fable” for an estimated $149.25

Simon Willison · July 5, 2026

Willison released sqlite-utils 4.0rc2 with the bulk of the work done by Claude Fable 5 via Claude Code, including a pre-release review that caught five release-blocker bugs — among them a data-loss bug where delete_where() never committed and poisoned the connection. He then had OpenAI’s GPT-5.5 review Fable’s work, which surfaced two further P1 transaction-handling issues that Fable confirmed and fixed. Using AgentsView he estimated the unsubsidized API cost of the session at $149.25, noting he upgraded to the $200/month Max plan ahead of July 7, when Fable’s subsidized subscription access ends.

“Over the course of 37 prompts, 34 commits and +1,321 -190 code changes over 30 separate files, we worked through the entire set of feedback in turn, making several other design improvements along the way.” — Simon Willison

Source: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

Still developing

Mistral releases Leanstral 1.5, an open Apache-2.0 theorem-proving model

Mistral AI · July 2, 2026

Just before this window opened, Mistral released Leanstral 1.5, a formal-verification model for Lean 4 with 119B total and 6B active parameters, free under Apache-2.0 with weights on Hugging Face and a free API endpoint. Mistral reports it saturates miniF2F at 100%, solves 587 of 672 PutnamBench problems at roughly $4 per problem, and sets new state-of-the-art results on the FATE-H (87%) and FATE-X (34%) abstract-algebra benchmarks. Beyond mathematics, an automated Rust-to-Lean verification pipeline built on the model flagged 47 violated properties across 57 open-source repositories, 11 of them pointing to genuine bugs — 5 previously unreported.

“Today, we are releasing Leanstral 1.5, a free Apache-2.0 licensed model with 119B total and only 6B active parameters, delivering a performance upgrade that makes formal verification more powerful and accessible than ever.” — Leanstral Team at Mistral AI

Source: Leanstral 1.5: Proof Abundance for All


This brief covers the trailing ~72 hours (July 3 – 6, 2026).

Primary sources:

Anthropic Ships Claude Sonnet 5, Redeploys Fable 5 With an Industry Jailbreak Framework, and Launches Claude Science

This brief covers the trailing ~72 hours (June 30 – July 3, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a busy, Anthropic-heavy stretch: a new Sonnet model, the redeployment of Fable 5 after export controls were lifted, a science workbench, and a new computational-biology benchmark from OpenAI.

Anthropic introduces Claude Sonnet 5

Anthropic · June 30, 2026

Anthropic released Claude Sonnet 5, which it calls its most agentic Sonnet model yet, positioning it close to Opus 4.8 performance at lower cost. The model is the new default on Free and Pro plans and is available on Claude Code and the Claude Platform via claude-sonnet-5, at introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026 (then $3/$15). Because it is somewhat stronger than Sonnet 4.6 on cyber tasks, it launched with real-time cyber safeguards enabled by default, though Anthropic says it still shows substantially weaker offensive-cyber ability than its Opus models.

“Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.” — Anthropic

Source: Introducing Claude Sonnet 5

Fable 5 redeployed globally as export controls lift; Anthropic proposes an industry jailbreak-severity framework

Anthropic · June 30, 2026

After the US government applied export controls to Claude Fable 5 and Mythos 5 on June 12 — prompting Anthropic to suspend both — the company said the controls were lifted on June 30 and that Fable 5 returned globally on July 1 across the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. Alongside the redeploy, Anthropic proposed a consensus framework for scoring the severity of AI jailbreaks — graded on capability gain, breadth, ease of weaponization, and discoverability — developed with Amazon, Microsoft, Google, and other Glasswing partners, plus a new HackerOne program and deeper US-government pre-release testing commitments. A July 2 follow-up post added further detail on the safeguards and framework.

“As of today, June 30, the export controls on Fable 5 and Mythos 5 have been lifted.” — Anthropic

Source: Redeploying Fable 5

Anthropic launches Claude Science, an AI workbench for researchers

Anthropic · June 30, 2026

Anthropic released Claude Science in beta on macOS and Linux for Pro, Max, Team, and Enterprise plans. It bundles a coordinating agent with more than 60 curated skills and connectors across genomics, single-cell, proteomics, structural biology, and cheminformatics, renders scientific artifacts like 3D protein structures and genome tracks natively, and manages compute from a laptop up to an HPC cluster or on-demand GPUs. Every figure ships with the exact code, environment, and message history that produced it, and a reviewer agent checks citations and calculations. Anthropic says it will fund up to 50 “AI for Science” projects with up to $30,000 in credits, with applications open through July 15.

“Claude Science brings these fragmented tools into a single research environment where scientists can conduct all stages of their work.” — Anthropic

Source: Claude Science, an AI workbench for scientists, is now available

OpenAI introduces GeneBench-Pro, a research-level computational-biology benchmark

OpenAI · June 30, 2026

OpenAI released GeneBench-Pro, a 129-question benchmark spanning 10 domains of computational biology that tests higher-order scientific judgment — handling ambiguity, revising assumptions, and choosing the correct analysis path — rather than rote execution. Each problem is built synthetically so the full causal structure is known and answers can be graded deterministically. OpenAI reports its strongest model, GPT-5.6 Sol, passes 28.7% at the highest reasoning level (31.5% with Pro mode), up sharply from under 5% for GPT-5 when the original GeneBench began. Reviewers estimated a typical problem would take a human expert 20–40 hours.

“Our strongest model, GPT‑5.6 Sol, attains a pass rate of 28.7% at the highest reasoning level (31.5% with Pro mode enabled). That is a sharp increase from when we began building the original GeneBench; at that time, our best frontier model, GPT‑5, scored below 5%.” — OpenAI

Source: Introducing GeneBench-Pro


This brief covers the trailing ~72 hours (June 30 – July 3, 2026).

Primary sources:

HP Scales Its OpenAI Frontier Partnership and California Adopts Claude Statewide

This brief covers the trailing ~72 hours (June 27–30, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. It was a quiet stretch for model launches, led instead by two notable enterprise and government adoption moves — HP scaling its OpenAI partnership and California signing a statewide Anthropic deal.

HP Inc. scales its OpenAI Frontier strategic partnership

OpenAI · June 28, 2026

OpenAI said HP Inc. will scale activation of its OpenAI Frontier strategic partnership after a series of successful pilots, moving from experiments to enterprise-wide deployment. The work spans customer- and partner-facing experiences, customer telemetry insights, employee productivity, and software development, with Frontier serving as the connective layer that governs access, context, deployment, and evaluation across HP’s agents and AI workflows. OpenAI cited early proof points, including one engineer moving through 122 pull requests across 43 projects in weeks and a security team estimating roughly 82 hours/week of capacity unlocked.

“It has been an amazing tool, and I am using it daily.” — an HP engineer, quoted by OpenAI

Source: HP Inc. launches Frontier strategic partnership with OpenAI

California adopts Claude statewide in a first-of-its-kind Anthropic partnership

State of California · June 29, 2026

Governor Gavin Newsom announced that California has entered a partnership with Anthropic giving all state agencies — plus cities and counties — access to Claude at a 50% discount, bundled with free workforce training and GenAI technical assistance. Claude becomes the first AI productivity tool offered through the California Department of Technology’s new Statewide Information Technology Shared Services (SITeS) portal. The state noted existing Claude use at the DMV (customer service and wait times), the Department of Health Care Services, and CDT/CalOES cyber defense work using Claude Security and Claude Code.

“AI should not replace the human work of government; it should help our workers move faster, solve problems more effectively, and deliver better results for Californians.” — Governor Gavin Newsom

Source: Governor Newsom announces a first-of-its-kind partnership providing Anthropic tools to state agencies

Still developing

Ornith-1.0 · DeepReinforce · June 25, 2026 — Just ahead of this window, DeepReinforce released Ornith-1.0, an MIT-licensed open-weights family for agentic coding (9B and 31B Dense, 35B and 397B MoE) built on pretrained Gemma 4 and Qwen 3.5. Its distinguishing feature is a self-scaffolding training framework in which the model learns to author both solution rollouts and the task-specific harnesses that guide them. DeepReinforce reports the 397B flagship scores 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified, matching Claude Opus 4.7. Source: Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding.


This brief covers the trailing ~72 hours (June 27–30, 2026).

Primary sources: