Google Ships Gemini 3.7 Flash, OpenAI Previews a 14X-Faster Ultrafast Mode on Cerebras, and Anthropic Details Claude’s Text Watermark

This brief covers the trailing ~72 hours (August 12–15, 2026). Every item below was confirmed on the originating organization’s own page, with a published date inside the window. Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash and halved the introductory token price; OpenAI previewed an Ultrafast API tier running GPT‑5.6 Sol at up to 14× the speed on Cerebras hardware; Anthropic published a detailed explainer on the text watermark coming to future Claude models under the EU AI Act; and Google DeepMind put a sign-language translation model into consumer products for the first time.

Google introduces Gemini 3.7 Flash at half the introductory price of 3.6 Flash

Google · August 13, 2026

Google released Gemini 3.7 Flash, positioned as its most intelligent “workhorse” model for coding and agents, arriving only three weeks after Gemini 3.6 Flash. The company reports substantial gains over 3.6 Flash on production-code quality (FrontierCode 1.1 Main, 43.6% vs. 34.4%), long-horizon software engineering (DeepSWE v1.1, 65.3% vs. 49.0%), complex document comprehension (GDP.pdf, 34.0% vs. 22.0%), and business workflow automation (AutomationBench, 30.4% vs. 17.0%), plus a WebDev Arena Elo of 1588 vs. 1538. Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, after which it doubles. The model ships with updated CBRN and cyber-offense safeguards and is available in Google Antigravity, AI Studio, Android Studio, Gemini Enterprise, and—for consumers—via Gemini Spark for AI Pro and Ultra subscribers.

“This release comes just three weeks after Gemini 3.6 Flash, and is a direct result of developer feedback and algorithmic innovations that we look forward to bringing to future models.” — Tulsee Doshi, Senior Director, Product Management, on behalf of the Gemini team

Source: Introducing Gemini 3.7 Flash

OpenAI previews Ultrafast: GPT‑5.6 Sol at up to 750 output tokens per second

OpenAI · August 13, 2026

OpenAI shared an early look at Ultrafast, a new API service tier that runs GPT‑5.6 Sol up to 14× faster than standard processing, generating up to 750 output tokens per second. The tier is powered by Cerebras and is in limited preview with a selected group of customers spanning coding, commerce, financial research, and support. OpenAI frames the point as removing the usual trade-off in which real-time latency meant dropping to a smaller model, and cites internal use in incident response—reading logs, analyzing traces, and preparing fixes while an outage is still unfolding—and in research, where overnight experiment batches compress into same-day iteration loops. Access expands as capacity grows.

“Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.” — OpenAI

Source: Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed

Anthropic explains the text watermark coming to future Claude models

Anthropic · August 14, 2026

Anthropic published a detailed explainer on the watermark that future Claude models will embed in generated text, implemented to comply with the EU AI Act after Anthropic and roughly 190 other signatories signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. The method is a version of Google DeepMind’s SynthID‑Text: rather than inserting hidden characters or extra tokens, it changes the source of randomness used when the model picks among equally good next words, leaving a key-detectable statistical pattern. Anthropic says the watermark carries no identifying information, costs nothing extra to serve, and is applied globally at launch because there is no durable way to scope it by region yet. Coverage is thin on factual passages, code, and light proofreading, where there are few free choices to encode into; a detection API is planned, and files such as .png or .svg get C2PA content credentials instead.

“Watermarking carries no identifying information and can’t be traced to a specific person, organization, or chat.” — Anthropic

Source: How Claude’s text watermark works

Google DeepMind ships SL2T, bringing ASL dictation to Gboard and Live Transcribe

Google DeepMind · August 12, 2026

DeepMind introduced SL2T, a massively multilingual sign-language-to-text translation model, and shipped it into consumer products for the first time: sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English. The model was trained on more than 100,000 hours of data across 50+ sign languages and scores 70 BLEURT zero-shot on the FLEURS‑ASL benchmark, which DeepMind says is well above any previously reported result. For privacy, an on-device MediaPipe Holistic model converts video into pose-landmark coordinates and the raw camera feed is discarded before anything reaches the server. DeepMind convened an AI Sign Language Advisory Committee of Deaf organizations and co-authored a joint impact report for the release.

“Sign languages aren’t simply ‘English on the hands.’ They require complex visual perception of fine-grained whole-body movements and full-fledged language translation.” — Google DeepMind Sign Language Team

Source: Putting sign language AI into users’ hands

OpenAI research finds the enterprise “frontier gap” tripling as work shifts to agents

OpenAI · August 12, 2026

OpenAI published two complementary studies—Enterprise Signals and a working paper, How Organizations Use AI: Evidence from ChatGPT—arguing that enterprise AI is moving from assistance to execution. As of June, Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers. Firms in the top 10% of usage now produce 8.3× as many output tokens per active user as median firms, up from 2.6× in January. Advanced capabilities track the same divide: 21% of weekly active users at frontier firms use Plugins and 19% use skills, versus 9% and 3% at typical firms. Codex adoption is spreading well beyond engineering—since February, weekly active enterprise users grew 108× in legal, 41× in sales, and 41× in recruiting, against 5× in engineering—and administrative data shows early-career employees sending 13 more messages per week than executives six months after adoption.

“Frontier firms—those in the top 10% of AI usage each month—now generate 8.3× as many output tokens per active user as typical firms, up from 2.6× in January.” — OpenAI

Source: From assistance to execution: How enterprises put AI to work


This brief covers the trailing ~72 hours (August 12–15, 2026).

Primary sources:

Leave a Reply