- Google's Gemini triple-drop + Gemini 4 confirmed: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber shipped July 21, with Gemini 3.6 Flash hitting 63.9% on MLE-Bench using 17% fewer tokens — and Google confirmed its most ambitious pre-training run yet for Gemini 4.
- OpenAI's GPT-5.6 powers ChatGPT Work: the Sol/Terra/Luna family (from $1/$6 to $5/$30 per 1M tokens) adds an autonomous agent that builds finished docs, sheets, and slides across your apps.
- Grok 4.5 and open coding models escalate: xAI ships Grok 4.5 at $2/$6 per 1M tokens and open-sources Grok Build, while DeepSeek V4, Qwen3.6, and Kimi K2.6 push open weights close to the frontier.
- Compute goes sovereign: Japan turns on a national AI factory on 27,500 NVIDIA Rubin GPUs, and the Model Context Protocol goes stateless in its biggest revision yet.
AI News (Top Updates)
1. Google ships the Gemini 3.6 Flash trio and confirms Gemini 4 is training
On July 21, Google released three models at once — Gemini 3.6 Flash (better coding, knowledge, and multimodal, using 17% fewer output tokens than 3.5 Flash), Gemini 3.5 Flash-Lite (the fastest 3.5-class model at 350 tokens/sec), and Gemini 3.5 Flash Cyber (a limited-access vulnerability-detection model paired with the CodeMender agent). Google also confirmed it has begun its 'most ambitious pre-training run yet' for Gemini 4. Gemini 3.6 Flash scores 63.9% on MLE-Bench (up from 49.7%) and 83.0% on OSWorld-Verified at $1.50 in / $7.50 out per 1M tokens; Flash-Lite runs at $0.30 / $2.50.
2. OpenAI's GPT-5.6 family powers the new ChatGPT Work agent
On July 9, OpenAI released GPT-5.6 in three tiers across ChatGPT, Codex, and the API — Sol (flagship, with an ultra mode that runs agents in parallel), Terra (balanced), and Luna (budget) — alongside ChatGPT Work, an autonomous 'finish-the-job' agent that sustains multi-hour projects and produces finished slides, sheets, and docs across Slack, Teams, Gmail, Drive, SharePoint, and Salesforce. Sol scores 52.7% on Agents' Last Exam and 92.2% on BrowseComp; pricing runs from Luna at $1 in / $6 out to Sol at $5 in / $30 out per 1M tokens.
3. Anthropic's Claude Fable 5 leads coding, and Claude for Teachers goes free
Claude Fable 5 (general use with safeguards) and Mythos 5 (gated to authorized cyber/bio researchers) headline the coding frontier — Stripe used Fable 5 to complete a 50-million-line codebase migration in a single day, priced at $10 in / $50 out per 1M tokens. On July 14, Anthropic launched Claude for Teachers, giving verified U.S. K-12 educators free premium Claude, teaching skills, and curriculum connections to state standards, with a pilot in the Detroit Public Schools Community District and no training on student data.
4. xAI ships Grok 4.5, open-sources Grok Build, and adds Automations
On July 16, xAI released Grok 4.5, its most capable model for coding, agentic tasks, and knowledge work — scoring 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro at $2 in / $6 out per 1M tokens, serving at ~80 tokens/sec and available in Grok Build and Cursor. A day earlier it open-sourced Grok Build, its agent tooling, and it launched Automations for scheduled, autonomous task execution — following a Voice Agent Builder (July 1) and 21 new flagship voices (July 6).
5. Japan turns on the world's first national AI infrastructure on NVIDIA
On July 16, Japan's government and industrial leaders partnered with NVIDIA to build a national 'physical AI' factory using 13,750 Vera CPUs and 27,500 next-generation Rubin GPUs. NVIDIA also reframed Vera Rubin around agentic post-training as the dominant compute workload (July 17) — training the largest models with a quarter of the GPUs of the Blackwell generation — and introduced new Jetson Thor edge modules (T3000 at 865 FP4 TFLOPS, T2000 at 400) for on-device robotics.
6. Meta reorganizes its frontier stack under 'Muse' and opens a hosted API
Meta Superintelligence Labs launched Muse Spark 1.1 (July 9), a multimodal agentic model with a 1M-token context, behind a new OpenAI-compatible Meta Model API in public preview — effectively its first developer-facing hosted frontier model. Two days earlier it shipped Muse Image (ranked No. 2 on LMArena for text-to-image and editing) and previewed Muse Video with native audio (No. 3 for text-to-video). No new open Llama model shipped in the window.
7. DeepSeek V4 lands as a fully open-weight, trillion-scale flagship
DeepSeek V4 reached its formal release with V4-Pro (1.6T total / 49B active) targeting top closed models on agentic coding, reasoning, and STEM, plus a fast, cheap V4-Flash tier (284B / 13B active). Both ship a 1M-token context window and support the OpenAI ChatCompletions and Anthropic API formats with thinking and non-thinking modes, with open weights on Hugging Face.
8. Hugging Face discloses an agent-driven intrusion and updates LeRobot
On July 16, Hugging Face disclosed that an autonomous AI agent breached its production infrastructure via a remote-code dataset loader and template injection, accessing some internal datasets and service credentials — though it found no tampering with public models, datasets, or Spaces, and used the open-weight GLM 5.2 to analyze 17,000+ attack events. Separately, LeRobot v0.6.0 (July 7) added world-model policies, new VLAs, six sim benchmarks, and a deployment CLI, closing the open robot-learning loop.
9. Microsoft adds GPT-5.6 to Copilot and expands its in-house MAI models
Microsoft made OpenAI's GPT-5.6 available in Microsoft 365 Copilot for agentic, end-to-end reasoning (July 9), while expanding its own MAI family in Foundry — MAI-Image-2, MAI-Voice-1, and MAI-Transcribe-1, which leads the FLEURS benchmark in 11 core languages from $0.36/hr and generates 60 seconds of audio in about one second. Copilot in SharePoint (July 16) also gained agent-driven generation of sites, interactive reports, and Office files.
10. The MCP spec goes stateless in its biggest overhaul yet
The Model Context Protocol shipped a release candidate for the 2026-07-28 spec that drops the initialize handshake and session management, so any request can hit any server instance behind an ordinary load balancer. It adds an Extensions framework with official MCP Apps (server-rendered UIs in sandboxed iframes) and a redesigned Tasks extension, and deprecates Roots, Sampling, and Logging under a new 12-month policy. Tier-1 beta SDKs are out for Python, TypeScript, Go, and C#.
11. Mistral pushes into robotics, documents, and prompt governance
Europe's flagship lab shipped Robostral Navigate (July 8), an 8B model that lets robots navigate complex indoor spaces from a single RGB camera and plain-language instructions — 76.6% success on unseen environments — plus Mistral OCR 4 (June 23), whose new include_blocks parameter returns paragraph-level bounding boxes for RAG pipelines. Mistral AI Studio also added version control and a 'system of record' for prompts and skills (July 9).
12. Open coding models keep closing the gap: Qwen3.6 and Kimi K2.6
Alibaba's Qwen3.6-27B is an Apache-2.0, dense 27B multimodal coder that beats a much larger MoE predecessor — 77.2% on SWE-bench Verified, 86.2% on MMLU-Pro, and 94.1% on AIME 2026 — with a native 262k context extensible toward 1M. Moonshot's open-weight Kimi K2.6 targets long-horizon autonomy with 4,000+ tool calls and 12+ hours of continuous execution, scoring 58.6 on SWE-Bench Pro and 83.2 on BrowseComp.
13. Media generation moves toward agents: Runway, Midjourney, and ElevenLabs
Runway shipped Agent Skills (July 2), a command-driven creative agent that executes full ad campaigns and localized variants from simple instructions, plus a speed-optimized image model. Midjourney V8.1 became the default model (June 10), rendering standard jobs 4-5x faster and producing 2K HD images without upscaling. ElevenLabs rolled out July updates across Music, its Speech Engine, and ElevenAgents, which now supports nested agent transfers and per-agent sentiment analysis.
14. The money and tooling around AI keep moving: Databricks, wearables, LangChain
Databricks is raising a Coatue-led strategic round valuing it at $188B (announced July 16), funding its Unity AI Gateway, the Genie AI coworker, and Lakebase. AI wearables are heating up too — smart-glasses maker Even Realities raised $150M at a $1B valuation, led by Meituan and Tencent (July 6). On the tooling side, LangChain's Interrupt 2026 shipped a LangSmith Engine that monitors production traces and opens PRs with fixes, plus SmithDB, an observability database up to 15x faster.
Top 5 New / Popular AI Products
1. Gemini 3.6 Flash
New Flash ModelGoogle's newest fast, cheap frontier-class model (July 21) — 63.9% on MLE-Bench and 83.0% on OSWorld-Verified using 17% fewer output tokens, at $1.50 in / $7.50 out per 1M tokens.
2. GPT-5.6 Sol
New Flagship ModelOpenAI's flagship (July 9), scoring 52.7% on Agents' Last Exam with an ultra multi-agent mode, priced at $5 in / $30 out per 1M tokens across ChatGPT, Codex, and the API.
3. Grok 4.5
Coding Agent ModelxAI's most capable model for coding and agentic work (July 16) — 64.7% on SWE-Bench Pro at $2 in / $6 out per 1M tokens, with the open-sourced Grok Build agent alongside it.
4. DeepSeek V4
Open-Weight Frontier ModelDeepSeek's V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B active) offer a 1M-token context with open weights on Hugging Face and OpenAI- and Anthropic-compatible APIs.
5. Muse Spark 1.1
Agentic Multimodal ModelMeta Superintelligence Labs' multimodal agentic model (July 9) with a 1M-token context, now available through the new OpenAI-compatible Meta Model API in public preview.
Discussion