- GPT-5.6, three tiers: OpenAI's new family — Sol, Terra, and Luna — leads with Sol, its best coding model yet (80 on the Artificial Analysis Coding Agent Index), plus the new agentic "ChatGPT Work."
- Open weights catch up: Thinking Machines' Apache-2.0 Inkling and Moonshot's Kimi K3 land within ~5 points of Claude Fable 5 — the open-vs-closed gap is now roughly one generation.
- Price war: xAI's Grok 4.5 claims Opus-class quality at $2/$6 per million tokens, undercutting the frontier and shipping straight into Cursor.
- Platform shifts: Microsoft starts routing Copilot workloads to its own MAI models even as GPT-5.6 becomes Copilot's "preferred model," and Google rebrands NotebookLM as Gemini Notebook.
AI News (Top Updates)
1. OpenAI ships the GPT-5.6 family (Sol, Terra, Luna) and launches ChatGPT Work
On July 9, OpenAI released GPT-5.6 in three tiers — Sol (flagship), Terra (balanced), and Luna (cost-efficient) — targeting enterprise work, coding, and science. Sol posts 80 on the Artificial Analysis Coding Agent Index (2.8 points above Claude Fable 5) while using under half the output tokens and costing about a third less. Pricing runs from Luna at $1 in / $6 out to Sol at $5 in / $30 out per 1M tokens. Alongside it, "ChatGPT Work" turns ChatGPT into an enterprise companion that drafts documents, builds spreadsheets, and assembles presentations across desktop, web, and mobile.
2. xAI releases Grok 4.5, an "Opus-class" coding model that undercuts on price
On July 8, SpaceXAI (xAI) shipped Grok 4.5, "built for coding, agentic tasks, and knowledge work" and pitched as "an Opus-class model, but faster, more token-efficient and lower cost." It reportedly beats Anthropic's Claude Opus 4.8 on several benchmarks, is priced at $2 in / $6 out per 1M tokens, and is available in Grok Build, Cursor (all plans), and the console — though not yet in the EU.
3. Thinking Machines releases Inkling, a fully open frontier-class model
On July 15, Thinking Machines published Inkling on Hugging Face — a decoder-only multimodal Mixture-of-Experts model with ~975B total parameters (41B active across 256 experts), a 1M-token context window, and native text/image/audio processing. It ships under Apache 2.0 with day-0 support in transformers, SGLang, and llama.cpp, in a full BF16 checkpoint (2TB VRAM) and a quantized NVFP4 version (600GB VRAM).
4. The open-weight gap narrows to about one model generation
Between July 4–17, the open-weight landscape shifted sharply. Moonshot's Kimi K3 (2.8T parameters, 1M context; weights promised July 27) trails Claude Fable 5 by just 5.4 points on FrontierSWE and only 1 point on Terminal Bench 2.1. Combined with Inkling, open models moved from "China-only" to "US-plus-China" territory, and DeepSeek V4 graduated mid-July with legacy endpoints retiring July 24.
5. Microsoft starts routing Copilot workloads to its own MAI models
Starting in July 2026, Microsoft began diverting some Microsoft 365 Copilot workloads away from OpenAI toward its in-house MAI models, with Excel as the first app (Outlook, Word, and PowerPoint expected to follow). Microsoft frames it as a gradual hybrid transition, driven by vendor independence, data governance, and per-token cost control.
6. GPT-5.6 becomes the preferred model inside Microsoft 365 Copilot
On the same day it launched (July 9), Microsoft made GPT-5.6 the "preferred model" powering Microsoft 365 Copilot — putting the newest frontier model in front of hundreds of millions of enterprise seats immediately, even amid public "breakup" chatter about the two companies' relationship.
7. Google rebrands NotebookLM as "Gemini Notebook"
Announced July 16, Google renamed NotebookLM to Gemini Notebook, giving it a blue-and-purple Gemini gradient and availability inside the Gemini app (and, soon, Search's AI Mode). A Gemini 3.5 + Antigravity upgrade adds secure cloud computing for code execution and data analysis, rolling out to AI Pro subscribers. The tool now serves 30M+ users and 600,000+ organizations.
8. Nvidia pushes edge AI with Jetson Thor and scales Nemotron open models
On July 15, Nvidia introduced new Jetson Thor computers — compact, power-efficient AI supercomputers that run foundation models directly on robots — for mainstream robotics and edge AI. In parallel, Japanese enterprises are building industry-specialized AI on Nvidia's open Nemotron models, and Nemotron 3 Ultra posted benchmark-leading results driving LangChain Deep Agents over proprietary models.
9. Hugging Face ships ML Intern, an autonomous ML research agent
Hugging Face released ML Intern, an open-source agent that autonomously discovers arXiv papers, analyzes citation graphs, retrieves datasets, writes code, and launches GPU training jobs via Hugging Face Jobs. It scores 32% on GPQA scientific reasoning (vs Claude Code's 23%) and beats Codex by 60% on HealthBench, running up to 300 iterations per task via a CLI plus mobile/desktop web apps.
10. Anthropic expands into classrooms and strengthens governance
On July 14, Anthropic launched Claude for Teachers, a free offering for verified K-12 and higher-ed educators, and committed $10M to Canadian AI research the same day. Earlier, on July 9, former Federal Reserve chair Ben Bernanke joined Anthropic's Long-Term Benefit Trust, and a new "reflect on how you use Claude" feature rolled out — all after the U.S. lifted export controls on Fable 5 / Mythos 5 in late June.
11. Meta opens the Muse family with Muse Image, Muse Video, and Muse Spark 1.1
On July 7, Meta introduced Muse Image and Muse Video — generative media models where "Muse Image follows instructions faithfully, edits with precision, composes from multiple references." Two days later (July 9), Muse Spark 1.1 updated Meta's personal-superintelligence line. Separately, reports say Meta delayed its next flagship Llama amid an internal reorganization, leaning more closed-source.
12. Mistral pushes into physical AI with Robostral Navigate and ships prompt governance
On July 8, Mistral introduced Robostral Navigate, its first model built for embodied navigation from a single camera, as the lab's valuation reportedly nears $23B. A day later, Mistral Studio added a "system of record" for prompts and skills — versioned, owned, and traceable — and reports place a new frontier-targeting open-weight sparse MoE family in July early access.
13. Ollama v0.32 turns the local runner into an agent host
Ollama v0.32.0 (July 14) introduced a built-in interactive agent experience that launches when you run Ollama, with chat, code, web search, and work-delegation features. The follow-up v0.32.1 (July 16) improved Gemma 4 tool calling and multi-turn reasoning, while earlier v0.31.1 delivered nearly 90% faster Gemma 4 token generation on Apple Silicon via multi-token prediction.
14. Google expands the open + multimodal Gemini stack
Google's recent Gemini 3.5 wave includes Live Translate (speech-to-speech across 70+ languages), 3.5 Flash with built-in computer use for desktop/mobile/browser automation, and Gemini Omni Flash (natively multimodal, public preview). On the open side, Gemma 4 12B runs locally on 16GB with vision and voice, and Nano Banana 2 Lite is Google's fastest, most cost-efficient image model.
15. Business & policy: SpaceXAI's identity crisis, Mistral's ~$23B, and open agent stacks
Fresh off going public and acquiring Cursor, SpaceXAI shipped Grok 4.5 and open-sourced Grok Build (July 15) — while Bloomberg reported an internal "identity crisis" at the merged company. Mistral's robotics push accompanies a reported valuation nearing $23B, and open agent stacks are proliferating: Grok Build, Hugging Face's ML Intern, and LangChain Deep Agents now all run on open or self-hostable foundations.
Top 5 New / Popular AI Products
1. GPT-5.6 Sol
New Flagship ModelOpenAI's new best coding model (July 9), scoring 80 on the Artificial Analysis Coding Agent Index while using under half the output tokens of rivals. Priced at $5 in / $30 out per 1M tokens, available in ChatGPT, Codex, and the API.
2. Grok 4.5
Coding Agent ModelxAI's "Opus-class but cheaper" model (July 8) at $2 in / $6 out per 1M tokens, shipping directly into Cursor for all plans — reportedly beating Claude Opus 4.8 on several benchmarks.
3. Inkling (Thinking Machines)
Open Frontier ModelA ~975B-parameter multimodal MoE (41B active) with a 1M-token context, released Apache 2.0 on Hugging Face (July 15) with day-0 support in transformers, SGLang, and llama.cpp.
4. Hugging Face ML Intern
Autonomous Research AgentAn open-source agent that discovers papers, retrieves datasets, writes code, and launches GPU training jobs — scoring 32% on GPQA (vs Claude Code's 23%) and beating Codex by 60% on HealthBench.
5. Ollama v0.32 (built-in agent)
Local AI AgentsOllama v0.32 (July 14) launches an interactive agent when you run it — chat, code, web search, and work delegation — turning the local model runner into an agent-capable platform.
Discussion