Thapa Technical — Dev News
AI News Daily Briefing July 18, 2026

AI News Today: GPT-5.6 Ships in Three Tiers & Open-Weight Models Close the Gap – July 18, 2026

By Vinod Thapa 6 min read
Today's AI Briefing (TL;DR)
  • GPT-5.6, three tiers: OpenAI's new family — Sol, Terra, and Luna — leads with Sol, its best coding model yet (80 on the Artificial Analysis Coding Agent Index), plus the new agentic "ChatGPT Work."
  • Open weights catch up: Thinking Machines' Apache-2.0 Inkling and Moonshot's Kimi K3 land within ~5 points of Claude Fable 5 — the open-vs-closed gap is now roughly one generation.
  • Price war: xAI's Grok 4.5 claims Opus-class quality at $2/$6 per million tokens, undercutting the frontier and shipping straight into Cursor.
  • Platform shifts: Microsoft starts routing Copilot workloads to its own MAI models even as GPT-5.6 becomes Copilot's "preferred model," and Google rebrands NotebookLM as Gemini Notebook.

AI News (Top Updates)

1. OpenAI ships the GPT-5.6 family (Sol, Terra, Luna) and launches ChatGPT Work

On July 9, OpenAI released GPT-5.6 in three tiers — Sol (flagship), Terra (balanced), and Luna (cost-efficient) — targeting enterprise work, coding, and science. Sol posts 80 on the Artificial Analysis Coding Agent Index (2.8 points above Claude Fable 5) while using under half the output tokens and costing about a third less. Pricing runs from Luna at $1 in / $6 out to Sol at $5 in / $30 out per 1M tokens. Alongside it, "ChatGPT Work" turns ChatGPT into an enterprise companion that drafts documents, builds spreadsheets, and assembles presentations across desktop, web, and mobile.

Why it matters: You can match the model tier to the job — cheap Luna for bulk work, Sol for hard coding — and hand ChatGPT the actual multi-step work instead of just questions.

2. xAI releases Grok 4.5, an "Opus-class" coding model that undercuts on price

On July 8, SpaceXAI (xAI) shipped Grok 4.5, "built for coding, agentic tasks, and knowledge work" and pitched as "an Opus-class model, but faster, more token-efficient and lower cost." It reportedly beats Anthropic's Claude Opus 4.8 on several benchmarks, is priced at $2 in / $6 out per 1M tokens, and is available in Grok Build, Cursor (all plans), and the console — though not yet in the EU.

Why it matters: Near-frontier coding help keeps getting cheaper and faster, with three labs now competing directly on the coding tier.

3. Thinking Machines releases Inkling, a fully open frontier-class model

On July 15, Thinking Machines published Inkling on Hugging Face — a decoder-only multimodal Mixture-of-Experts model with ~975B total parameters (41B active across 256 experts), a 1M-token context window, and native text/image/audio processing. It ships under Apache 2.0 with day-0 support in transformers, SGLang, and llama.cpp, in a full BF16 checkpoint (2TB VRAM) and a quantized NVFP4 version (600GB VRAM).

Why it matters: A genuinely open, customization-first model at frontier scale — you can self-host and fine-tune it instead of renting a closed API.

4. The open-weight gap narrows to about one model generation

Between July 4–17, the open-weight landscape shifted sharply. Moonshot's Kimi K3 (2.8T parameters, 1M context; weights promised July 27) trails Claude Fable 5 by just 5.4 points on FrontierSWE and only 1 point on Terminal Bench 2.1. Combined with Inkling, open models moved from "China-only" to "US-plus-China" territory, and DeepSeek V4 graduated mid-July with legacy endpoints retiring July 24.

Why it matters: The open-weight labs keep closing the gap with closed frontier models — a real second-source option for teams that want control and lower cost.

5. Microsoft starts routing Copilot workloads to its own MAI models

Starting in July 2026, Microsoft began diverting some Microsoft 365 Copilot workloads away from OpenAI toward its in-house MAI models, with Excel as the first app (Outlook, Word, and PowerPoint expected to follow). Microsoft frames it as a gradual hybrid transition, driven by vendor independence, data governance, and per-token cost control.

Why it matters: The Microsoft–OpenAI relationship is shifting, and the AI powering your Office apps may increasingly run on Microsoft's own models.

6. GPT-5.6 becomes the preferred model inside Microsoft 365 Copilot

On the same day it launched (July 9), Microsoft made GPT-5.6 the "preferred model" powering Microsoft 365 Copilot — putting the newest frontier model in front of hundreds of millions of enterprise seats immediately, even amid public "breakup" chatter about the two companies' relationship.

Why it matters: If you use Microsoft 365 Copilot, you automatically got OpenAI's newest model the day it shipped — while Microsoft also hedges with MAI.

7. Google rebrands NotebookLM as "Gemini Notebook"

Announced July 16, Google renamed NotebookLM to Gemini Notebook, giving it a blue-and-purple Gemini gradient and availability inside the Gemini app (and, soon, Search's AI Mode). A Gemini 3.5 + Antigravity upgrade adds secure cloud computing for code execution and data analysis, rolling out to AI Pro subscribers. The tool now serves 30M+ users and 600,000+ organizations.

Why it matters: Your research notebooks now run real data analysis and code — not just summaries — under one unified Gemini brand.

8. Nvidia pushes edge AI with Jetson Thor and scales Nemotron open models

On July 15, Nvidia introduced new Jetson Thor computers — compact, power-efficient AI supercomputers that run foundation models directly on robots — for mainstream robotics and edge AI. In parallel, Japanese enterprises are building industry-specialized AI on Nvidia's open Nemotron models, and Nemotron 3 Ultra posted benchmark-leading results driving LangChain Deep Agents over proprietary models.

Why it matters: AI inference is moving off the cloud and onto physical devices, and Nvidia is becoming a model provider, not just a chip maker.

9. Hugging Face ships ML Intern, an autonomous ML research agent

Hugging Face released ML Intern, an open-source agent that autonomously discovers arXiv papers, analyzes citation graphs, retrieves datasets, writes code, and launches GPU training jobs via Hugging Face Jobs. It scores 32% on GPQA scientific reasoning (vs Claude Code's 23%) and beats Codex by 60% on HealthBench, running up to 300 iterations per task via a CLI plus mobile/desktop web apps.

Why it matters: An open agent that outperforms proprietary coding agents on scientific reasoning — and you can run it yourself.

10. Anthropic expands into classrooms and strengthens governance

On July 14, Anthropic launched Claude for Teachers, a free offering for verified K-12 and higher-ed educators, and committed $10M to Canadian AI research the same day. Earlier, on July 9, former Federal Reserve chair Ben Bernanke joined Anthropic's Long-Term Benefit Trust, and a new "reflect on how you use Claude" feature rolled out — all after the U.S. lifted export controls on Fable 5 / Mythos 5 in late June.

Why it matters: Anthropic is racing OpenAI and Google into education while adding heavyweight oversight as its frontier models scale.

11. Meta opens the Muse family with Muse Image, Muse Video, and Muse Spark 1.1

On July 7, Meta introduced Muse Image and Muse Video — generative media models where "Muse Image follows instructions faithfully, edits with precision, composes from multiple references." Two days later (July 9), Muse Spark 1.1 updated Meta's personal-superintelligence line. Separately, reports say Meta delayed its next flagship Llama amid an internal reorganization, leaning more closed-source.

Why it matters: Meta is doubling down on consumer creative AI — direct challengers to Nano Banana, Midjourney, and Sora — even as its open Llama leadership wobbles.

12. Mistral pushes into physical AI with Robostral Navigate and ships prompt governance

On July 8, Mistral introduced Robostral Navigate, its first model built for embodied navigation from a single camera, as the lab's valuation reportedly nears $23B. A day later, Mistral Studio added a "system of record" for prompts and skills — versioned, owned, and traceable — and reports place a new frontier-targeting open-weight sparse MoE family in July early access.

Why it matters: Europe's flagship lab is chasing robotics and enterprise AI ops, not just chat and documents.

13. Ollama v0.32 turns the local runner into an agent host

Ollama v0.32.0 (July 14) introduced a built-in interactive agent experience that launches when you run Ollama, with chat, code, web search, and work-delegation features. The follow-up v0.32.1 (July 16) improved Gemma 4 tool calling and multi-turn reasoning, while earlier v0.31.1 delivered nearly 90% faster Gemma 4 token generation on Apple Silicon via multi-token prediction.

Why it matters: Private, offline agents with no API bills — the local-first AI stack keeps maturing fast.

14. Google expands the open + multimodal Gemini stack

Google's recent Gemini 3.5 wave includes Live Translate (speech-to-speech across 70+ languages), 3.5 Flash with built-in computer use for desktop/mobile/browser automation, and Gemini Omni Flash (natively multimodal, public preview). On the open side, Gemma 4 12B runs locally on 16GB with vision and voice, and Nano Banana 2 Lite is Google's fastest, most cost-efficient image model.

Why it matters: Google is pushing capable multimodal agents both onto laptops (open Gemma) and into developer APIs.

15. Business & policy: SpaceXAI's identity crisis, Mistral's ~$23B, and open agent stacks

Fresh off going public and acquiring Cursor, SpaceXAI shipped Grok 4.5 and open-sourced Grok Build (July 15) — while Bloomberg reported an internal "identity crisis" at the merged company. Mistral's robotics push accompanies a reported valuation nearing $23B, and open agent stacks are proliferating: Grok Build, Hugging Face's ML Intern, and LangChain Deep Agents now all run on open or self-hostable foundations.

Why it matters: Money and open tooling are moving fast — the infrastructure to build and self-host serious agents is now widely available.

Top 5 New / Popular AI Products

1. GPT-5.6 Sol

New Flagship Model

OpenAI's new best coding model (July 9), scoring 80 on the Artificial Analysis Coding Agent Index while using under half the output tokens of rivals. Priced at $5 in / $30 out per 1M tokens, available in ChatGPT, Codex, and the API.

Why it's trending: Top coding performance with dramatically better token efficiency — better results for less spend.

2. Grok 4.5

Coding Agent Model

xAI's "Opus-class but cheaper" model (July 8) at $2 in / $6 out per 1M tokens, shipping directly into Cursor for all plans — reportedly beating Claude Opus 4.8 on several benchmarks.

Why it's trending: Near-frontier coding quality at a fraction of Opus pricing, right inside the editor.

3. Inkling (Thinking Machines)

Open Frontier Model

A ~975B-parameter multimodal MoE (41B active) with a 1M-token context, released Apache 2.0 on Hugging Face (July 15) with day-0 support in transformers, SGLang, and llama.cpp.

Why it's trending: A truly open, customization-first frontier model you can self-host and fine-tune.

4. Hugging Face ML Intern

Autonomous Research Agent

An open-source agent that discovers papers, retrieves datasets, writes code, and launches GPU training jobs — scoring 32% on GPQA (vs Claude Code's 23%) and beating Codex by 60% on HealthBench.

Why it's trending: An open agent that out-reasons proprietary coding tools on scientific tasks.

5. Ollama v0.32 (built-in agent)

Local AI Agents

Ollama v0.32 (July 14) launches an interactive agent when you run it — chat, code, web search, and work delegation — turning the local model runner into an agent-capable platform.

Why it's trending: Private, offline agents with no API bills — local-first AI keeps gaining momentum.

Discussion

Leave a Reply
Ready to Break Into Tech?
Build Real, Job-Ready Skills

Join our live online classes and master full-stack web development, databases, and modern AI coding tools — step by step with expert mentorship.

Explore Our Courses
Also check: /blog/daily-tech-news-2026-07-16.php — Daily Tech News