Thapa Technical — Dev News
AI News Daily Briefing July 25, 2026

AI News Today: Claude Opus 5 Lands, GPT-5.6 Race & Kimi K3 Goes Open – July 25, 2026

By Vinod Thapa 6 min read
Today's AI Briefing (TL;DR)
  • Anthropic ships Claude Opus 5: The new flagship (claude-opus-5) comes “close to the frontier intelligence of Claude Fable 5 at half the price,” holding pricing flat at $5/$25 per million tokens, scoring roughly 3x the next-best model on ARC-AGI 3, and adding a 2.5x-speed Fast mode.
  • The GPT-5.6 vs Gemini Flash frontier reset: OpenAI's GPT-5.6 family (Sol/Terra/Luna, from $5/$30 down to $1/$6) posts big agent and cybersecurity gains, while Google's Gemini 3.6 Flash becomes the new default workhorse using ~17% fewer output tokens at $1.50/$7.50.
  • Kimi K3 becomes the largest open-weight model: Moonshot's 2.8-trillion-parameter model tops Arena's Frontend Code leaderboard and self-reports beating Opus 4.8 and GPT-5.5, with free weights (~1.4TB) due July 27.
  • AMD bets up to $5B on Anthropic: AMD agreed to invest in Anthropic and supply tens of billions in Instinct-class AI servers — a direct challenge to Nvidia's data-center dominance.

AI News (Top Updates)

1. Anthropic launches Claude Opus 5, its new flagship

Released July 24, Claude Opus 5 (claude-opus-5) is described as coming “close to the frontier intelligence of Claude Fable 5 at half the price.” Anthropic claims state-of-the-art results on Frontier-Bench and GDPval-AA coding, about three times the next-best model on ARC-AGI 3 (novel problem-solving), and it surpasses Fable 5 on OSWorld 2.0 “at just over a third of the cost.” Pricing stays flat versus Opus 4.8 at $5 per million input tokens and $25 per million output, a new Fast mode runs at 2.5x speed, and Anthropic calls it its most aligned model to date.

Why it matters: The best-in-class option for agentic coding and reasoning just got cheaper to run — and it's already the strongest model on the Claude Pro tier.

2. Meta opens Muse Spark 1.1 to consumers

On July 24, Meta made Muse Spark 1.1 — its Superintelligence Labs flagship — available to consumers inside the Meta AI chatbot, adding prompts of up to 1 million tokens and the ability to spin up groups of AI agents for multi-minute research tasks. It shipped alongside two Facebook features: “Facebook Verified” (a free video-selfie checkmark for adults) and “Seller,” an AI-assisted Marketplace listing app. A developer/API preview of Muse Spark 1.1 landed earlier on July 9.

Why it matters: Meta's proprietary frontier model — the successor to open-weight Llama — is now reaching the billions of users on WhatsApp, Instagram, and Facebook.

3. OpenAI's GPT-5.6 family resets the coding and agents frontier

Announced July 9 and now the anchor of OpenAI's July wave, the GPT-5.6 family spans three tiers — Sol (flagship), Terra (balanced), and Luna (fastest) — across ChatGPT, Codex, and the API. Sol scores 53.6 on Agents' Last Exam (13.1 points over Claude Fable 5), 80 on the Artificial Analysis Coding Agent Index, and 73.5% on ExploitBench (versus GPT-5.5's 47.9%). Pricing per million tokens: Sol $5/$30, Terra $2.50/$15, Luna $1/$6.

Why it matters: You can now match the model tier to the task and the budget instead of overpaying for every call.

4. Google refreshes the Gemini Flash tier

On July 21 Google shipped three Flash-tier models at once: Gemini 3.6 Flash (the new default, cutting output tokens about 17% while improving coding and knowledge work at $1.50/$7.50 per million), Gemini 3.5 Flash-Lite (roughly 350 output tokens/sec at $0.30/$2.50), and Gemini 3.5 Flash Cyber, a vulnerability-detection model limited to governments and trusted partners. Separately, Gemma 4 12B brings native text, vision, and audio to an Apache-2.0 open model that runs on a 16GB laptop.

Why it matters: The model tier most apps actually ship on just got cheaper and faster — lower bills and quicker responses.

5. xAI ships Grok 4.5, adds Automations, and open-sources Grok Build

On July 16 xAI released Grok 4.5, its most capable model, built for coding and agentic work in collaboration with Cursor: 83.3% on Terminal Bench 2.1, 62.0% on DeepSWE 1.0, and a claimed ~4.2x token efficiency, priced at $2/M input and $6/M output. The same day it launched Automations for scheduled, autonomous Grok workflows, and a day earlier it open-sourced Grok Build, its app-building tool.

Why it matters: Near-frontier coding keeps getting cheaper, and the agent harness itself is now open for you to run and inspect.

6. Moonshot's Kimi K3 becomes the largest open-weight model yet

Moonshot AI's Kimi K3, unveiled July 16, is a 2.8-trillion-parameter model — more than double its prior 1T system — that tops Arena's Frontend Code leaderboard and self-reports beating Claude Opus 4.8 and GPT-5.5. It's live now via web and API at $3/M input and $15/M output, with open weights (roughly 1.4TB) scheduled for July 27. Alibaba has separately previewed Qwen3.8-Max, a 2.4T multimodal MoE.

Why it matters: Frontier-level capability is increasingly something you can self-host, not just rent — though provenance and diligence still matter.

7. AMD to invest up to $5B in Anthropic and supply AI servers

Reported July 22, AMD agreed to invest up to $5 billion in Anthropic and sell it tens of billions of dollars of Instinct-class AI servers, pairing an equity stake with a massive compute-supply deal. It arrives as AI capital keeps accelerating: Databricks is raising at a $188 billion valuation, and Crunchbase pegs global startup investment at a record ~$510 billion in the first half of 2026.

Why it matters: More chip competition means more compute and, eventually, better pricing across the tools you build on.

8. The Model Context Protocol goes stateless

The upcoming MCP 2026-07-28 spec removes the initialize/initialized handshake and the session-ID header, making the protocol core stateless so any request can hit any server instance. It adds routing headers, OAuth/OIDC authorization hardening, and an Extensions framework (including MCP Apps for server-rendered UI in sandboxed iframes), while deprecating Roots, Sampling, and Logging with 12-month-plus removal windows.

Why it matters: Stateless requests remove sticky routing and shared session stores — a major scalability unlock for enterprise agent deployments.

9. Black Forest Labs launches FLUX 3 — unified image, video, and audio

Debuting July 23 in limited access, FLUX 3 generates 20-second video clips with synchronized native audio alongside images and even robotics action prediction, using BFL's “Self-Flow” unified architecture. It was preferred over Runway Gen-4.5 in 77% of comparisons, and BFL disclosed a $300M Series B at a $3.25B valuation. Runway, meanwhile, shipped a unified Dev API and a Media Router that auto-selects the best model per generation.

Why it matters: Creating video, image, and audio from a single prompt is becoming one seamless workflow — and the open image lab is now well-capitalized for it.

10. Mistral rebrands Le Chat to “Vibe”

Mistral restructured its assistant into Vibe Work (agentic productivity), Vibe Code (a developer mode via CLI, VS Code extension, and web), and Vibe Chat, all at the same chat.mistral.ai URL with accounts and history preserved. The lab also introduced Robostral Navigate, its first embodied-navigation model for robotics, and released Mistral OCR 4 for document intelligence.

Why it matters: Mistral is repositioning from chatbot to a unified work-and-coding agent platform — and pushing into embodied AI.

11. Microsoft turns Copilot into apps and adds Claude on Azure

Microsoft shipped SharePoint Copilot Apps in public preview (July 9), letting developers embed interactive UI — data grids, forms, approval panels, KPI dashboards — directly inside the Microsoft 365 Copilot canvas using standard web frameworks. Anthropic's Claude also reached general availability in Microsoft Foundry, hosted on Azure via NVIDIA Blackwell systems with prompt caching and extended thinking, and Foundry agents can now publish straight to Copilot and Teams.

Why it matters: Copilot is moving from text answers to actionable apps, while Azure keeps hedging beyond OpenAI with native Claude access.

12. Hugging Face: a 1T open MoE and an open-robotics push

Thinking Machines released “Inkling” (July 15), a roughly 1-trillion-parameter (41B active) decoder-only Mixture-of-Experts model that natively handles text, image, and audio with a 1M-token context and day-0 support in transformers, SGLang, vLLM, and llama.cpp. Nvidia and Hugging Face also expanded the open-robotics LeRobot project with Isaac GR00T 1.7 and Isaac Teleop, while Ollama's v0.32 line improved local Gemma 4 tool calling.

Why it matters: The Hub remains the distribution center of the open-weight race, and open robotics is now its fastest-growing category.

13. Nvidia at SIGGRAPH: edge world models, MCP in creative tools, and Jetson Thor

At SIGGRAPH 2026 (July 20) Nvidia launched Cosmos 3 Edge, a 4B-parameter world foundation model for Jetson, RTX PRO, and DGX; a Synthetic Video Detector NIM microservice (~92% accuracy, processing 1080p in ~22ms); and MCP support across Firefly, Blender, Houdini, and Unreal Engine. It also positioned the Vera Rubin platform for post-training/agentic workloads and introduced new Jetson Thor modules for mainstream robotics.

Why it matters: World-foundation models are moving to the edge, and agents are getting standardized access to professional creative tools.

14. Policy & business: EU AI Act enforcement powers go live August 2

The European Commission's supervision and enforcement powers over general-purpose AI providers activate August 2, 2026 — the ability to demand documentation, run model evaluations, order market restrictions, and levy fines up to 3% of worldwide turnover or €15 million. It's the first hard enforcement mechanism aimed squarely at foundation-model providers, landing amid a record AI-led venture boom.

Why it matters: Provenance, documentation, and compliance now sit alongside raw capability — build model-agnostic and keep diligence tight.

Top 5 New / Popular AI Products

1. Claude Opus 5

New Flagship Model

Anthropic's new frontier model (claude-opus-5), near Fable-5 intelligence at flat $5/$25 pricing, about 3x the next-best model on ARC-AGI 3, with a 2.5x-speed Fast mode.

Why it’s trending: Best-in-class agentic coding and reasoning, now cheaper to run than the top competing flagships.

2. GPT-5.6 (Sol / Terra / Luna)

New Frontier Family

OpenAI's three-tier lineup across ChatGPT, Codex, and the API, from a $5/$30 flagship down to a $1/$6 speedster, with big gains on agents and cybersecurity benchmarks.

Why it’s trending: Match the model tier to the task and budget instead of paying flagship rates for everything.

3. Kimi K3

Largest Open-Weight Model

Moonshot's 2.8-trillion-parameter model that tops Arena's Frontend Code leaderboard, live via web and API at $3/$15, with open weights (~1.4TB) due July 27.

Why it’s trending: Frontier-class capability you can eventually self-host, escalating the open-weight race past the 1T mark.

4. FLUX 3

Unified Media Model

Black Forest Labs' new model generating images plus 20-second video clips with synchronized native audio via a “Self-Flow” architecture, preferred over Runway Gen-4.5 in 77% of comparisons.

Why it’s trending: One prompt for image, video, and audio — from the leading open image/video lab, now backed by a fresh $300M round.

5. Grok 4.5

New Coding Model

xAI's most capable model, built with Cursor for coding and agentic work: 83.3% on Terminal Bench 2.1, a claimed ~4.2x token efficiency, at $2/$6 per million tokens.

Why it’s trending: Near-frontier coding at aggressive token efficiency, with an open-sourced agent harness (Grok Build) to match.

Discussion

Leave a Reply
Ready to Break Into Tech?
Build Real, Job-Ready Skills

Join our live online classes and master full-stack web development, databases, and modern AI coding tools — step by step with expert mentorship.

Explore Our Courses
Also check: /blog/daily-tech-news-2026-07-24.php — Daily Tech News