Google’s Gemini 2.5 Pro Hits 82.4% on GPQA Diamond—Beats OpenAI’s GPT-5.5 by 6.1 Points on Graduate-Level Science Reasoning
Google just seized the reasoning crown from OpenAI with a 6.1-point margin on graduate-level science problems. The company that spent 18…
Straiker Raises $64 Million Series A on June 29—Agentic Security Startup Grows Revenue 15× in Under a Year Protecting Fortune 500 AI Agents
A security startup protecting AI agents from AI agents just grew revenue 15× in under a year. The category didn’t exist 24 months…
Anthropic Launches Claude Tag on June 23—Slack AI Agent Writes 65% of Product Team’s Code, Runs on Claude Opus 4.8
Anthropic just revealed that 65% of its own product team’s code comes from an AI agent—and they’ve opened that same system to…
Google DeepMind’s GenCast Beats ECMWF ENS on 97.2% of 1,320 Weather Targets—AI Forecasts 15 Days in 8 Minutes, Outperforms Europe’s Gold Standard by 12.6% on Extremes
Europe’s operational weather forecasting system—the one MeteoSwiss and national meteorological services actually rely on—just got beaten…
Generalist AI Raises $400 Million at $2 Billion Valuation—NVIDIA and Bezos Back Physical AGI Startup’s 99% Success Rate Robot Models
A robotics company founded in 2024 just raised $400 million by solving the problem that has stalled physical AI for decades: teaching robots…
Inception Labs Launches Mercury AI at 700+ Tokens/Second—Matches GPT-4.1 Nano Performance While Outpacing Commercial LLMs
A diffusion-based language model just hit 700+ tokens per second while matching GPT-4.1 Nano on benchmarks. The speed gap between diffusion…
Qualcomm in Talks to Acquire Tenstorrent for $8-10 Billion—Jim Keller’s RISC-V AI Chip Startup Valuation Triples in One Year
Qualcomm just offered 3x last year’s valuation for a RISC-V chip startup. The bid tells you more about NVIDIA’s moat than…
DeepSeek Launches Janus-Pro-7B on January 27—Open-Source 7B Model Beats DALL-E 3 and Stable Diffusion on GenEval and DPG-Bench
A 7-billion-parameter open-source model just outperformed DALL-E 3 on two major image generation benchmarks. DeepSeek dropped Janus-Pro-7B on…
NSA to Run Classified AI Benchmarking for Military ‘Frontier Models’—Trump’s Executive Order 14409 Gives Government 30-Day Pre-Release Access
The NSA now gets first look at frontier AI models before anyone else—up to 30 days before public release. Executive Order 14409, signed June…
Visa Embeds Payment Network Into ChatGPT on June 10—AI Agents Can Now Shop Across 175 Million Merchant Locations
ChatGPT can now spend your money. Visa just wired 175 million stores into OpenAI’s agent layer—with spending limits and fraud checks,…