The Harness, Explained: Why the Same Scaffold Scores 65% or 74% on SWE-bench
A 100-line Python harness scored 65% on SWE-bench Verified in July 2025. The same project now claims above 74% — and boots faster than Claude…
Tencent Open-Sources Hy4 Preview — 770B MoE, 49B Active, 1M-Token Context, Apache 2.0
Tencent shipped a 770-billion-parameter model under a real Apache 2.0 license — no MAU caps, no bespoke community terms. And it published zero…
MCP, Explained: How the Model Context Protocol Actually Works — and Why 91.8% of Servers Ship Without Auth
Every major AI vendor adopted MCP within five months of its release. Then researchers scanned 21,000 public MCP servers and found 91.8% had no…
StackOne Defender 0.8.2 in Practice: An Apache-2.0 Prompt-Injection Filter for Tool Calls — What Works, What Does Not
A bundled ONNX classifier now sits between your agent and its tools, catching a vendor-reported 88.7% of prompt injections. The interesting…
AI This Week: The Verification Layer — 5 Stories Where AI’s Weak Point Was Knowing What Was Real
Two startups raised $80.5M this week to do the least glamorous job in enterprise software: watch what employees actually do all day. The…
172 Billion Tokens Show Every LLM Fabricates Above 10% at 200K Context
The best of 35 open-weight models still invented answers 1.19% of the time about entities that provably did not exist in the source document.…
Skan AI Raises $63M to Map How Work Actually Happens—Because Enterprise Agents Keep Automating Processes That Don’t Exist
$681.5M went into AI agent startups in the first 15 days of August 2026. The largest cheques didn’t go to agents. They went to the layer…
Grok Still Leaks Full Chat Histories 11 Weeks After Disclosure—AES-256 Encrypted Prompts Beat Guardrails 40% of the Time
Grok’s safety classifier reads your prompt. It does not read what Grok decrypts afterwards. Adversa AI shipped attack instructions as…
NLPatent Rebrands to Clerq with Agentic AI That Compresses Patent Research from Weeks to 10 Minutes
Patent research that took your junior associates three weeks now runs in 10 minutes. The billable hour just got a lot more expensive to justify.
Sola Raises $17.5M Series A From Andreessen Horowitz After 5× Revenue Growth—Screen-Recording Automation Replaces Traditional RPA
Two MIT dropouts convinced a16z that watching someone work is better training data than writing code. Their $17.5M bet suggests RPA’s…