AI Model Comparisons Machine Learning Prompt Engineering8 Min Read Artur MarkusonSeptember 1, 2026 The Harness, Explained: Why the Same Scaffold Scores 65% or 74% on SWE-bench A 100-line Python harness scored 65% on SWE-bench Verified in July 2025. The same project now claims above 74% — and boots faster than Claude…
AI Model Comparisons Machine Learning Natural Language Processing6 Min Read Artur MarkusonAugust 28, 2026 172 Billion Tokens Show Every LLM Fabricates Above 10% at 200K Context The best of 35 open-weight models still invented answers 1.19% of the time about entities that provably did not exist in the source document.…
AI Model Comparisons AI News & Updates Machine Learning10 Min Read Artur MarkusonAugust 12, 2026 NVIDIA Nemotron 3.5 Lightning Launches at 1,200 Tokens/Second—30B MoE Outpaces Gemma 4 by 29× and Qwen 3.6 by 35% on Agent Tasks NVIDIA just shipped a 30B parameter model that outputs 1,200 tokens per second—29 times faster than Gemma 4 26B on identical hardware. This…
AI Model Comparisons AI News & Updates Open Source AI12 Min Read Artur MarkusonAugust 7, 2026 LG AI Research Launches K-EXAONE 2.0 with 750B Parameters—Korea’s Largest Open-Source Foundation Model Beats GLM-5.1 on Long-Context Benchmarks A Korean electronics conglomerate just shipped a 750-billion-parameter model under Apache 2.0 and beat China’s best on long-context…
AI Coding & Development AI Model Comparisons AI News & Updates10 Min Read Artur MarkusonAugust 5, 2026 Alibaba’s Qwen3.8-Max Launches with 2.4 Trillion Parameters—But Zero Official Benchmarks Alibaba just dropped a 2.4-trillion-parameter model claiming it’s “second only to Claude Fable 5.” Five days later,…
AI Model Comparisons AI News & Updates Generative AI10 Min Read Artur MarkusonJuly 28, 2026 Black Forest Labs Launches FLUX 3 on July 24—12.4B Parameter Multimodal Model Generates 20-Second Video Clips with Native Audio, Beats Runway Gen-4.5 in 77% of Comparisons A single 12.4-billion-parameter model now generates video with synchronized audio in one pass—and it’s already running robot arms on…
AI Model Comparisons AI News & Updates AI Startups & Companies10 Min Read Artur MarkusonJuly 26, 2026 OpenAI Launches GPT-4.5 on February 27—Intermediate Model Bridges GPT-4 and GPT-5 as Company Seeks $40 Billion at $340 Billion Valuation OpenAI is raising money at a valuation larger than McDonald’s while simultaneously admitting their flagship model isn’t ready yet.…
AI Model Comparisons AI News & Updates Generative AI10 Min Read Artur MarkusonJuly 6, 2026 Claude Fable 5 Reclaims #1 Spot with 100/100 Quality Index—First Model to Top July 2026 Leaderboards After 19-Day Export Control Suspension A perfect 100/100 quality score across 357 ranked models. That’s what Claude Fable 5 posted within 48 hours of coming back online after…
AI Model Comparisons AI News & Updates Machine Learning10 Min Read Artur MarkusonJuly 1, 2026 Google’s Gemini 2.5 Pro Hits 82.4% on GPQA Diamond—Beats OpenAI’s GPT-5.5 by 6.1 Points on Graduate-Level Science Reasoning Google just seized the reasoning crown from OpenAI with a 6.1-point margin on graduate-level science problems. The company that spent 18…
AI Model Comparisons AI News & Updates Machine Learning9 Min Read Artur MarkusonJune 28, 2026 Google DeepMind’s GenCast Beats ECMWF ENS on 97.2% of 1,320 Weather Targets—AI Forecasts 15 Days in 8 Minutes, Outperforms Europe’s Gold Standard by 12.6% on Extremes Europe’s operational weather forecasting system—the one MeteoSwiss and national meteorological services actually rely on—just got beaten…