benchmarks
2 articles
Gemma 4 Arrives — Google's Open-Source Quartet Crushes Rivals on Token Efficiency but Still Lags in Intelligence
Google’s Gemma 4 open-source family comes in four sizes, all with multimodal support, reasoning mode, and 256K context. The flagship 31B uses 2.5× fewer tokens than Qwen3.5 27B but trails it by 3 intelligence points, while the tiny E2B can run on a phone.
Anthropic Exposes AI Benchmarks' Dirty Secret — Leaderboard Gaps Might Just Mean 'Bigger VM'
Anthropic found that agentic coding benchmark scores can swing by up to 6 percentage points based on hardware configuration alone — often more than the gap between top models on leaderboards. Next time someone claims a 2-3% lead, ask them what VM they ran on.