on-device-ai
2 articles
Gemma 4 Arrives — Google's Open-Source Quartet Crushes Rivals on Token Efficiency but Still Lags in Intelligence
Google’s Gemma 4 open-source family comes in four sizes, all with multimodal support, reasoning mode, and 256K context. The flagship 31B uses 2.5× fewer tokens than Qwen3.5 27B but trails it by 3 intelligence points, while the tiny E2B can run on a phone.
AI Doesn't Need to Memorize the Times Table Anymore: How Reasoning and Tool Calling Let Small Models Punch Above Their Weight
Apple MLX creator Awni Hannun argues intelligence-per-watt is rising because models no longer need to memorize answers they can compute. Reasoning and tool use may let 5B-15B models approach today's frontier, though the ceiling is still unknown.