llama.cpp's True Power — Georgi Gerganov Demonstrates 300 t/s on a Three-Year-Old Mac

llama.cpp creator Georgi Gerganov demonstrates Gemma 4 26B running at a blistering 300 tokens/s on a three-year-old Mac Studio M2 Ultra with speculative decoding. Throw in a WebUI and MCP support, and the whole ecosystem has become almost absurdly mature.