local-inference
1 articles
llama.cpp's True Power — Georgi Gerganov Demonstrates 300 t/s on a Three-Year-Old Mac
llama.cpp creator Georgi Gerganov demonstrates Gemma 4 26B running at a blistering 300 tokens/s on a three-year-old Mac Studio M2 Ultra with speculative decoding. Throw in a WebUI and MCP support, and the whole ecosystem has become almost absurdly mature.