llama-cpp
2 articles
llama.cpp's True Power — Georgi Gerganov Demonstrates 300 t/s on a Three-Year-Old Mac
llama.cpp creator Georgi Gerganov demonstrates Gemma 4 26B running at a blistering 300 tokens/s on a three-year-old Mac Studio M2 Ultra with speculative decoding. Throw in a WebUI and MCP support, and the whole ecosystem has become almost absurdly mature.
llama.cpp Hits 100k Stars — Georgi Gerganov's Love Letter to Local AI
llama.cpp just crossed 100k stars. Creator Georgi Gerganov reflects on the progress of local LLMs, the agentic era, 'good enough intelligence,' and why he believes an open, portable software stack is the only path that makes sense.