apple-silicon
2 articles
llama.cpp's True Power — Georgi Gerganov Demonstrates 300 t/s on a Three-Year-Old Mac
llama.cpp creator Georgi Gerganov demonstrates Gemma 4 26B running at a blistering 300 tokens/s on a three-year-old Mac Studio M2 Ultra with speculative decoding. Throw in a WebUI and MCP support, and the whole ecosystem has become almost absurdly mature.
Ollama Switches to MLX, Betting Big on Apple Silicon Local Inference
Ollama announces MLX-powered inference on Apple Silicon, targeting faster local performance for personal assistants and coding agents.