llama.cpp just hit 100,000 stars on GitHub. Georgi Gerganov marked the milestone by looking back at the project and the state of local AI.

But he wasn’t just celebrating. He opened by acknowledging that opinions on local LLMs are deeply polarized. Many nuances get glossed over, and the discourse tends to follow hype waves rather than actual analysis.

Against that backdrop, he chose not to argue — just to share facts and observations.

That Opening Joke Wasn’t Really a Joke

Georgi’s thread started with a half-serious prediction:

Since now 90% of the world’s code is written by AI agents, I predict that in 3-6 months, 90% of AI agents will run llama.cpp locally 😄

He immediately followed with “Jokes aside,” signaling clearly that the prediction was tongue-in-cheek. The real point: local LLM usage is growing, and he expects that trend to continue.

Mogu PSA:

My read is that the joke’s underlying logic isn’t that far-fetched. If the number of AI agents really does explode, the cost and latency pressure of routing all of them through cloud APIs gets serious. Local inference is at least a path worth taking seriously — though Georgi himself didn’t commit to any guarantees ╰⁠(⁠°⁠▽⁠°⁠)⁠╯


Local Agentic AI: Nobody Believed It a Year Ago — Now It’s Running

Georgi admits he didn’t expect the agentic era to reach local LLMs this fast. A year ago, the situation was:

  • Models were too large; long-context tasks were impractical
  • Memory and compute requirements were massive
  • There was no clear path toward meaningful agentic applications

Things started shifting last summer with the release of gpt-oss. In Georgi’s own words — fairly measured — that was the first glimpse of tool calling working reasonably well within the resource constraints of everyday devices.

Better models kept dropping, and now useful local agentic workflows are a reality.

Georgi expects this trend to continue, and thinks 2026 could become one of the most important years for the local AI movement. Note the phrasing: expectation, not proclamation.

Mogu PSA:

He used “likely” and “one of the most important” — not “definitely” or “the most important.” In an industry awash with absolutist pronouncements, this kind of hedging actually makes me trust that he’s speaking seriously rather than doing marketing (⁠⌐⁠■⁠_⁠■⁠)


You Don’t Need Frontier Intelligence to Do Useful Things

This is the most important section of the entire thread — and the easiest to overlook. Georgi listed three examples:

  • Automatically searching and sending emails → doesn’t need frontier intelligence
  • Summarizing articles or technical docs → doesn’t need trillion-parameter models
  • Controlling home appliances, turning off the garage light → doesn’t need massive GPU data centers

The core of this section is a view Georgi explicitly framed as personal judgment: he believes there’s a threshold of AI intelligence that humans can effectively comprehend and meaningfully utilize. Beyond that threshold, more intelligence is unnecessary at best — and counterproductive at worst.

Here’s the original:

I believe that there is a certain level of intelligence we as humans can comprehend and meaningfully utilize to improve our working process. Beyond that level, access to more intelligence becomes unnecessary at best and counterproductive at worst.

He also believes this “good enough” level of artificial intelligence is fully achievable locally — what’s been missing all along is simply the right software stack.

Mogu going off-topic:

My take is that he’s not arguing about whether bigger models are always better. He’s asking whether humans can stably integrate that intelligence into their workflows. If a system is so powerful that people can’t understand it or collaborate with it, that may become a different kind of friction. The phrase “counterproductive at worst” is the one worth pausing on ┐⁠(⁠ ̄⁠ヘ⁠ ̄⁠)⁠┌


The Only Technical Path That Makes Sense: Run on Every Device

Next, Georgi laid out a piece of technical philosophy — and his tone was emphatic:

From technical point of view, I think that llama.cpp + ggml is the only solution that actually makes sense.

His reasoning: AI’s software stack must run efficiently across all kinds of devices, hardware, and operating systems. This technology is too important to be locked into any single vendor. It has to be built in the open, with the community and independent hardware vendors. He believes that’s the only way to make a real long-term impact.

Mogu going off-topic:

“The only solution that actually makes sense” — Georgi really wasn’t being modest here (⁠๑⁠•⁠̀⁠ㅂ⁠•⁠́⁠)⁠و⁠✧

But his “only” isn’t about raw performance. It’s about direction. The core of his argument is vendor lock-in risk: if AI infrastructure gets tied to a single vendor’s ecosystem, the industry’s long-term development is playing at someone else’s table. Open, portable, community-driven — that’s the bet he’s making.


1,500+ Contributors, Still Accelerating

Georgi mentioned that llama.cpp now has over 1,500 contributors, and the project is still growing steadily. Beyond thanking maintainers and contributors, he specifically noted that reliable partners have supported the project throughout its journey.

His gratitude toward the team felt especially genuine:

I feel extremely lucky to be able to work together with so many talented contributors. Every day I learn something new and I feel there is so much more cool stuff that we are going to build.

A project where the founder says “I learn something new every day” — regardless of star count — says something about the quality of the community.


Closing

Georgi’s ending was restrained. No grand conclusion, no attempt to convince everyone right now:

I won’t try to convince you about what is currently and will be possible with local AI. We will just continue to build as usual.

Just keep building as usual. His bet is that once the smoke clears and people look objectively at what they’ve built together, the benefits will speak for themselves:

I am confident that after the smoke clears and we look objectively at what we have built together, the benefits will be obvious to everyone.

No arguing, no preaching — just keep shipping.

Mogu going off-topic:

“We will just continue to build as usual.”

In an industry full of “we will change the world” and “this is a turning point for human civilization,” this line is almost rebelliously quiet. But looking back, the open-source projects that actually make long-term impact tend to sound exactly like this: talk less, ship more, let the code speak (⁠ง⁠ ⁠•⁠̀⁠_⁠•⁠́⁠)⁠ง