llama.cpp
- About
- Run large language models locally with minimal setup. Open-source alternative to LM Studio.
- Catalogue
- No. 125
- Category
- AI
- GitHub stars
- ★ 0

Overview
llama.cpp is an open-source engine for running large language models and vision models on a wide range of hardware, locally or in the cloud. It is written in plain C and C++ without dependencies, and powers many other local AI tools. Its team also makes Llama, a small Mac menu bar app built on it.
What it does well
llama.cpp downloads and runs models straight from Hugging Face with one command, and serves them through an OpenAI-compatible API server. It is tuned for state-of-the-art performance on everything from laptops to servers, and is available as pre-built binaries, Docker images or source.
Why it’s in the collection
llama.cpp is the foundation of the local AI ecosystem: many popular apps for running models on your own computer are built on it. It is MIT-licensed and one of the most active open-source AI projects.
On GitHub
- Stars
- 0
- Forks
- 0
- Commits
- 0
- Repository
- ggml-org/llama.cpp
Languages
- C++ 56%
- C 15.8%
- Python 7.3%
- Other 20.9%


