llama.cpp
How vibe coders and indie builders use llama.cpp: 32 projects built with it and 4 built for it since May 2026.
Projects built with it
32
Share of stated stacks, 4 wks
0.1%
Last 4 weeks vs 8 before
too few to tell
Add-ons built for it
4
Share of stated stacks, per week
Share rather than counts, so changes in how many posts a subreddit lets through don't show up as trends.
Often used with
- Ollama 15
- LM Studio 12
- OpenRouter 7
- Python 6
- OpenAI API 6
- Claude API 6
- Gemini API 5
- vLLM 5
What's built with it
Latest projects built with llama.cpp
- Xencode — Runs a terminal coding assistant that talks to local or cloud models and edits a workspace with approval checkpoints.“I built an offline-first AI coding assistant in Rust because I didn't want my code sent to cloud APIs”
- Xtream Agentic Harness — Runs a terminal AI assistant that keeps tool outputs out of long-term chat history, uses XML-based tool calling that works across models, and archives each turn for restoring.“Spent 9 months building my own harness”
- Summon — Provisions disposable cloud GPU workers on demand, attaches a cached model, runs inference, and destroys the worker afterward.“I got annoyed at paying for idle cloud GPUs or fighting for inventory, so I built a system that does the searching for me.”
- Orihai — Runs long-form interactive fiction over weeks, keeping characters, locations and facts consistent by compiling each turn into a budgeted context.“I made an interactive app for long-form fiction you play over weeks”
- Crucible — Runs a terminal AI agent harness that lets local LLMs call a built-in constraint solver and rewind file and conversation state via git checkpoints.“I built Crucible – A terminal AI agent harness powered by a custom functional logic programming language with a built-in constraint solver”
- Have AI API keys / self-hosted… — Lets users chat with multiple cloud and self-hosted LLMs on iOS using their own API keys, with MCP connectors, RAG memory, and parameter controls.“Have AI API keys / self-hosted models but no iOS client? I have an app for you!”
- Otaku — Runs LLM-powered roleplay chats with automatic summarization, character extraction, and transparent context building, through a web UI or terminal UI.“Otaku — an LLM frontend for roleplay”
- Warren — Captures pages you read in Chrome and builds a local, searchable markdown wiki of them using a local vision model.“Warren: self hosted browsing memory. Local LLM reads what you read, you get a searchable wiki out of it.”
- Living Web: AI Assistant & Content Summarizer — Runs local or cloud LLMs in a browser sidebar to summarize pages, chat with documents, explain or translate text, and automate basic form filling and table exports.“I built a free Chrome extension to run local LLMs (Ollama, LM Studio) and cloud APIs in a browser sidebar”
- I shipped my first… — Generates invoices with an AI assistant that runs locally on the user's machine.“I shipped my first llama.cpp-powered desktop app — an invoice generator where the AI (and your data) never leaves your machine”
- warpdrv — Runs local LLMs through a desktop harness with built-in tools, sub-agents, voice chat, and a second AI that reviews outputs.“Built a Harness for LLMs using locally-run Qwen”
- Kallilex — Corrects, shortens, or rephrases selected text via a global hotkey and inserts the result back in place using any configurable AI endpoint.“(Open Source): Kallilex - Correct text directly, without opening a tab!”
- Musiclyse — Analyzes a song's audio with multiple models and lets a local LLM discuss and compare tracks with the user in a terminal chat.“I'm developing a music to LLM chatbot terminal in Python using Qwen 3.8 27B”
- Otaku — Runs a terminal roleplay chat client for LLMs that shows the exact context sent, auto-summarizes older messages, and extracts characters from chats.“Otaku — a roleplay terminal client”
- LLMs Gateway — Turns a local machine into an OpenAI-compatible endpoint that installs and serves GGUF models from Hugging Face in per-capability llama.cpp containers.“I built an open-source LLM inference gateway — search HF, download GGUF models, and serve them via llama.cpp with per-capability Docker containers”
Latest add-ons built for llama.cpp
- local-llmup — Installs and switches local LLM models across runtimes such as Ollama, llama.cpp, MLX, and LM Studio.“Local-llmup : package to manage localllm , install or switch in any runtime : ollama , llama.cpp , mlx , lm studio”
- LlamaForge — Provides a graphical control panel for launching and managing llama.cpp and vLLM local model servers and setting up agents.“Update to LlamaForge, my GUI control panel for llama.cpp: runs on Linux/macOS now, plus vLLM and agent setup”
- I built a GUI control panel for… — Provides a graphical control panel for configuring and launching llama-server with llama.cpp model settings.“I built a GUI control panel for llama.cpp so I'd stop hand-editing models.ini and llama-server flags”
- claudely — Launches Claude Code against local LLM servers like LM Studio, Ollama, and llama.cpp without modifying the user's real Claude config.“claudely: launch Claude Code against Local LLM provider like LM Studio / Ollama / llama.cpp without trashing your real claude config”
Pitched as alternatives to llama.cpp
- v7multiplikator — Runs transformer language model inference directly on the CPU using hand-written x64 assembly with AVX2 kernels.“LLM Inference in X64 Assembler”