vLLM
How vibe coders and indie builders use vLLM: 20 projects built with it since May 2026.
Projects built with it
20
Share of stated stacks, 4 wks
<0.1%
Last 4 weeks vs 8 before
too few to tell
Add-ons built for it
0
Share of stated stacks, per week
Share rather than counts, so changes in how many posts a subreddit lets through don't show up as trends.
Often used with
- Ollama 13
- OpenAI API 10
- LM Studio 8
- Claude API 8
- OpenRouter 6
- Gemini API 6
- llama.cpp 5
- Docker 4
What's built with it
Latest projects built with vLLM
- Hide The Annoying — Uses a fast classification model to collapse X/Twitter posts matching unwanted topics into summary bars that can be expanded or allowlisted per author.“I got tired of broken keyword muting on X/Twitter, so I built an open-source extension powered by fast decision models (Jev & Cloudflare Clef)”
- OpenTerminalUI — Provides a self-hosted stock research terminal with charts, fundamentals, screening, backtesting, and an AI agent that summarizes a stock from live data and filings.“Spent way too many weekends building my own "Bloomberg terminal" for Indian + US stocks. Here's the AI agent doing a quick read on CCL”
- TvT (ThisVideoThing) — Serves and organizes a local video library for offline streaming with automatic captions, thumbnails, and tagging.“TvT (ThisVideoThing) - is a free media server. I built it after buying Plex and its organization system, cloud logins, and premium prices drove me a little crazy to be honest. I looked elsewhere, but couldn't find an easy to use system, with no subscriptions or logins, so I built TvT.”
- Crucible — Runs a terminal AI agent harness that lets local LLMs call a built-in constraint solver and rewind file and conversation state via git checkpoints.“I built Crucible – A terminal AI agent harness powered by a custom functional logic programming language with a built-in constraint solver”
- Have AI API keys / self-hosted… — Lets users chat with multiple cloud and self-hosted LLMs on iOS using their own API keys, with MCP connectors, RAG memory, and parameter controls.“Have AI API keys / self-hosted models but no iOS client? I have an app for you!”
- FullPrice.LOL — Searches nearby retailers and auctions for overstock, clearance, and customer return items and sends alerts when a saved search matches.“I hate paying retail so I built a search engine for overstock and customer returns”
- OpenLawAI — Answers questions about laws with cited sources using local semantic search and a cloud LLM, and analyzes documents and drafts contracts.“I open-sourced a self-hosted legal AI assistant — local RAG with cloud LLM”
- FutureOS — Runs a shared AI coding and personal agent backend that is reachable from terminal, desktop, mobile, CLI, and chat bots, with per-tool-call approval.“I built an open-source agent that runs in my terminal and reports to my phone — one backend, five interfaces”
- aiops-fabric / ViewSense AI — Governs AI agents with scoped permissions, tool access, memory, execution budgets, human approvals, audit logs, and model provider abstraction.“AI agent platform fully local/self-hosted and Looking for developers”
- Prisma — Runs several AI models in parallel and merges their answers into one synthesized reply.“I built Prisma — a tool that runs multiple AI models at once and merges their answers into one”
- notetaker-canvas — Lets users sketch on an infinite whiteboard and have an AI read the canvas and draw its answer at a chosen spot.“Made a whiteboard where you can point at something and ask "what goes here" and the AI actually draws it there”
- LlamaForge — Provides a graphical control panel for launching and managing llama.cpp and vLLM local model servers and setting up agents.“Update to LlamaForge, my GUI control panel for llama.cpp: runs on Linux/macOS now, plus vLLM and agent setup”
- Arqon — Offers web tools that refactor code, let users query blockchains in natural language, edit images precisely, and generate long coherent books with AI.“Hello World! I would like to share a bit about my vibecoding journey”
- Council — Sends one question to several AI models in parallel, has them critique each other, and shows where their answers agree and disagree along with a synthesis.“Council — got tired of asking the same question to 3 different AIs and comparing by hand, so I built a Mac app that does it for me (solo, open source)”
- Inferix — Runs user-supplied model Docker images as serverless inference endpoints on AMD MI300X GPUs, scaling to zero when idle and billing per second.“I built a POC for serverless inference platform on AMD GPUs — 5-min demo, would love feedback before opening up”