vLLM is an open-source serving engine for running language models on your own hardware. Deploying it in production is less about installing it and more about three decisions made first: which model, how many people at once, and how fast answers must feel. Everything else follows from those.
What vLLM is, and where it sits in your stack
vLLM is the piece that takes a language model you have downloaded and turns it into something many people can use at the same time. It loads the model onto your graphics cards, keeps many requests in flight efficiently rather than handling them one after another, and exposes them through an API that most existing tools already know how to talk to. If you have run a model on a laptop and wondered what changes when a hundred colleagues use it, this is the answer to that question.
It is not the whole system. vLLM serves models; it does not decide who is allowed to ask, keep a record of what was asked, or route between several models. Those belong to the layer you build around it, and skipping that layer is the single most common reason an in-house deployment stalls before it reaches real users. The on-premise AI page describes what typically gets built on top.
The reason teams reach for it specifically is throughput. A naive setup processes requests in turn and leaves expensive hardware idle between them. vLLM was designed around keeping the card busy - batching requests continuously as they arrive and managing the memory each conversation consumes far more carefully than a simple loop would. On a fixed amount of hardware, that difference decides how many people you can serve.
Three decisions to make before you install anything
Which model, exactly
Not a model family - a specific model at a specific size and precision. This single choice determines your memory requirement, your hardware bill, and the quality your users experience. Teams that skip it end up sizing hardware for an imagined workload and discovering the mismatch after the purchase order. If you are still choosing, our comparison of the best open-source LLMs for business is the place to start.
How many people at once
Concurrency, not headcount. A thousand employees with access might produce a handful of simultaneous requests; twenty developers running a coding assistant can produce more load than all of them. Estimate the peak you actually expect and size for that, with honest headroom rather than an imagined worst case.
How fast it has to feel
An overnight document-processing job and a live chat assistant place completely different demands on the same hardware. Interactive use is judged on how quickly the first words appear; batch work is judged only on whether it finishes by morning. Decide which you are building, because tuning for one degrades the other.
Sizing memory without guesswork
GPU memory is the binding constraint, and it holds two things. The first is the model weights, which are fixed once you have chosen the model and precision. The second is working memory for every conversation currently in flight - the running state of each request, which grows with how long the conversation is and how many are happening at once. People size for the first and get surprised by the second.
Long context is where this bites hardest. A model asked to read a fifty-page contract consumes far more working memory per request than one answering a one-line question, so a deployment that felt comfortable in testing can run out of room the week somebody starts pasting in documents. Decide the longest input you will genuinely allow and size for that, rather than leaving it open-ended and hoping.
Quantisation - running the model in a compressed format - is the usual lever. It reduces the memory the weights occupy substantially in exchange for a small quality cost that most business workloads never notice. It is very often the difference between a model fitting on hardware you already own and needing to buy more. Test it against your own tasks rather than trusting a general claim either way. For the hardware side of this decision, our guide to sizing a GPU server works through the trade-offs.
The honest process is empirical. Load the exact model you intend to run, put realistic traffic through it at the concurrency you expect, with inputs as long as the ones your users will actually send, and watch what happens. Published figures age quickly and rules of thumb do not survive contact with a specific workload.
Planning a self-hosted model deployment?
Tell us which model you want to run, how many people will use it and what your data rules are. We will size it honestly, build the gateway and logging layer around it, and leave your team able to run it.
Talk to an On-Premise AI SpecialistThe layer you need around it
A gateway, not a direct connection
Applications should talk to one stable internal endpoint, not to the serving process itself. That indirection is what lets you swap models, add a second machine, or move from rented hardware to your own without touching every application that depends on it. Open models move fast enough that you will make that swap sooner than you expect.
Authentication tied to your directory
Access should come from the identity system you already run, so that when somebody leaves the company their access to the model ends the same way their access to everything else does. A shared key passed around a team is not access control, and it is the thing an auditor will find first.
Logging you can answer questions from
Who asked what, when, against which model version. In a regulated environment this is not optional, and in every environment it is what lets you work out why quality changed after an upgrade. Decide early how long these records are kept and who may read them - prompts often contain exactly the sensitive material the deployment was meant to protect. Our note on GDPR-compliant AI deployment covers why that retention decision matters more than the hosting choice.
Running it after launch
A production deployment needs the ordinary operational discipline any internal service gets: health checks, automatic restart, monitoring of memory and utilisation, and an alert when the queue grows faster than it drains. None of this is specific to AI, which is precisely why it gets forgotten - the project is treated as a model rollout rather than as a service going live.
Model upgrades deserve a deliberate routine. Keep the previous version available, run the new one alongside it against a fixed set of your own real questions, and compare the answers before switching. A newer model is not automatically better at your tasks, and discovering that after a cutover is a bad way to learn it. Keeping enough storage for two or three model versions at once is cheap insurance.
Then there is capacity. Utilisation tells you when you are close to the limit, but only if somebody is watching it. Assign that, along with patching the host and updating the serving software, to a named person or a support arrangement. A deployment with no owner degrades quietly.
vLLM or something simpler?
If a handful of people are experimenting, a lighter tool is faster to stand up and easier to reason about - we have written about where Ollama fits in a business and where teams outgrow it. The point at which vLLM earns its extra complexity is when you have real concurrent users, expensive hardware you want fully used, or a service other applications depend on. Before that, it is overhead.
Both approaches can sit behind the same internal endpoint, which means starting simple does not commit you to staying there. Keep the interface stable and the serving layer becomes an implementation detail you are free to change.
What teams get wrong
Three mistakes account for most stalled deployments. Sizing for the model weights alone and being surprised by concurrent load. Exposing the serving process directly to applications, so the first model change becomes a migration project. And launching without logging, then being unable to answer a security question three months later.
A fourth is subtler: treating the deployment as finished at launch. Models improve, usage grows, and the serving software updates. Whoever owns it needs time allocated, not goodwill. For the wider picture of what running this in-house costs against a cloud API, see our breakdown of on-premise versus cloud AI cost.
Getting started
Pick one workload with a clear owner - document question-answering for a single department is the usual first choice. Rent capacity rather than buying, so the first month teaches you what you actually need. Put the gateway, authentication and logging in from the beginning; retrofitting them is far more work than building to them. Then measure real usage for a quarter before committing to hardware.
Inwizards has been building software since 2009 and deploys and supports self-hosted AI systems from teams in the US, UAE and India. If you want the serving layer, the gateway and the handover documentation built together rather than assembled later, the on-premise AI page explains how a deployment runs and what your team owns afterwards. For agents that take actions in your systems on top of it, see AI agent development.