Deploying DeepSeek on-premise means downloading the open-weight models published by the Chinese lab DeepSeek and serving them on hardware you control. Doing that is technically the same as any self-hosted deployment. What is different is the governance conversation, and it is worth having before the hardware order.
The distinction that decides this whole discussion
There are two entirely different things people mean by "using DeepSeek", and conflating them causes most of the confusion. One is the hosted app and API operated by the company, where your prompts travel to their infrastructure under their terms. The other is downloading published model weights and running them on your own servers, where nothing leaves your network.
A number of governments and large organisations have restricted the hosted app on their devices and networks. Those restrictions are about the service, not about the mathematics in the weights file. That does not make a self-hosted deployment automatically acceptable in your organisation - but it means the question your security team should be asked is specific, and answerable, rather than a blanket yes or no.
What a self-hosted DeepSeek deployment actually sends outside
If you download the weights and serve them behind your own gateway, the answer is nothing. There is no callback, no telemetry, no licence check. The model file is inert until your own inference server loads it. Prompts, documents and outputs stay wherever you run it, which is the same guarantee you get from any other open-weight family on your own hardware.
Your security team will still want to see that demonstrated rather than asserted. Egress rules on the inference host, network capture during a test period, and a written record of where the weights were downloaded from and which version they are. That evidence is cheap to produce and it is what turns an argument into a sign-off.
Licensing: check the terms for your exact version
Open weights do not mean one licence, and terms differ between releases from the same lab. Some DeepSeek releases have been published under permissive open-source terms that allow commercial use and derivative models; others carry their own model licence with additional conditions. Distilled variants add another layer, because a smaller model built on top of a different base family may inherit that family's licence rather than DeepSeek's.
Read the licence text attached to the specific version you intend to deploy, record which version that was, and have counsel confirm it if the deployment will touch customer data or a regulated process. This is the step teams skip, and it is the one with legal consequences rather than technical ones.
Evaluating open models for your own servers?
We will test the candidates on your real tasks, size the hardware for the one that wins, and build the deployment so swapping models later is a configuration change rather than a rebuild.
Talk to an On-Premise AI SpecialistSizing the hardware for a mixture-of-experts model
DeepSeek's flagship releases use a mixture-of-experts design, and that architecture is routinely misread when people plan hardware. Only a fraction of the model's parameters are used for any single token, so throughput can be better than the total parameter count suggests. But the whole model still has to be resident in memory to be available for routing. Teams who read only the speed half of that trade-off order too little memory and discover it at load time.
The practical consequence is that the largest DeepSeek releases are multi-GPU deployments, not something to put on a single workstation card. Check the memory requirement stated on the model card for the exact variant and quantisation you plan to run, then add headroom for context length and concurrency. Our guide to choosing a GPU server for company AI covers how to work that out, and running vLLM in production covers the serving layer.
When the distilled variants are the better choice
Alongside the flagship models, DeepSeek has released much smaller distilled versions trained to imitate the reasoning behaviour of the large one. These run on modest hardware and are often the right answer for real business workloads - routing tickets, extracting fields, classifying documents, drafting routine replies.
The honest way to choose is to test both on your own examples rather than on a public benchmark. A distilled model that handles ninety per cent of your volume on one affordable server, with the occasional hard case escalated, frequently beats a flagship deployment that is impressive in a demo and expensive every day. Our comparison of the best open-source LLMs for business covers the wider field.
What these models are good at, and where to be careful
DeepSeek's reputation rests largely on reasoning - models trained to work through a problem in steps before answering, which shows up most clearly in maths, code and multi-step logic. If your workload involves reconciling figures, transforming data, or following a procedure with branches, that behaviour is genuinely useful and worth testing.
Two cautions go with it. The first is cost per answer: a model that thinks in steps produces far more tokens than one that answers directly, which means more compute and slower responses for the same question. On a high-volume routine task that trade-off usually goes the wrong way, and a plain instruction-following model is the better tool. The second is that reasoning quality does not transfer evenly across languages or domains - a model strong on code may be ordinary on your contracts.
So test on the work you actually have, and separate your workloads before you choose. High-volume classification, routing and extraction want a small fast model. A smaller number of genuinely hard requests can be sent to something heavier. Most deployments that work well in practice run both, which is another argument for the gateway described below.
The questions your security review will ask
Expect three, and prepare for them rather than improvising. Where did the weights come from, and can you prove the file has not changed since - answered with a checksum and a recorded source. What leaves the network at inference time - answered with egress rules and a capture. And who can query the model and what is logged - answered with a gateway that authenticates users, not an open port on a GPU box.
There is a fourth question that is about policy rather than engineering: is a model from this lab acceptable to us at all, given where it was trained and published. That is a legitimate organisational decision and it is not ours to make. What we would say is that it should be made explicitly, by the people accountable for it, rather than settled by whoever happened to write the pilot.
Keep the model replaceable
The single most useful design decision is to treat the model as the part most likely to change. Applications should talk to one internal gateway that handles authentication, routing, rate limits and logging, and the gateway decides which model serves a request. Done that way, trying a different family next quarter is a configuration change.
Done the other way - each application integrating directly with one model's API surface - every future change becomes a migration. That is the difference between a deployment that ages well and one that has to be rebuilt. The same reasoning applies whichever family you pick, which is why our Mistral on-premise guide and Llama on-premise guide reach the same conclusion from different starting points.
Getting started
Pick one workload with an owner. Settle the licence and the governance question in writing before hardware is ordered. Rent capacity for a month so real usage tells you what to buy. Test the flagship and a distilled variant against your own tasks, judged by the people who will use the output. Then build the gateway, authentication and logging before anyone outside the pilot group gets access.
Inwizards has been building software since 2009 and deploys self-hosted AI systems from teams in the US, UAE and India. We are model-agnostic by design: the on-premise AI page explains how a deployment runs, and AI agents covers what gets built on top once the model is serving reliably.