You can run AI without sending data to a public cloud, and there is more than one way to do it. Self-hosting on your own servers is the strictest. Private tenancies, in-country regions and no-retention contracts sit in between. Which one you need depends on the rule you are actually trying to satisfy.
Write down the rule before you choose the architecture
Almost every conversation that starts with "we cannot send this to the cloud" turns out to be one of five different conversations, and they have different answers. It might be a clause in a customer contract. It might be a regulator's expectation about where personal or patient data is processed. It might be a client's security questionnaire that your sales team has to pass. It might be an internal policy written years ago for a different kind of system. Or it might be an instinct - a reasonable one - that nobody has written down yet.
Find the sentence. If it exists in a contract or a policy, read what it actually prohibits: naming a country, naming a category of provider, requiring that data not be used for training, requiring that it not be retained, or requiring that you can prove deletion. Those are five different requirements and only one of them forces hardware into your own building. Teams that skip this step routinely buy servers to satisfy a clause about retention that a contract amendment would have handled.
The four arrangements people call "not the cloud"
Ranked from strictest to most permissive. Each one is a legitimate answer to some version of the question, and the gaps between them are where most of the confusion lives.
On-premise, self-hosted AI
Open-weight models running on hardware you own, inside your own network. Prompts and documents are processed on machines you can point at. This is the only arrangement where the answer to "where is our data?" is a room rather than a contract, and it is the right answer when the data is genuinely restricted or when a client's security review will not accept a third party in the path. The cost is that patching, access control, backups, monitoring and physical security all become yours. The strictest version of this - no internet connection at all - is covered in our air-gapped AI deployment guide.
Private tenancy in infrastructure you control
The model runs on compute rented from a cloud provider, but inside an account and network you control, rather than through a shared service that many customers call. Data sits in storage you own, inside your own network boundary, and the provider is a landlord rather than a participant in the conversation. For a lot of teams this satisfies the real requirement while avoiding a hardware purchase. It is not, however, "not the cloud" if the sentence you are trying to satisfy names cloud providers - which is exactly why you read the sentence first.
An in-country or sovereign region
A hosted service, but one contractually pinned to processing and storing data in a named country or region. This is the right answer when the requirement is about geography rather than about who operates the machines - a common shape in European and Gulf procurement. Confirm two things in writing: that processing as well as storage stays in region, and what happens to support access from outside it. Our notes on on-premise AI in Europe and GDPR-compliant AI deployment cover the documentation side.
A hosted API with no-training and no-retention terms
The data does travel to a third party, but under terms that it will not be used to train models and will not be retained beyond the request. For a great deal of ordinary business content this is a sensible, well-governed choice, and pretending otherwise leads teams to spend heavily for no real gain. It is the wrong choice when you need to be able to say the data never left, because here it did - you are relying on a promise rather than on physics.
Need AI that keeps company data inside your own network?
We will map where your data would travel in each option, tell you which one actually satisfies the rule you are working to, and build the one you choose.
Talk to an On-Premise AI SpecialistThe hops nobody checks
Whichever arrangement you pick, the leak is rarely the model. It is the components around it, and three in particular. Embeddings: to make documents searchable, every passage is converted by a separate model, and if that step calls a hosted service while the chat model runs locally, the entire document set has already left before anyone asks a question. Telemetry: serving frameworks, interface layers and agent libraries often ship with analytics or crash reporting on by default, and some of it includes prompt content. Third-party stores: a managed vector database or a hosted logging tool can hold the same text as the documents themselves.
Draw the data path on one page before you buy anything, and write against every hop whose machine it is. The hops that leave your network are your decision list. We go through this in detail in open-source LLM data privacy, and the document-level version is in running a local LLM over confidential documents.
What to get in writing
For anything other than a fully self-hosted deployment, the protection is contractual, so the words matter more than the architecture diagram. Ask for four things explicitly: that your content is not used to train or improve models; that it is not retained after the request, or is retained for a stated period you have agreed; where processing happens, not only where data is stored; and who can access it for support purposes, from which countries. Then ask what happens on termination and how deletion is evidenced.
Vague reassurance in a sales call is not a term. If a provider will not put a sentence in the agreement, treat the absence as the answer. This is not legal advice - have your own counsel review anything you are relying on.
What self-hosting asks of you in return
Self-hosting is not automatically the responsible choice, and it is worth being blunt about the trade. You take on hardware, patching, access control, backups, monitoring and someone to call when the service stops on a Sunday. A small team running an under-maintained server with broad internal access is not more private than a well-governed hosted arrangement with a proper data processing agreement - it only feels that way, which is the more dangerous of the two states.
The deciding factor is usually whether the data is genuinely restricted and whether you have someone to own the infrastructure. If both are true, self-hosting is the honest answer. If only the first is, a private tenancy buys you the compliance position while you build the capability. The on-premise versus cloud AI cost comparison covers the money, and sizing a GPU server covers what you would actually be buying.
Running AI without sending data to the cloud: how to decide
The short decision guide. If a contract or regulator names a country, start with an in-country region. If it names cloud providers or requires that data never leave your premises, you are self-hosting. If the requirement is about training and retention only, a no-retention agreement may satisfy it outright. If the data is restricted and you have an infrastructure owner, self-host. If it is restricted and you do not, take a private tenancy now and revisit in a year. And if none of the five rules from the first section actually exists in writing, the honest first task is to get it written down rather than to buy servers.
What this looks like once it is built
In practice, most of what companies want from AI without the cloud is unremarkable: an assistant that answers from internal documents, a system that reads incoming email or invoices and files them correctly, a chat interface staff can use without pasting confidential text into a public tool. All of that runs perfectly well on your own hardware. Our private AI chatbot guide covers the staff-facing version, and self-hosted LLMs for business covers the operational side of keeping it running.
Getting started
Do it in this order: find the rule, draw the data path, pick the strictest arrangement the rule genuinely requires rather than the strictest one available, then build one workflow and measure it. The architecture decision is much easier once the sentence you are satisfying is written on the whiteboard next to it.
Inwizards has been building software since 2009 and deploys private AI systems from teams in the US, UAE and India. The on-premise AI page covers self-hosted deployment, AI agent development covers what gets built on top, and AI agents explains the workflows involved.