·3 min read
Private AI: Running LLMs On-Premise Without Sending Data to the Cloud
When customer data cannot leave your network, a locally hosted language model is often the only realistic way to use AI. Here is what it takes: models, hardware, integration and the trade-offs.
Many companies in Germany want to use AI, and many of them are not allowed to send customer or company data to an external AI service. Contracts, GDPR obligations, works council agreements or simply a cautious IT department say no. That does not have to be the end of the AI project. Open-source language models can run on your own servers, inside your own network, and many practical use cases work very well with them.
This article summarises what I have learned setting up locally hosted LLMs for business applications: when it makes sense, what you need, and where the limits are.
When a private LLM is the right choice
A locally hosted model is worth it when at least one of these is true:
- The data is sensitive. Customer records, contracts, health or financial data, internal documents.
- The data must stay in the EU or on-site because of contracts or regulation.
- The use case is narrow and repeatable: classifying documents, extracting fields, answering questions from your own knowledge base, guiding users through a workflow.
- You expect steady, high volume, where per-token cloud pricing adds up.
If your use case needs the very best reasoning available today, or the data is not sensitive, a cloud model is usually faster to start with. Many projects use both: a local model for anything that touches personal data and a cloud model for the rest.
What you actually need
A model that fits the task
You do not need the largest model. For structured tasks such as extraction, classification, summarising or tool calling, models in the 7–14 billion parameter range are often good enough, especially with a clear prompt and examples. Larger models (around 70B) are noticeably better at open-ended reasoning but need far more hardware.
The deciding factors are language quality (test it in German if your users speak German), support for structured output or function calling, and the licence.
Hardware
The rule of thumb is GPU memory. Quantised to 4 bit, a model with 7–8 billion parameters fits into roughly 6–8 GB of VRAM; a 70B model needs around 40 GB or more. Plan headroom for the context window and for several users at once. A single workstation GPU is enough for a pilot; production with many parallel users needs proper sizing.
A serving layer
Tools like Ollama or vLLM load the model and expose an HTTP API that looks much like the cloud APIs. Your applications talk to that endpoint instead of the internet. Put it behind a reverse proxy with TLS and authentication, log requests, and monitor latency and GPU usage like any other service.
Integration with your systems
The model alone does not create value. The value comes from connecting it to your data and processes:
- Retrieval (RAG): index your documents and let the model answer from them, with sources.
- Tool calling: let the model call your existing APIs, for example to look up a customer or create an order, while the business rules stay in your backend.
- Guardrails: validate every model output before it changes data, and keep a human in the loop for anything critical.
Trade-offs to be honest about
- Quality: a well-chosen local model is very good at focused tasks, but it is not a top cloud model. Measure it on your real data before you promise anything.
- Operations: you own updates, monitoring and capacity. Budget time for that.
- Cost profile: hardware is an up-front investment; in return there is no per-request bill and no data leaving the building.
How I approach these projects
- Pick one use case with clear value and measurable success criteria.
- Build a small evaluation set from real (anonymised) examples.
- Compare two or three candidate models on that set, on the target hardware.
- Integrate step by step: read-only first, then actions with validation.
- Hand over with documentation, monitoring and a runbook for your team.
In a recent project this approach let users on handheld terminals, the web and Windows work with an AI assistant in German, without any customer data leaving the company network.
If you are thinking about AI but data protection keeps blocking the discussion, a private LLM is often the way forward. I am happy to look at your use case in a free call.
Get the AI Readiness Checklist
10 questions to answer before you start an AI project, plus occasional notes on applied AI and SaaS. Free.