
When sensitive information cannot leave the company
Private AI for businesses, in the cloud or on premises
Open-source models on cloud servers or hosted at your company, connected to your information. Neither the data nor the questions leave the organization.
What is private AI?
Private AI is an artificial intelligence system that runs within infrastructure controlled by the company, instead of sending the information to an external service. The model runs on a private server or in a private cloud, and the data it processes stays there.
It is not the opposite of using ChatGPT or Gemini: in many projects the sensible approach is to combine the two, decide for each process which information can leave and which cannot, and leave the application ready to switch models without modifying the code.
Why companies want private AI
The data does not leave the company
Contracts, technical documentation, customer data, or financial information are processed within the company's infrastructure.
Regulatory compliance
Keeping data and processes on private servers makes it easier to comply with the GDPR and with internal rules on where information can be kept.
Independence
The company decides which model it uses, in which version, and when it is updated, without depending on a provider changing the service, the price, or the terms. In the past, what we did not want was to depend on closed software; now the risk is depending on an AI provider.
Commercial use license
Some open-source models have licenses that allow commercial use without paying per query, and others set conditions. That is why the license is reviewed when choosing the model.
High but predictable cost
The investment is in the hardware, not in consumption per query. With intensive, continuous use, the cost does not grow with each question, until the volume requires more capacity.
Works without an external connection
Installed on premises, it keeps answering even if the internet connection or a provider's service goes down.
Private AI or commercial AI: what each one does
They do not compete on the same ground. A commercial model is general-purpose; a private AI is sized for what the company needs.
Commercial AI: general-purpose
For example, ChatGPT, Gemini, or Claude
Huge models that run in data centers with a very high infrastructure cost, spread across millions of users. That is why they are paid for per use.
- They know about almost everything: they write, program, translate, and reason through long tasks.
- They process text, images, and audio.
- The information sent to them leaves the company.
Private AI: specialized
Smaller, open-source models
Most companies do not have the budget for infrastructure equivalent to that of a commercial model, nor do they need it. A smaller model is chosen, focused on the tasks the company needs.
- If it only has to understand text, it does not need to process images.
- If it has to analyze numerical data, it does not need to know world history.
- With fewer parameters, it runs on an affordable server and performs well in its field.
- The information does not leave the company.
Which information can go to a commercial model and which cannot
With a commercial service, GDPR compliance is ensured by contract: the provider acts as data processor (Article 28), and that contract sets out where it processes the data, whether it uses the data to train its models, and how long it keeps it. With private AI, all of that depends only on the company.
Information that can go to a commercial model
Writing a marketing text, summarizing a public article, translating a product sheet, generating ideas or generic code. None of these tasks identifies people or compromises the company.
Information that should not leave the company without checking the GDPR and internal rules
CVs, case files, customer records, medical records, internal reports with real figures, or any document with special categories of data.
Three ways to implement private AI
In a private cloud
The model runs on infrastructure dedicated to the company, with a provider that does not share resources with other customers. It is the middle ground: there is no need to buy servers, but the data does not go to a public service.
On premises, on the company's servers
The model is installed on the company's servers, with the GPU that the volume of use requires. It is the option with the most control: the information does not leave the network, not even encrypted. It requires investing in hardware and maintenance.
Hybrid architecture
The most common approach in practice: processes with sensitive information stay inside and the rest use commercial models, which tend to be more capable and require no infrastructure. The boundary is decided process by process.
An AI that knows your company's documentation
It is one of the cases companies ask for most: being able to ask questions in natural language about manuals, procedures, technical data sheets, or contracts, and getting answers that come from their own documents, not from the model's general knowledge.
This does not require training a model from scratch, which takes enormous amounts of data and computing power. An existing model is used and given access to the company's documentation at the moment of answering: when a procedure or a price list changes, the model does not need to be retrained.
The six components of a private AI for internal documentation
From document to answer, everything within the company's infrastructure. The technologies for each component are examples: in each project they are chosen according to the type of documentation, the volume, and the available hardware, and they can be replaced when a better alternative appears.
1. Document repository
For example, ownCloud or Nextcloud
The company's document management system: manuals, procedures, contracts, case files. The documents stay where they are.
2. Chunking and embeddings
For example, bge-m3
Documents are split into chunks, and each one is converted into a vector (embedding) that represents its meaning, not its exact words.
3. Vector database
For example, Qdrant
It stores those vectors and, given a question, returns the chunks closest to it in meaning.
4. Language model
For example, gpt-oss, Mistral, or Qwen
An open-source language model, with a license that allows commercial use, that writes the answer with those chunks in front of it.
5. Model server
For example, Ollama or vLLM
It runs the model and the embeddings on the server's GPU. It is chosen according to how many people use it at the same time, and it can be changed without modifying the rest.
6. Query interface
For example, AnythingLLM or a custom interface
The chat interface: who has access, which workspaces (collections of documents) they ask about, and which documents each answer comes from.
Real case: private AI at a nonprofit organization
When information cannot leave the organization
A nonprofit organization wanted to query its documentation with AI: case files with highly sensitive data, internal protocols, and codes of conduct. This is precisely the information that cannot leave the organization.
The AI runs on a server belonging to the organization: the documents remain in its document management system, the model is open source and runs on that same server, and each team asks questions from a chat interface about the workspaces assigned to it. Nothing leaves the organization.
Permissions managed in a single system
Users, workspaces, and permissions are managed in the document management system. A service we developed replicates them in the query interface as soon as they change, and each uploaded document is automatically indexed in its workspace.
Validated with a consumer GPU
A consumer graphics card, not a data center one, runs the model, reduced in size to take up less memory, and the embeddings, for a few conversations at a time. With more memory, the context can be extended and more chunks retrieved for each answer.
Background indexing
Documents are indexed without interrupting use of the assistant. In testing, indexing several thousand public documents took almost a day, without the server running out of memory.
How we implement private AI in a company
Before starting: well-organized information
Private AI works with the company's data as it is. Before starting, it must be clear what information there is, where it is, what state it is in, and who can see it. If it is not well organized, the first step is data governance and, in the case of documents, document management.
Analysis of use and processes
What it will be used for: querying documentation, classifying information, or working within agents and automations. This analysis determines whether an on-premises server, a private cloud, or a hybrid architecture is needed.
Choice of model and infrastructure
Which model fits the use and the language, and what hardware it needs to answer in a reasonable time.
Connection to the data, without training the model
The model is not trained with the company's information: it is connected to it. If it has to answer from documentation, the documents are split into chunks, converted into vectors, and indexed by workspace; if it works with the ERP or the CRM, it queries their data with each user's permissions.
Integration with the company's tools
The query interface, access for each user, or the connection with the applications, agents, and automations that will use it. Users and permissions are synchronized from the system where they are already managed, so they do not have to be maintained twice.
Testing and adjustments
With real cases from the company, checking that each answer is correct and, when it comes from a document, that it cites its source.
Frequently asked questions about private AI
What is the difference between private AI and local AI?
Does any regulation require AI to be on private servers?
With private AI, does each person query only their own documentation?
Can private AI also be used for agents and automations?
What hardware does private AI need?
How much does private AI cost?

What documentation should your team be able to consult?
Tell us what documentation there is and what you want to be able to ask. We will tell you what is needed and whether an on-premises server or a private cloud is the better option.




