Menu

Private AI for business: where your data goes

The question always arrives at the same moment: a team wants to use an assistant on its documents, and someone asks where the content goes. It is the right question, and it deserves more than a reassuring answer.

The word "private" covers three realities that are regularly conflated. Telling them apart takes five minutes and saves months of misunderstanding.

Private AI for business: where your data goes

What "private" means, and what it does not

Private AI
Your content is read by nobody else, is not used to train a model, and is not kept beyond what you decided. A matter of contract and guarantees, not of servers.

Local AI
The model runs on your machines. That neither stops detailed logs from leaving for an outside service, nor makes a badly configured setup safer than a well-contracted hosted one.

Open model
Its weights are published. That says what you may do with it, not where your data goes. Where the computation runs is an architecture detail; what matters is who is allowed to read.

The three questions to ask a vendor

Who can read my content? The vendor, its subcontractors, its support teams. The answer must be in writing.

Is my content used to train a model? The most important question and the most often sidestepped. A "no" belongs in the contract, with its exceptions if any.

How long is it kept, and can I have it deleted? A professional conversation holds names, amounts, sometimes whole case files.

What leaves and what stays

Data moves at four moments: the question asked, the documents read to answer, the answer itself, and the technical logs that keep a trace.

People reason about the first and forget the other three. Yet it is often the log, kept "for debugging", that holds the most for the longest. Ask for the log retention policy, not just the conversation one: it is the question that most often catches vendors off guard.

The guardrails that actually hold

Framing an assistant with written instructions helps with tone and is not enough for confidentiality. An instruction is an intention: it holds in most cases and gives way in the ones that matter.

The limits that hold are technical. Access: an assistant cannot reveal what it never could read. Action: the list of what it can trigger is written, everything else forbidden by default. Approval: anything binding goes through a person.

When it is your customers' data

When the assistant talks to your customers, you become responsible for content that is not yours. Two reflexes cover most of it: collect only what is needed to answer, and set a short retention period by default.

What compliance actually asks for

The European framework does not require a model in your basement. It requires knowing which data is processed, why, by whom, for how long, and being able to demonstrate it. Confidentiality is a matter of organisation and evidence before it is a matter of servers.

Frequently asked questions

Not necessarily. In-house hosting is one way to obtain guarantees, not the guarantee itself. An outside service committing in writing on reading, training and retention can offer a higher level than a poorly maintained local setup.

Ask this first, and get a contractual answer. A written no-training commitment, with its exceptions, beats any sales statement.

GDPR does not oppose using an assistant. It requires knowing which data is processed, for what purpose, for how long, and being able to prove it. That work belongs to project scoping, not to the aftermath.

Your documents, your rules

Let's talk about what you want to entrust to an assistant, and what must stay out of its reach.