Menu

AI guardrails: what actually keeps an assistant in check

2 min read

Definition

A guardrail is a limit placed on an assistant to keep it inside its scope. The guardrails that hold are not written instructions but technical restrictions: what the assistant can access, and what it is allowed to trigger.

The most frequent design mistake

People think they are containing an assistant by writing instructions: do not talk about this, never answer that. Useful for tone, insufficient for safety.

A written instruction is an intention. It holds in most cases and gives way in the ones that matter, when the question is phrased differently than expected. A guardrail that counts is not a sentence, it is a closed door.

The three levels that hold

  • Access: an assistant cannot reveal what it was never able to read. The strongest guardrail, and the easiest to verify.
  • Action: the list of things it can trigger is written, limited, and anything absent from it is forbidden by default.
  • Approval: binding actions go through a person. The system prepares, the human confirms.

All three can be verified. You can demonstrate that a document is out of scope, that an action is not declared, that approval is required. A style instruction cannot be demonstrated.

What a guardrail is not

It is not a banned-word filter, which blocks legitimate requests and lets rephrased ones through. Nor is it a sales promise: "our system never says anything silly" is not a guardrail, it is an argument.

And it is not an obstacle to usefulness. Well placed, guardrails make an assistant more useful, because they let you trust it with sensitive subjects knowing exactly how far it will go.

The link with confidentiality

The question "what can it say" is really the question "what can it read". An assistant with access to an entire workspace will eventually hand someone information they were not entitled to.

Which is why the access scope is decided at the same time as the content scope, not afterwards. Private AI in a company is built on that discipline more than on statements of intent.

What to test before go-live

Three trials reveal the essentials: a question out of scope, a request for an undeclared action, and a sideways question aiming at restricted information. If the system holds on those three, the rest is tuning.

Pro tip

Ask to see the list of sources the assistant can reach, not the list of its instructions. The first is verifiable, the second is declarative.

Frequently asked questions

Limits placed on an assistant to keep it inside its scope. The ones that hold are not written instructions but technical restrictions: what the assistant can access, and what it is allowed to trigger.

No. A written instruction is an intention: it holds in most cases and gives way in the ones that matter. An assistant cannot reveal what it never could read, and that is the strongest guardrail.

Three trials are enough: a question out of scope, a request for an undeclared action, and a sideways question aiming at restricted information. Also ask to see the list of reachable sources, which is verifiable, rather than the list of instructions, which is declarative.

At KERN-IT

We define what an assistant may read, say and keep, before building one.

Private AI for business: where your data goes

Got a project in mind?

Let's talk about how we can help you turn your ideas into reality.