Knowledge base: where the assistant gets its answers
Definition
A knowledge base is the body of content an assistant relies on to answer. It is not a folder of files: it is a corpus that is dated, owned by someone and reviewed, otherwise the assistant will answer correctly from information that has become wrong.The misunderstanding at the start
The request often arrives in this form: "we have a shared folder with everything in it, just plug it in". It is technically possible and almost always a bad idea.
A shared folder holds three versions of the same document, drafts, forgotten exceptions and information nobody applies any more. An assistant plugged into that will not be wrong: it will answer faithfully from the outdated version. That is what makes the error hard to catch.
What makes a base usable
- A date on every piece of content, and the rule saying when it becomes suspect.
- An owner per domain, able to decide when two documents contradict each other.
- A written scope: what the assistant may read, and what it must never see.
- A revision path: how wrong information gets corrected, and how fast.
None of these four require technology. They require an organisational decision, which is why they are so often postponed.
Less is more
The temptation is to ingest everything to "cover all cases". Experience points the other way: a small, maintained base gives better answers than a broad, neglected one.
Thirty up-to-date documents beat three thousand where nobody knows which ones hold. Starting small also makes gaps visible immediately, because unanswered questions surface instead of drowning.
The knowledge that does not exist yet
An assistant project systematically reveals questions the organisation never answered in writing. It is uncomfortable, and it is the main benefit of the exercise.
The right mechanism treats each unanswered question as a proposed addition, which a competent person validates, dates and publishes. The base then grows through real use rather than anticipation, and every addition has already proven its worth.
The role of the source in the answer
A well-kept base lets the assistant show where its statements come from. That possibility changes the nature of the tool: from an answer taken on trust to an answer verifiable in one click.
It is also what makes content correctable. Without a visible source, nobody knows which document to revise when an answer is wrong. A custom assistant is judged as much on its ability to cite as on its ability to answer.
Before ingesting anything, ask each owner to name the document that holds in their domain. That single question often removes half the intended corpus.
Frequently asked questions
The body of content the assistant relies on to answer. Not a folder of files: a corpus that is dated, owned by someone and reviewed, otherwise the assistant will answer correctly from information that has become wrong.
Fewer than expected. Thirty up-to-date documents beat three thousand where nobody knows which ones hold. Starting small makes unanswered questions visible instead of drowning them.
Technically yes, and it is almost always a bad idea. A shared folder holds several versions of the same document and information nobody applies any more. The assistant will answer faithfully from the outdated version, which makes the error hard to catch.
Related terms
At KERN-IT
We build assistants that answer from your documents, cite their sources and know when to hand off.
→ Enterprise AI assistant: on your documents, in your tone