An internal AI knowledge system is an assistant that answers questions about how your company works, with a citation to the document, interview or recording that supports each answer, and that says when it does not know. Building one that people trust takes four things most projects skip: capturing the knowledge that is not written down, a review step between what the system extracts and what it may say, privacy rules that are visible to the people whose work it describes, and a training loop that turns approved knowledge into learning for new joiners. This is how TAI Labs builds one, and it is the system Prometheus, our company brain, is built on.
The three layers of company knowledge, and the one nobody captures
Most internal assistants read documents. Some also watch activity: tickets, chat, clicks. Both are crowded, well-funded categories, and both miss the layer that explains why the work is done the way it is. That layer only comes from asking people, and it is the one a new joiner needs most.
| Layer | Examples | How to capture it |
|---|---|---|
| Written | Handbooks, SOPs, product docs, wikis, drives | Read-only connections over approved sources, with the version recorded |
| Exhaust | Tickets, chat threads, emails, meeting recordings | Selectively, with consent, and never as the only evidence |
| Tacit | Why we do it this way, the workarounds, the risks, what people wish they could hand over | Discovery interviews, recorded programme sessions and check-ins, cut into cited moments |
How to build one, step by step
Seed it with discovery, not with a document dump
Interview the people who do the work: text, voice or video answers to a short set of questions about how the work really happens, where it hurts and what they would hand over. Record discovery calls. This is the tacit layer, and it is the reason the system will know things the wiki does not.
Extract claims, each linked to its evidence
The unit of the system is a claim (a workflow, a pain point, a process step, an outcome) linked to the exact span of a source that supports it: an interview excerpt, a transcript timestamp, a document section. No evidence, no claim. Models propose and label; code validates every citation and drops anything uncited.
Put a human review step between extracted and approved
Claims move through a lifecycle: raw, extracted, reviewed, approved, published. Only approved knowledge reaches the assistant for anyone outside the review team. Regeneration may replace extracted claims, never approved ones.
Make privacy a visible rule, not a setting
Every claim carries a visibility tier. Aggregates need a minimum group size. People appear by role, never by name, in anything the assistant says. Consent is captured at the moment of recording, and a 'who can see this' label sits on every input.
Answer with citations, or abstain
The assistant cites at the sentence, shows coverage ('based on 6 interviews and 1 call; finance not yet interviewed'), and says 'I don't know yet' in plain words. Every gap becomes a question it can ask the team.
Turn approved knowledge into learning
Approved claims, process cards and cited clips are assembled into short learning paths for new joiners and for people moving into a new workflow. This is the step that changes the economics: the knowledge system becomes the training system.
Keep it alive with ingestion, and expose it where people work
Recorded sessions, uploaded documents and, later, read-only connectors keep adding evidence. One agent with one memory is reached from web chat, voice, a weekly email brief and, over MCP, from inside Claude, ChatGPT, Copilot or Teams.
Why these systems die in security review, and how to avoid it
Internal assistants fail security review more often than they fail on model quality, and the reasons are predictable: broad read access, no record of who saw what, transcripts treated as trusted instructions, and an agent that can read private data and take an outbound action in the same run. Each of these has a design answer, and they have to be in the design from the start rather than retrofitted.
- Every read goes through a server-side organisation and role gate, with a negative test that tries to read another organisation's data
- Transcripts, answers, documents and connector content are treated as untrusted input in every prompt
- No autonomous run combines private reads, untrusted content and an outbound action without a person confirming
- One audit row per question answered and per tool call, with undo for anything the agent changed
- Connectors are read-only, allow-listed by an admin, with permission re-checked at read time
- No customer data in fixtures, demos or screenshots; a fictional sample company for those
The tools
- NotionThe written layer: SOPs, handbooks, the review queue
- Google DriveApproved source documents, versioned
- SlackWhere the assistant is asked, and where check-ins arrive
- Custom MCPsThe assistant inside Claude, ChatGPT, Copilot or Teams; and read-only connectors
- ClaudeExtraction, labelling and cited answers
- ElevenLabsVoice check-ins and interview answers
- LangGraphThe extraction and review pipeline, with evals
- GeminiTranscription of recorded interviews and sessions
Do it with us once
We build the system with you, once, and your team runs it.
Discovery interviews, the extraction and review pipeline, the privacy rules, the learning paths and the channels, with your programme as the source. Bring one team to a 15-minute call and we will scope it.
Honest empty states, and why they build trust
A knowledge system will be a fifth full for a long time. The temptation is to hide that behind confident answers, and it is the fastest way to lose the team. The alternative is to show coverage on every answer, say what has not been interviewed yet, and let the gap become a question the assistant asks the right person. Teams trust a system that admits what it does not know and gets better every week, and distrust one that answers everything on day one.
How the knowledge system becomes the training system
The reason to build the system this way, rather than as another search box, is what happens after the programme. Every workflow a team learned live is captured as a process card and a set of cited clips. Every new joiner is onboarded from those, in the company's own words, with the reasons attached. Every check-in adds evidence about whether the new way stuck. The training was done once; the system keeps teaching. Bersin's 2026 research describes the shift in corporate learning as moving from training to dynamically sharing information, and this is what that looks like in practice.
- Process cards: 'How this company does X', with the clip that proves it
- Learning paths of three to six lessons, assembled from approved cards and clips, assigned on publish
- Weekly check-ins that record whether the new workflow is being used, and what got in the way
- Baselines and outcomes per workflow, so the renewal conversation has numbers in it
Frequently asked questions
Is this a RAG chatbot over our documents?
Retrieval over documents is one layer of it. The difference is the tacit layer captured through interviews and recorded sessions, the review lifecycle between extraction and publication, the visible privacy tiers, and the learning paths built from approved knowledge.
What data does it need to start?
Discovery interviews with the people who do the work, and the handful of documents they point to. It does not need every drive connected on day one, and it works better if it is not.
How do we stop it exposing something a colleague said in confidence?
Consent at the moment of recording, visibility tiers on every claim, aggregates only above a minimum group size, and people described by role rather than name in anything the assistant says. Answers given under an anonymous framing are never re-identified.
Can it act on our systems?
It drafts freely and asks before acting. Messaging people, assigning learning or changing anything client-visible needs a person's approval, is logged and can be undone.
How does this relate to custom AI training?
The programme is the source: discovery seeds the system, the live sessions are recorded into it, the workflows the team builds become its process cards. The custom AI training page describes the programme; Prometheus is the system that keeps it alive.
Who owns it afterwards?
You do. One named organisation admin can see the sensitive views; your champions maintain the approved knowledge; the learning paths are yours to assign. We build it with you once.
Capture the knowledge people carry in their heads through discovery, link every claim to its evidence, review before you publish, make privacy visible, cite or abstain, and turn what is approved into learning. That is how TAI Labs builds an internal AI knowledge system, and it is what Prometheus does for the teams we train.