OmniKB
A knowledge base that answers from your own documents, and shows its sources.
Our own build, and the argument we make to clients turned on ourselves: intake, retrieval, tenancy and deployment as one system instead of a model with a chat box in front of it.
Request a walkthrough- Intake
- PDF, Word, text and CSV files up to 50MB. Text is extracted, split into overlapping passages, and embedded before anything is asked of it.
- Retrieval
- A question is embedded, matched against the passages in your own library, and answered from the ones that actually matched.
- Isolation
- Clients, documents and passages are scoped to one tenant and enforced in the database, not just in the interface.
- Deployment
- Run it on your own infrastructure, or have us host it and supply the model access. The choice changes where your documents sit, so it is made explicitly.
The answer exists. It is in four documents and nobody knows which.
OmniKB is built for teams carrying a lot of client material — contracts, scopes, notes, correspondence — who have lost track of what is actually in it.
The question is rarely find the file. It is what did we agree, when did it change, and where does it say so. That answer is usually spread across several documents and one person’s memory.Every answer arrives with its receipts.
An answer you cannot check is a guess with better grammar. Pick a question and watch the passages come back with it — the ones actually retrieved, not a citation added afterwards.
Ask a question
Two documents matched
Answer
The signed scope fixes the first milestone at six weeks from kickoff with two review points. A later email moved the second review by three days under the notice clause; the milestone date itself did not change.
Source · Scope of work
…delivery of the first milestone is due within six weeks of the kickoff date, subject to two scheduled review points…
Source · Email thread
…pushing the second review to the 14th, three days later. Milestone stays where it is.
What happens to a document.
Most document chat is a file picker wired to a model. The work that decides whether an answer can be trusted happens before the question is ever asked.
Upload
A file lands in storage under its tenant, a document record is created, and processing starts in the background.
Extract
Text is pulled out per format — PDF, Word, plain text and CSV each take a different path, with timeouts and failure states.
Split
Sentence-aware passages of roughly a thousand tokens, overlapping by two hundred, so an answer is never cut off mid-thought.
Embed
Passages are embedded in batches and stored alongside the document, which moves from pending to processing to done.
Retrieve
A question is embedded on the fly and ranked against the passages the asker is allowed to see.
Answer
The matched passages become the context for the answer, and the sources come back with it.
Why not just paste it into a chatbot.
A fair question, and for a one-off it is the right answer. This is about the hundredth question, over a library that keeps growing, when someone has to defend what the answer was based on.
| Capability | OmniKB | A general assistant | Enterprise search | Build it yourself |
|---|---|---|---|---|
| Answers across several documents at once | Yes | Yes | No | Yes |
| Cites the passages the answer came from | Yes | No | No | Yes |
| Can run on infrastructure you control | Yes | No | Yes | Yes |
| Per-tenant isolation enforced in the database | Yes | No | Yes | Yes |
| Your library persists between questions | Yes | No | Yes | Yes |
| Someone else builds and operates it | Yes | Yes | Yes | No |
Category comparison, not a product benchmark. Building it yourself gets every row — the question is whether that is the work you want.
Two ways to run it.
The difference is who holds the infrastructure and the model account. It changes where your documents sit and who gets woken up at three in the morning, so it is worth deciding early.
Self-hosted
You operate it
You own the infrastructure and the model account. We build it, deploy it into your environment, and hand it over.
- Deployed into your cloud, under your own accounts
- Your model provider keys, billed directly to you at cost
- Documents, database and index never leave your infrastructure
- Documentation and handover so your team can run it
Run by Audacity
We operate it
We host it and provide the model access. You upload documents and ask questions; the infrastructure, keys, upgrades and on-call are ours.
- Hosting, monitoring and upgrades on our infrastructure
- Model usage on our account, metered and passed through
- Your documents processed by us under a written agreement
- Evaluation and quality work continues after launch
Both are quoted per engagement rather than from a price list. Building and integrating it is scoped once; running it is ongoing — infrastructure and model usage either billed to you directly, or metered through us with an operating fee. Usage scales with the documents you index and the questions people ask, so the running cost tracks how much the thing is actually being used.
The questions we get asked.
Including the ones with awkward answers. You would find them out on the first call anyway.
That depends on which way you run it. Self-hosted, the storage, database and index are all on your infrastructure and never leave it. Run by us, they sit on ours under a written data agreement. Either way, passage text is sent to a model provider to be embedded and to generate answers, so it leaves your network at that step in both cases. If that is a problem for your data, say so on the first call — it changes the deployment, and it is better to find out before anyone signs anything.
Available on enquiry.
No sign-up, no trial, no self-serve. OmniKB is set up as part of an engagement — into your environment, or onto ours with the model access included. Which of those fits is usually clear within one conversation.
A walkthrough is a call. We show you the working system against real documents, work out which way of running it suits you, and tell you plainly whether it fits.
Where it stands
The document pipeline, retrieval, answers with citations, and tenant isolation work today. Search, analytics and parts of the admin surface are unfinished.
We would rather show you the working parts and be straight about the rest than demonstrate something you cannot run.