Built by Audacity OneAvailable on enquiry · In active development

OmniKB

A knowledge base that answers from your own documents, and shows its sources.

Our own build, and the argument we make to clients turned on ourselves: intake, retrieval, tenancy and deployment as one system instead of a model with a chat box in front of it.

Request a walkthrough
Intake
PDF, Word, text and CSV files up to 50MB. Text is extracted, split into overlapping passages, and embedded before anything is asked of it.
Retrieval
A question is embedded, matched against the passages in your own library, and answered from the ones that actually matched.
Isolation
Clients, documents and passages are scoped to one tenant and enforced in the database, not just in the interface.
Deployment
Run it on your own infrastructure, or have us host it and supply the model access. The choice changes where your documents sit, so it is made explicitly.

The answer exists. It is in four documents and nobody knows which.

OmniKB is built for teams carrying a lot of client material — contracts, scopes, notes, correspondence — who have lost track of what is actually in it.

The question is rarely find the file. It is what did we agree, when did it change, and where does it say so. That answer is usually spread across several documents and one person’s memory.

Every answer arrives with its receipts.

An answer you cannot check is a guess with better grammar. Pick a question and watch the passages come back with it — the ones actually retrieved, not a citation added afterwards.

Ask a question

Two documents matched

Answer

The signed scope fixes the first milestone at six weeks from kickoff with two review points. A later email moved the second review by three days under the notice clause; the milestone date itself did not change.

Source · Scope of work

…delivery of the first milestone is due within six weeks of the kickoff date, subject to two scheduled review points…

Source · Email thread

…pushing the second review to the 14th, three days later. Milestone stays where it is.

Illustrative. The shape of an answer, not a product screenshot.

What happens to a document.

Most document chat is a file picker wired to a model. The work that decides whether an answer can be trusted happens before the question is ever asked.

Why passages overlapA clause split down the middle answers nothing. Consecutive passages share their last two hundred tokens, so a sentence that straddles a boundary survives in both — shown below in the same run of text, read twice.
  1. Upload

    A file lands in storage under its tenant, a document record is created, and processing starts in the background.

  2. Extract

    Text is pulled out per format — PDF, Word, plain text and CSV each take a different path, with timeouts and failure states.

  3. Split

    Sentence-aware passages of roughly a thousand tokens, overlapping by two hundred, so an answer is never cut off mid-thought.

  4. Embed

    Passages are embedded in batches and stored alongside the document, which moves from pending to processing to done.

  5. Retrieve

    A question is embedded on the fly and ranked against the passages the asker is allowed to see.

  6. Answer

    The matched passages become the context for the answer, and the sources come back with it.

Why not just paste it into a chatbot.

A fair question, and for a one-off it is the right answer. This is about the hundredth question, over a library that keeps growing, when someone has to defend what the answer was based on.

How OmniKB compares with a general assistant, enterprise search, and building the system yourself.
CapabilityOmniKBA general assistantEnterprise searchBuild it yourself
Answers across several documents at onceYesYesNoYes
Cites the passages the answer came fromYesNoNoYes
Can run on infrastructure you controlYesNoYesYes
Per-tenant isolation enforced in the databaseYesNoYesYes
Your library persists between questionsYesNoYesYes
Someone else builds and operates itYesYesYesNo

Category comparison, not a product benchmark. Building it yourself gets every row — the question is whether that is the work you want.

Two ways to run it.

The difference is who holds the infrastructure and the model account. It changes where your documents sit and who gets woken up at three in the morning, so it is worth deciding early.

Self-hosted

You operate it

You own the infrastructure and the model account. We build it, deploy it into your environment, and hand it over.

  • Deployed into your cloud, under your own accounts
  • Your model provider keys, billed directly to you at cost
  • Documents, database and index never leave your infrastructure
  • Documentation and handover so your team can run it

Run by Audacity

We operate it

We host it and provide the model access. You upload documents and ask questions; the infrastructure, keys, upgrades and on-call are ours.

  • Hosting, monitoring and upgrades on our infrastructure
  • Model usage on our account, metered and passed through
  • Your documents processed by us under a written agreement
  • Evaluation and quality work continues after launch

Both are quoted per engagement rather than from a price list. Building and integrating it is scoped once; running it is ongoing — infrastructure and model usage either billed to you directly, or metered through us with an operating fee. Usage scales with the documents you index and the questions people ask, so the running cost tracks how much the thing is actually being used.

The questions we get asked.

Including the ones with awkward answers. You would find them out on the first call anyway.

That depends on which way you run it. Self-hosted, the storage, database and index are all on your infrastructure and never leave it. Run by us, they sit on ours under a written data agreement. Either way, passage text is sent to a model provider to be embedded and to generate answers, so it leaves your network at that step in both cases. If that is a problem for your data, say so on the first call — it changes the deployment, and it is better to find out before anyone signs anything.

Available on enquiry.

No sign-up, no trial, no self-serve. OmniKB is set up as part of an engagement — into your environment, or onto ours with the model access included. Which of those fits is usually clear within one conversation.

A walkthrough is a call. We show you the working system against real documents, work out which way of running it suits you, and tell you plainly whether it fits.

Where it stands

The document pipeline, retrieval, answers with citations, and tenant isolation work today. Search, analytics and parts of the admin surface are unfinished.

We would rather show you the working parts and be straight about the rest than demonstrate something you cannot run.