What people are trying to fix

  • “Nobody can find anything in our shared drive.”

    Search that understands the question, not only the words, and shows the passage it found.

  • “New staff ask the same people the same things.”

    An assistant that answers from your handbook and procedures, with the page cited.

  • “A chatbot made something up about our product.”

    Answers limited to your approved sources. When the source is missing, the system abstains and says what it could not find.

  • “Our documents are confidential.”

    A private deployment where indexing, search and the model run in your own environment.

  • “Our answers are out of date in three places.”

    Indexing that follows your documents, so a corrected page corrects the answers.

What we build

A knowledge system is more than a chatbot on top of a folder. These are the parts.

  1. Grounded question answering

    Answers written only from retrieved passages, each with a citation to the exact source.

  2. Abstain when the source is missing

    A check that every claim is supported by the retrieved text. If it is not, the answer is a refusal with a reason.

  3. Search over documents, tickets and data

    Hybrid retrieval that combines meaning-based and keyword search, then reranks the results.

  4. Ingestion and indexing

    Pipelines for PDFs, office files, wikis, tickets and databases, with updates and deletions handled.

  5. Access control

    People retrieve only what they are allowed to read.

  6. Evaluation and monitoring

    A question set from your real use, scored for correct answers, correct refusals and correct citations.

What you receive

  • The working assistant, in your environment, with a search interface or an API.
  • The ingestion pipeline and the schedule that keeps the index current.
  • An evaluation set of your questions, with scores and the failures listed.
  • A record of sources: what is indexed, what is not, and why.
  • Monitoring for unanswered questions, so you can see the gaps in your documents.
  • A handover session for whoever maintains the content.

Does it connect to what we already use?

Yes. The value is in reaching the places your knowledge already lives.

  • File shares and drives
  • Wikis and knowledge bases
  • Help desk and ticketing
  • Email archives
  • Databases
  • PDFs and scans
  • Intranet and chat tools
  • Identity, for permissions

Scanned documents need text extraction first, and the quality depends on the scan. We check a sample early so you know before we build.

Where it can run

The whole pipeline can run privately. The trade-off is model size, so we test on your questions first.

  1. On your devices

    Fits, with limits

    A single-user assistant over a folder of files, with a small model. Fine for personal or small-team use, not for a shared collection with permissions.

  2. On your servers

    Fits

    Embeddings, search, reranking and the answering model all run on your servers. Answer quality depends on model size.

  3. Private cloud

    Fits

    The same pipeline in your cloud account, with your keys and network rules.

  4. Hybrid

    Fits

    Retrieval stays local. Only the retrieved passages go to a cloud model, if you allow it, and the passages sent are logged.

  5. Cloud

    Fits

    The fastest to build. Documents and questions go to managed services under their terms.

Not every model or feature is available in every mode. We say which before you commit.

How it is built

For technical readers: what we choose from. The choice is made per system, from your data, your constraints and your workload.

Retrieval
Search across your documents with citations, combining keyword and vector search, and GraphRAG where the relationships between people, products and cases matter.
Documents
Parsing, OCR and structured extraction for PDFs, scans, spreadsheets and email.
Agents and tools
LangGraph and LangChain, tools through MCP, and an approval step before any action.
Evaluation
Test sets of real questions with known answers, rerun on every change, and answers that abstain when the sources do not support them.
Cost and speed
Caching, routing between smaller and larger models, and provider prompt caching where it applies.

The engine we run ourselves

We do not publish client work. ragX is the retrieval engine we operate for our own tools.

System we operate

ragX

Hybrid retrieval, reranking and a verifier that makes it abstain. It runs locally, with a local model for the critic step.

Internal and in daily use. Not offered as a hosted product; the same architecture is what we build for clients who need answers from their own documents.

Read the case study
  • Hybrid retrieval: dense vectors and keyword search fused by reciprocal rank fusion, then a reranker.
  • A critic pass on a local model, and a natural-language-inference verifier that must entail each claim or the answer abstains with a reason.
  • Indexing and evaluation jobs for the collections our tools use.

Questions we get

What is RAG?

Retrieval-augmented generation. The system first finds the relevant passages in your documents, then a model writes an answer from those passages only. The retrieval step is what ties the answer to your sources.

How do you stop it making things up?

Two ways. Answers are limited to retrieved passages, and a verification step checks that each claim is supported by them. If not, the system abstains. This lowers the rate of unsupported answers. It does not make errors impossible, so we measure it on your questions.

Can it run without sending our documents to a cloud?

Yes. Indexing, retrieval and the model can all run on your servers or in your private cloud. The trade-off is model size, which we test with you.

How many documents can it handle?

Size is rarely the limit. Document quality is. We test on a sample first and show you what the system does with your real files.

What happens when documents change?

The pipeline reindexes changed files on a schedule you set and removes deleted ones, so the answers follow your documents.

How do we know it works?

We build a question set from your real use and score answers, citations and refusals before launch. The same set runs after every change.

Bring us the documents nobody can search.

We start with the problem, then tell you whether it is worth building.

Bring us a bottleneck