AI & data products

Search that answers in sentences,
and always shows where the answer came from

Keyword search returns a list of links; a retrieval-augmented search answers the actual question and shows exactly which document it came from. We build the retrieval layer to be precise first, since a confident wrong answer is worse than a slow right one.

from$2,000
Timeline4 to 9 weeks depending on document volume and retrieval accuracy needed
What is includedRetrieval tuned against your real documents or catalogue, not a generic embedding setupAnswers in plain language with a citation to the exact source document or sectionA relevance check that declines to answer rather than guessing from weak matchesRe-indexing as documents are added, updated or retiredSearch delivered wherever your users already are: site, app, chat, or internal tool
4,096-dimsparse representation behind a two-level memory and retrieval engine we built
0.91routing accuracy on a model deciding which retrieval mode fits a given query
973pages indexed and served through a search-fed consumer AI product

What it is and who needs it

A RAG search product answers questions in plain language by retrieving relevant passages from your own documents or catalogue and generating an answer grounded in them, with a citation showing exactly where that answer came from. It fits any large document collection, product catalogue, or knowledge base where keyword search returns too many irrelevant results and a person wants an actual answer, not a list to sift through. It is not worth the complexity for a handful of pages a person can read directly; the value is in volume large enough that finding the right passage by hand takes real time.

What is inside

Documents are chunked and indexed with careful attention to chunk size and overlap, since retrieval quality depends heavily on getting this right for your specific content, not a one-size-fits-all default. Retrieved passages feed into the model to generate an answer grounded specifically in what was retrieved, with the source citation attached so a person can verify the claim directly rather than trusting it blindly. A relevance check on retrieved results catches the case where nothing in the index actually answers the question, declining to answer rather than generating a confident response from weak or unrelated matches. Re-indexing keeps the search current as documents change, so an outdated document never continues to surface as if it were still accurate.

How we build it

We start by understanding your actual document structure and content type, since chunking strategy differs significantly between, say, a legal document and a product catalogue, and getting this wrong is the most common cause of bad retrieval. We test retrieval quality against a set of real questions with known correct answers before connecting a generation model on top, since a retrieval problem disguised as a generation problem wastes time fixing the wrong layer. The relevance check gets tuned against deliberately irrelevant queries to confirm it declines appropriately rather than hallucinating from unrelated context. We launch against one document source, measure real answer quality, and expand once that foundation is solid.

What to watch

The real risk is a confident answer built on a weak or irrelevant retrieved passage, which is why the relevance check exists and gets tested against deliberately irrelevant queries before launch, not assumed to work from the retrieval architecture alone. Chunking strategy is the other place mistakes compound quietly, a chunk boundary that splits a critical sentence in half can make a document technically indexed but practically unretrievable for the exact question it should answer. Re-indexing discipline matters as documents change, since a RAG system is only as current as its last index update, and a stale index producing a plausible but outdated answer is worse than an obvious gap.

Timeline and price

Option Price What it covers Timeline
MVP from $2,000 One document source, citation-backed answers, relevance check 4 to 5 weeks
Production from $5,000 Multiple sources, access control, scheduled re-indexing, query logging 6 to 8 weeks
Full control (handover-ready) from $8,500 Everything in Production, plus a full handover package: architecture docs, test suite, admin access audit, and a walkthrough so your own team or another vendor can run it without us 8 to 9 weeks

Running cost on top of the build is usually $20 to $65 a month in vector-store and model costs, depending on document volume and query frequency.

What you own at the end

You own the document index, the retrieval pipeline, the citation logic and the full source code, running on your own infrastructure. The handover package documents the chunking and retrieval strategy, so adding new document sources later follows an established, tested pattern.

Pairs with the AI knowledge base assistant for staff-facing use and the AI support agent product for customer-facing use, both built on this same retrieval foundation. See the AI agents service page for the full range of agent and search builds. Real builds: the visa consulting centre support bots case study and the SENET memory engine case study, both built around retrieval that cites its sources. Have a document collection too big to search by keyword alone? Get in touch.

FAQ

How much does a RAG search product cost?

From $2,000 for retrieval over a single, reasonably organized document set with citation-backed answers. Multiple sources, access control and re-indexing on a schedule runs $5,000 to $8,500.

How long does it take?

Four to five weeks against a well-structured document set. Messy or very large document collections need more time upfront for chunking and indexing strategy, typically extending to seven to nine weeks.

What is the stack?

A vector store (pgvector on PostgreSQL, or a dedicated vector database for large corpora), Python for chunking and retrieval logic, and Claude or GPT for generating the final answer from retrieved context, with citations tied to source IDs.

Who owns the index and the retrieval logic?

You. The document index, the retrieval pipeline and the code run on your own infrastructure, with no dependency on a search-as-a-service subscription.

What happens when no document actually answers the question?

A relevance check on retrieved results catches weak matches and the system declines to answer confidently from them, rather than generating a plausible-sounding response from unrelated context.

Start here

Tell us the problem.
We bring the system.

A 30-minute call, a written plan with numbers within 48 hours, no obligation. If we are not the right fit, we will say so and point you to someone who is.