A RAG pipeline that answers from your documents
not from whatever the model happens to remember
An AI assistant that answers confidently from memory and an AI assistant that answers correctly from your own documents are very different products. We build the retrieval pipeline, chunking, embeddings, vector search, that grounds every answer in content you actually wrote, with a source it can point to.
What it is
Retrieval augmented generation, RAG, is the technique behind most AI assistants that answer accurately about your business specifically: your documents get split into chunks, each chunk gets converted into a vector embedding that captures its meaning, and those vectors live in a vector database. When a question comes in, the pipeline finds the chunks most similar in meaning to the question and hands them to the model along with the question, so the model answers from that retrieved text instead of from whatever it happened to learn during training.
When you need it (and when you do not)
You need this once an AI assistant has to answer questions about content specific to your business, your actual pricing, your policies, your product documentation, your internal knowledge base, content a general-purpose model was never trained on and cannot know. It is also the right build when accuracy and traceability matter: a support agent that can cite the exact policy document it answered from is trustworthy in a way a model’s confident-sounding guess is not.
You do not need a RAG pipeline for a general-purpose assistant with no proprietary knowledge to ground answers in, or for a narrow task with a small, fixed set of answers that fits comfortably in a prompt without retrieval at all. Building a vector database for twenty FAQ entries is more infrastructure than the problem needs.
How we build it
Chunking strategy is the part most RAG pipelines get wrong by defaulting to a fixed character count regardless of content structure; we chunk around your actual document structure, sections, headings, logical units, so a retrieved chunk is a coherent piece of meaning, not an arbitrary slice that cuts a sentence in half. We default to pgvector on PostgreSQL, which keeps vector search inside a database you likely already run, avoiding a separate system to operate, and move to a dedicated vector store like Qdrant only when scale or filtering needs genuinely call for it. Retrieval is tuned against a real evaluation set, a list of representative questions with known correct answers, rather than shipped on a default configuration and hoped to work; we adjust chunk count, similarity threshold and reranking until the evaluation set’s accuracy is actually good enough to launch. Every answer carries a citation back to its source chunk, both so a user can verify it and so we can see exactly where a wrong answer came from when one happens. We built retrieval pipelines on this pattern for an AI persona answering as a subject-matter expert and for a multi-channel sales agent that needs to answer from a real product catalogue, not an approximation of one.
What to watch
The most common RAG failure is not the model hallucinating, it is retrieval returning nothing relevant and the model answering anyway as if it had found something, we build an explicit “I don’t have information on that” fallback specifically to prevent the model from papering over a retrieval gap with a confident guess. Document freshness is the ongoing cost of ownership: a knowledge base that updates policies but never re-indexes the change will keep answering from the outdated version, so the refresh pipeline matters as much as the initial build. Lock-in is low, pgvector and open embedding models keep the pipeline portable across LLM providers, which matters because the model layer on top of retrieval is exactly where provider pricing and capability shift fastest.
Price and timeline
| Scope | Price | Timeline |
|---|---|---|
| Single knowledge base, pgvector | from $1,500 | 2 to 3 weeks |
| Multiple sources, reranking, evaluation set | from $4,000 | 4 to 5 weeks |
Related
Built as part of AI agents and custom development. Pairs with an LLM gateway with cost control and prompt and knowledge versioning to keep the whole AI layer maintainable. See it grounding an AI persona answering as a digital expert and a seven-channel AI sales agent. Tell us what documents your assistant needs to answer from: get in touch.
FAQ
How much does a RAG pipeline cost?
A setup covering a single knowledge base, chunking, embeddings, retrieval and citation, starts at $1,500. A fuller pipeline with reranking, multiple document types and an evaluation set runs $3,000 to $7,000.
How long does it take?
2 to 5 weeks. The pipeline itself goes together faster than the tuning: getting chunk size, retrieval count and similarity thresholds right against your actual documents takes real iteration, not a default configuration.
What vector database do you use?
pgvector on PostgreSQL for most setups, it avoids adding a separate database to operate and performs well at the scale most of our clients need. A dedicated vector store (Qdrant, Weaviate) makes sense at larger scale or with specific filtering needs.
Does this stop the AI from making things up?
It reduces it significantly by grounding answers in retrieved text rather than the model's trained knowledge, and citation lets a person verify the source, but no RAG setup eliminates the risk entirely. We build in explicit fallback behavior for when retrieval finds nothing relevant, rather than letting the model guess.
Who owns the knowledge base and the pipeline?
You do. The documents, the embeddings and the retrieval code live on your infrastructure, and the pipeline works with any LLM provider, nothing is locked to a single vendor's platform.