Service · Aarohii AI Solution Private Limited
RAG development
Retrieval-Augmented Generation (RAG) retrieves relevant passages from your sources and gives them to a language model before it answers. We build RAG when the product must cite, not invent.
When RAG is the right default
Policies, product docs, tickets, and contracts change. Training those facts into a model is slow and often wrong by the time it ships. RAG keeps the source of truth outside the weights, where your team can update it without a training run. See RAG vs fine-tuning for SaaS.
The pattern fits a specific shape of problem: a question arrives in language, the answer exists somewhere in documents you control, and someone needs to be able to check where the answer came from. Support assistants, internal search, policy Q&A, contract review, and operator copilots all have that shape.
It fits badly when the corpus does not exist yet, when you do not hold the rights to the documents, or when the real task is classification or form-filling dressed up as a chat feature. We will say so on the call rather than after the invoice.
The architecture we actually ship
There is no proprietary magic here, and any vendor claiming some should be asked to draw it.
- Ingest sources you have a documented right to use, with the access rules recorded per source rather than assumed.
- Chunk and tag so retrieval matches the questions operators actually type. This is where most quality is won or lost, and it is unglamorous work.
- Embed and store. We use whatever vector store fits your cloud and your team's ability to run it.
- Retrieve the top passages at query time, filtered by the same permissions the user has elsewhere in the product.
- Generate only with those passages in context, with the prompt versioned like any other code.
- Show sources in the interface, so a wrong answer can be audited instead of argued about.
- Measure retrieval hits and grounded answers as two separate numbers.
Reranking and hybrid search get added when simple retrieval fails the eval set — not as decoration on a diagram. Definition and background: what is RAG.
How we measure whether it works
Two tests, kept separate, because a system can pass the second while failing the first and still look like a success in a demo.
- Retrieval test: for a labelled set of real questions, did the pipeline fetch the passage that contains the answer?
- Grounding test: did the generated answer stay inside the passages retrieved, without adding plausible detail from nowhere?
The eval set starts at fifty to a hundred real questions, written down with expected outcomes by someone who does the job daily. It is deliberately small enough to build in the first week and useful enough that no prompt change ships without it. Ownership of that set transfers to your team at handover; a test suite only you can run is a dependency, not a deliverable.
What drives the cost
Three things, in order: the state of the corpus, the number of permission rules, and how much the interface has to explain itself. A clean set of documents with one access tier and a simple answer panel is a short project. Scanned PDFs, per-team visibility rules, and an interface that must show provenance inline are all real work.
Token spend per feature is logged from the first week, because the number that ends a project is usually discovered in a monthly bill by someone who was not in the design review. General buying notes: what drives cost.
What we will not do
We do not train foundation models. We do not promise zero hallucinations. We do not ship a retrieval feature without a way to show sources, because a feature nobody can audit is a feature nobody will trust by the second month. And we do not start a build when the answer to "do we have the rights to these documents?" is a shrug.
How an engagement runs
- 20-minute fit call — we say yes or no.
- Written plan — scope, success tests, and what we will not do.
- Build with evaluation hooks and weekly demos you can inspect.
- Support after launch — we stay for the messy month, not just demo day.
Open the proof without an NDA: PixellPeep, Captverse, Auvora, ViraQueue. Buying notes: how to choose a company.
Questions
Do you guarantee zero hallucinations?
No. RAG reduces made-up facts when retrieval is good and the UI shows sources. We measure the remaining errors and report them; we do not pretend they are zero.
Which vector database do you use?
Whatever fits the client's cloud and ops. The database is not the product. Retrieval quality, access rules, and evaluation are.
How much data do we need before RAG makes sense?
Enough that people currently answer questions by searching it. If the whole corpus is a handful of pages, put them in the prompt and skip the pipeline.
Can RAG respect our existing permissions?
Yes, and it must. Retrieval is filtered by the same rules that govern the documents elsewhere, so a user cannot reach through the assistant to read what they could not open directly.
What happens when the model vendor has an outage?
We write the answer into the plan before the build: degrade to search-only, fail over to a second provider, or show an honest error. Deciding during the outage is the expensive option.
If the problem maps to work we actually ship, we will say so in 20 minutes.
Request a fit call