Service · Aarohii AI Solution Private Limited
LLM development
We build the product around the model: retrieval, interface, evaluation, and operations. The model is a component you can swap. The product is what remains when you swap it.
What "LLM development" means here
It does not mean training a model. It means shipping a feature in your product that uses a hosted or open language model, and that still works in the third month — when the prompt has been edited eleven times, the vendor has deprecated a model version, and a customer has asked why the assistant said something odd on a Tuesday.
That is mostly ordinary engineering: a versioned prompt, a scoped tool list, a retrieval layer, a trace for every request, a cost number per feature, and an eval set that tells you whether last week's change helped. Background: what is an LLM.
The application layer we build
- Routing. Which model handles which call, and what the fallback is when it is slow or unavailable.
- Grounding. Retrieval where the answer must come from your documents, so the model is quoting rather than remembering. See RAG development.
- Structure. Schema-constrained output where another system consumes the result, because a downstream parser does not care how fluent the prose was.
- Interface. UI that shows sources, shows uncertainty, and makes the human's job review rather than proofreading.
- Operations. Traces, latency, and token spend per feature, visible from the first week rather than the first bill.
Choosing a model, usefully
Buyers usually open with "GPT or Claude or Llama?" It is the wrong first question, and answering it early locks in an architecture for a reason nobody can later reconstruct. The questions that actually constrain the build are narrower:
- What is the latency budget for this screen, measured at the 95th percentile rather than the average?
- Where is data allowed to be processed, and does that rule come from a contract or a preference?
- How reliable is tool calling for the specific call shapes this feature needs?
- What happens to the feature when the vendor is down — degrade, fail over, or stop?
- What does a thousand of these requests cost, and who watches that number?
We write those answers into the plan, then pick a model to fit them. Swapping models later should be a configuration change and a re-run of the eval set, not a rewrite. Longer argument: evaluating LLM vendors.
How we know a release is better than the last one
An eval set built from your own tasks, with pass/fail rules and a named owner. Fifty to a hundred real cases, chosen to include the awkward ones, is enough to make "the model got worse" a claim you can check before a release rather than a rumour you argue about after one. Detail: AI testing and the production checklist.
What we will not do
We do not train foundation models — nothing in a mid-market product roadmap justifies that cost. We do not treat fine-tuning as the starting point; it is optional and late, and the default comparison is written up in RAG vs fine-tuning. We do not ship a language-model feature with no trace, because the first support escalation turns that into an unanswerable question.
How an engagement runs
- 20-minute fit call — we say yes or no.
- Written plan — scope, success tests, model constraints, and what we will not do.
- Build with evaluation hooks and weekly demos you can inspect.
- Support after launch — we stay for the messy month, not just demo day.
Proof you can open without an NDA: Captverse, Auvora, PixellPeep, ViraQueue. Buying notes: what drives cost · build vs buy.
Questions
Do you train your own models?
No. We build the product layer around hosted and open models. Training a foundation model solves a problem mid-market software does not have.
Which model should we use?
Whichever one satisfies the latency, data-residency, tool-calling and failover constraints we write down first. We design so the choice can be revisited without a rewrite.
Can you work with open models on our own infrastructure?
Yes, where the operational cost is understood. Self-hosting moves spend from an API bill to your platform team; that is a trade, not a saving.
How do you control token cost?
Measure it per feature from the first week, cap context size deliberately, and cache what repeats. Cost surprises are almost always an unmeasured prompt growing quietly.
What if the vendor deprecates the model we launched on?
The eval set is re-run against the replacement and the differences are reported. This is routine when the tests exist and a crisis when they do not.
If the problem maps to work we actually ship, we will say so in 20 minutes.
Request a fit call