Aarohii AI Solution

AI concepts explained

Ten terms that decide what an AI project costs and whether it survives contact with real users. Each section answers the question in the first sentence, then gives the mechanism, the failure mode, and the case where the honest answer is "don't use this."

This is the vocabulary we use in scoping calls. It is written for the person signing the invoice, not for a model-training leaderboard. Where a concept maps to work we ship, we say so and link to it; where it does not, we say that too.

What this guide covers

Short one-line definitions live in the AI glossary. Longer arguments live in Insights.

What is RAG?

Retrieval-Augmented Generation (RAG) is an AI architecture that retrieves relevant information from external data sources and provides that information to a language model before it generates an answer. The model is not expected to memorise your policy PDF. It is expected to use the passages you retrieved.

The mechanism is ordinary software, and that is the point. You ingest sources you have a right to use. You split them into chunks and tag them so retrieval matches the questions operators actually type, not the questions a demo scripts. You embed those chunks and store them. At query time you retrieve the top passages, put them in the model's context, and generate an answer only from those passages. Then you show the sources, so a human can check the answer without trusting it.

That last step is what separates a product from a toy. If the interface cannot show which passage produced a sentence, nobody can audit a wrong answer, and the feature quietly loses the team's trust within a month of launch.

RAG is the default when facts change faster than you can retrain, which covers most business software: policies, pricing, tickets, contracts, product documentation. It is also where the unglamorous decisions live. Chunk size, tagging, and access rights determine quality far more than the brand of database underneath.

When not to use it. If there is no corpus, or you do not hold the rights to the documents, RAG is not a bargain—it is a blocker, and the project needs a data conversation before an engineering one. If the task is not "answer from documents" but "classify this" or "collect these five fields," a classifier or a plain form is cheaper, faster, and easier to defend.

Two measurements keep a RAG build honest, and they must stay separate: did retrieval fetch the right passage, and did the answer stay inside what was fetched. A system can score well on the second while failing the first, which looks like success right up to the first audit. Reranking and hybrid search are worth adding when simple retrieval fails your eval set—not as decoration on an architecture diagram.

Related: RAG development · RAG vs fine-tuning

What is an LLM?

A large language model (LLM) is a machine-learning model trained on text so it can predict and generate language. Products use LLMs for chat, summarisation, extraction, and tool calling.

An LLM is not a database and not your company. It will invent plausible text if you let it, because generating fluent language is exactly what it was trained to do—and fluency is not a truth signal. Grounding through retrieval, evaluation on your own tasks, and a UI that shows uncertainty are how products stay honest.

We treat the model as a component. The product is the retrieval, the interface, the evaluation, and the operations around it. That framing changes which questions matter during procurement. Buyers usually open with "GPT or Claude or Llama?" The useful questions are narrower: what is the latency budget for this screen, where is data allowed to be processed, how reliable is tool calling for our call shapes, and what happens to the feature when the vendor has an outage. We write those answers into the plan, because each one constrains the architecture more than the model name does.

What we do not do. We do not train foundation models. Nothing about a mid-market product roadmap justifies that cost, and any vendor offering it should be asked what problem it solves that retrieval cannot.

Related: LLM development · Evaluating LLM vendors

What is an AI agent?

An AI agent is software that can decide on steps and call tools—search, tickets, email, APIs—to complete a task, typically using a language model to choose the next action. A chatbot answers in language. An agent may change state in another system. That difference is the whole reason approvals and traces matter.

A useful agent has parts you can list on one page. A planner or loop, usually a single model with a tool list. Tools scoped to least privilege, so the worst case is bounded by what the credentials allow rather than by the model's judgement. A required human approval before any irreversible action—refunds, sends, deletions, anything a customer sees. Short memory or none, because long-lived memory is a product decision with a privacy cost, not a free upgrade. And traces, so a failure can be replayed instead of argued about.

If you cannot name the tools, the policy, and the evaluation that proves the agent stayed inside both, you do not have an agent. You have an autocomplete with API keys.

When not to use one. If the process has a fixed sequence of steps, write the sequence. A deterministic workflow with one model call inside it is cheaper to run, easier to test, and far easier to explain to an auditor than a loop that re-decides its own plan on every run.

Related: AI agent development · Agent vs chatbot

What is generative AI?

Generative AI is machine learning that produces new content—text, images, audio, or code—rather than only classifying or scoring existing data. In most business software today, that means large language models working inside a product workflow.

Generative does not mean unsupervised. Production systems still need evaluation, access control, and a person accountable when the output causes harm. The interesting design work is usually about constraint: what the feature is allowed to write, which sources it may draw on, and where a human signs off before anything leaves the building.

The common failure is scope. A generative feature that can produce anything is hard to test, so it never gets tested, so it never gets trusted. A generative feature that produces one artefact—a draft reply grounded in the ticket history, a summary of one contract, a first-pass job description—can be evaluated, improved, and defended.

Related: Generative AI development · When not to add AI

What is AI automation?

AI automation is the use of machine-learning or generative models to perform or assist steps in a repeatable business process—classify, extract, draft, or route—while a human remains responsible for irreversible outcomes.

It is not the same as classic RPA, which replays deterministic clicks and keystrokes. It is also not a replacement for process design, and this is where most automation budgets are lost: if the process is undefined, a model will automate the confusion, faster and at higher volume than the humans managed.

The sequence that works is unexciting. Write down the process as it is actually performed, including the exceptions people handle by habit. Find the steps that are high-volume, low-judgement, and reversible. Automate those first, with the model's output shown to a person before it commits. Measure how often the human overrides it. Expand only where the override rate is low enough that review has become theatre.

When not to use it. If the step is rare, or irreversible, or requires judgement your team cannot articulate, leave it alone. Automating a monthly task to save an hour costs more in maintenance than it returns, and nobody thanks the vendor who automated a decision that later needed defending.

Related: AI automation services

What is fine-tuning?

Fine-tuning is additional training of a language model on your own examples so it better matches a style or task. It is not a substitute for up-to-date documents, and it is usually a later decision than retrieval.

The economics are what people miss. Fine-tuning is expensive to maintain when facts change, because every meaningful change to the underlying knowledge means another training run, another evaluation cycle, and another version to keep track of. Retrieval moves that cost into a document update.

Start with RAG unless three things are true: the task is stable, you have labelled examples in volume, and you can say specifically why retrieval cannot do the job. Style and format compliance are the honest wins—getting output reliably into a house tone or a rigid schema. Teaching the model facts is usually the wrong reason, and it is the reason most often given.

Related: RAG vs fine-tuning for SaaS

What is a vector database?

A vector database stores numeric representations of text—embeddings—so a system can find similar passages at query time. In RAG, that retrieval step is what grounds the language model.

An embedding is a list of numbers that places a chunk of text in a space where similar meanings sit close together. Retrieval is then a nearest-neighbour search: convert the question into the same space, return the closest chunks. That is the entire trick, and it explains both the strength and the limits. Similarity is not relevance. A passage that looks like the question can still be the wrong passage.

The brand of database is rarely the decision that makes a product honest. Chunking, access rights, evaluation, and showing sources matter more, and all four are chosen by your team rather than by a vendor. Choose the store that fits your operational reality—what your team can run, back up, and monitor—and spend the saved attention on retrieval quality.

Related: RAG development · RAG architecture

What is LLM evaluation?

LLM evaluation is a repeatable way to check whether a language-model feature still produces acceptable outputs on the tasks you actually run—using a labelled set, pass/fail rules, and a person who owns the result.

Public benchmarks are not your product. A support assistant needs tests built from your tickets. A RAG feature needs two tests, not one: that retrieval fetched the right passage, and that the answer stayed inside it. An extraction feature needs the fields that matter scored separately from the fields nobody reads.

A workable eval set starts smaller than teams expect—fifty to a hundred real cases, chosen to include the awkward ones, with the expected outcome written down by someone who does the job. The value is not statistical rigour. The value is that "the model got worse" becomes a claim you can check before a release instead of a rumour you argue about after one.

Without this, every prompt change is a guess, and every vendor upgrade is a gamble taken in production.

Related: AI testing · Production checklist

What is AI observability?

AI observability is the ability to see what an AI system did after it ran: the prompts or retrieved passages used, the tools called, latency and cost, and enough context for a human to replay a failure—without dumping secrets into a slide.

Without traces, "the model was wrong" is not a bug report. It is a complaint. With traces, you can change a prompt, a tool policy, or a retrieval chunk and prove the next release is better than the last one.

Cost belongs in the same view as correctness. Token spend per feature is the number that tells you whether a launch is sustainable, and it is usually discovered late, in a monthly bill, by someone who was not in the design review. Logging it per request from day one costs almost nothing and prevents that conversation.

Related: Engineering notes · Launch checklist

What is AI governance?

AI governance is how an organisation decides allowed uses, data rights, evaluation, human review, and shutdown for AI systems—before and after they go live.

The shutdown clause is the part usually missing. Someone must be able to turn a feature off, and the team must know who that is before an incident rather than during one.

Our consulting work includes a light governance pass: a named owner for each AI feature, a written policy for high-impact or irreversible actions, and clear ownership of the eval set. That is deliberately modest. We do not sell a certificate. For a mid-market product, three written pages that people actually follow beat a framework nobody reads.

Related: AI consulting · Security & trust

Where to go next

If you are scoping a build: Services. If you are deciding whether to build at all: Build vs buy and When not to add AI. If you need one-line definitions to hand to a colleague: AI glossary. If you want to see this applied in a shipped product: Products.

If the problem maps to work we actually ship, we will say so in 20 minutes.

Request a fit call