de_DE Deutsch |
Data centre visualising a multi-step AI retrieval routine

Agentic RAG: When AI Decides for Itself What to Look Up

Why a planning, multi-step retrieval routine answers complex questions with evidence – and where the extra effort pays off.

A clerk asks your AI system: “Which of our supplier contracts expire in the fourth quarter, and which of those contain an automatic renewal clause?” Classic RAG systems fall flat on a question like this. They search once, grab a few roughly matching passages and guess the rest. Agentic RAG works differently. It breaks the question apart, searches several times, checks its own intermediate results, and answers only once the evidence holds up.

Why classic RAG hits a wall

The first generation of RAG systems follows a rigid routine. A question arrives, the system turns it into a vector, pulls the five or ten most similar text chunks from the database, and hands them to the language model along with the question. One pass, one search, one answer. For simple questions this works surprisingly well.

The moment a question has several parts, though, the model buckles. Our contract question is really two searches: Which contracts expire in Q4? And which of those renew on their own? A single vector search blends both intents and returns a compromise that answers neither cleanly. The language model then fills the gaps with plausible-sounding inventions. These are exactly the cases that land on someone’s desk later as a complaint.

The problem isn’t the model, it’s the routine. A person wouldn’t reach into the shelf just once for a question like this either. They’d find the expiring contracts, set them aside, and then check each one individually for the renewal clause. Agentic RAG mirrors that behaviour.

What agentic RAG does differently

At its core sits a controller model that takes charge of the search itself. Instead of retrieving once and stopping, it plans. It splits the question into sub-steps, writes its own search queries, weighs the hits, and decides whether it knows enough or needs to look again. Only at the end does it write the answer.

Four building blocks make the difference. The planner divides complex questions into answerable bites. The retriever runs a separate search per bite, often across several sources – the contract database, the wiki, the ticket history. A grader checks whether the evidence actually supports the sub-question and throws out the weak hits. And a loop lets the system go back out when there are gaps, rather than answering too soon.

One point that matters for your infrastructure: this approach isn’t tied to any single vendor. You can run it entirely on open-weight models inside your own data centre. The controller doesn’t need to be a giant – a solidly fine-tuned model in the seven-to-fourteen-billion-parameter range plans reliably enough, as long as the tools and the data quality hold up.

Analyst reviewing contract clauses with sourced evidence on screen
Sourced answers instead of guesses: every statement points to a concrete passage. · AI-Designed

A worked example: the multi-part question

Let’s stay with the contracts. A classic system searches once for “contracts Q4 renewal” and hopes for the best. An agentic system moves in stages. First it searches specifically for contracts with an end date in the fourth quarter and finds, say, twenty-three matches. Then it checks those twenty-three one by one for the automatic-renewal passage – each with its own narrow search in the full text of the contract at hand.

What comes out is not a vague summary but a list: eight contracts, each with its number, end date and the exact clause quoted as evidence. Whoever reviews the answer clicks straight through to the source. For regulated industries that traceability isn’t a bonus, it’s a requirement.

The payoff shows up above all in comparisons and conditions. “Show me the differences between the March version and the current one” or “Which incidents last week involved the same assembly?” – questions like these demand several looks into different documents. That’s exactly where the agentic approach earns its keep.

Where companies feel the difference

In customer service, an agentic system answers even tangled requests that pull together contract status, product documentation and past tickets. The first-level agent gets a sourced answer instead of a hunch and escalates less often.

In legal and compliance work, every source counts. A team that has to check a hundred framework contracts for a specific liability clause saves days – and still keeps full control, because every statement points to a concrete passage. In manufacturing, the same mechanism helps when an incident report needs to be matched against maintenance history, manual and spare-parts catalogue before anyone picks up a wrench.

Finance benefits too. Questions like “How has our gross margin developed over the last four quarters, broken down by division?” require several retrievals from separate reports. A single vector hit can’t manage that. A planning system can.

Team analysing a knowledge graph built from multiple document sources
Complex questions need several looks into separate sources. · AI-Designed

The price: latency, cost, control

Let’s be honest about it: agentic RAG costs more. Several search passes and grading steps mean more compute per answer. Where a classic system finishes in two seconds, an agentic routine can take six or eight. For a chat where the user is waiting, that may be too long. For a thorough contract review that would otherwise eat half a day, it’s a bargain.

So the rule is: don’t make everything agentic. Route simple questions to the fast single retrieval and reserve the heavier routine for the complex cases. A small classification model at the entrance decides in milliseconds which path a request takes. That way you pay the higher cost only where it earns its return.

Running it in-house gives you real leverage here. You see every search pass, every intermediate step and every cost line. You can cap the loops, set a time budget and catch outliers before an agent burns through twenty searches. That control is what you hand away with a closed cloud API.

How to start without over-engineering

Don’t build the full orchestra on day one. Start with a clean classic RAG and measure which questions fail on it. A pattern usually shows up fast: the multi-part ones, the comparative ones, the conditional ones. Those are the first you hand to an agentic routine.

Then extend in small steps. First the decomposition of multi-part questions. Next the grader that filters out weak hits. Last the retrieval loop with a hard limit. Measure after each step against real questions from your operation, not demo examples. A system that shines in the lab and misses in daily use costs you trust – and that’s hard to win back.

Want to find out whether agentic RAG makes the difference for your documents and your operation? At AI-Designers we plan and run self-hosted AI systems inside your own data centre – from the first feasibility check to production use. Talk to us about your concrete use case.

Images: AI-Designed

Leave a Reply

Your email address will not be published. Required fields are marked *