Take Sales
Back to blogOperations

The knowledge base is your agent’s ceiling

· Updated on 7 min read

Teams who swap models hoping for better answers are usually disappointed. The model is rarely the bottleneck. The bottleneck is what it has to read.

The quality of an agent’s answers is capped by the quality of the knowledge base, not by the model. If the right passage does not exist, or is not found by the search, no model recovers it. Improving the base (coverage, freshness, absence of contradiction) returns more than any model swap.

What belongs in it

The practical rule is to start with what is already maintained for people, because those sources have an owner and age slowly.

  • The pricing and plans pages, because it is the most expensive question to get wrong.
  • The FAQ and support articles, already written as answers to real questions.
  • The product catalogue, with the attributes people actually decide on: delivery, size, compatibility.
  • Policies: exchanges, returns, cancellation, warranty.
  • Technical documentation, when your audience asks integration questions.

And what stays out, at least at first: the PDF exported once, the spreadsheet on someone’s desktop, the sales deck from two years ago. Not because they are useless, because nobody is responsible for updating them, and unowned content becomes a wrong answer with a date on it.

Contradiction costs more than a gap

A question with no answer in the base produces an “I do not know, let me check”: annoying, but honest. The same fact written two different ways in two places produces two plausible answers and no way to choose. RAG retrieves both, the model picks one, and you find out which when the customer complains.

Before growing the base, it is usually worth shrinking it: one pass removing duplicates and stale information improves retrieval more than doubling the number of documents.

Structure beats volume

Search does not read whole documents, it retrieves chunks. A document with a descriptive title, short sections and one idea per paragraph chunks well. A ten-page run of prose does not: the retrieved fragment arrives without the context that was on the previous page.

1

One idea per paragraph is the cheapest heuristic there is for improving retrieval. It is not sophisticated, and it works better than most things that are.

Measure coverage by the unanswered questions

The most useful part of operating an agent is the log of questions where it found nothing. It is a list, ordered by real demand, of the content missing from your site, and it serves the content team well beyond the agent. Each item resolved improves the agent, the SEO and the support queue at once.

Guardrails are not a substitute

A faithfulness guardrail stops the agent asserting something no passage supports. That prevents the error, but it does not produce the answer: the result is an “I do not know” instead of a wrong price. That is the correct behaviour, and it is also a reminder that the only way to turn that “I do not know” into an answer is to write the content.

#Knowledge base#RAG#Operations

Start the conversation today

Launch an AI agent that qualifies leads and books meetings around the clock. Live in under ten minutes.

A guided demo · your agent live in under 10 minutes