[UTC+3]

Answer Spread and the Cost of Checking

Six numbers — and it becomes clear how many answers a month will need a second look and how many hours go into review. From there the decision is made on the sum, not on a hunch.

Your numbers

6 000 pcs

The entire flow: answers to clients, prompts for operators, draft documents.

25 %

The ones where an error costs money or reputation: prices, deadlines, contract terms.

6 min
1 500 RUB

Fully loaded hourly cost — taxes and overhead included.

Monthly cost of review
225 000 RUB

This counts people's time for full review of the critical share of the flow. Building automated checks and fixing the knowledge base are counted separately.

A working review routine

A typical cost for a mid-sized flow. At these numbers it pays off to move part of the checking into automation: matching numbers against the knowledge base, mandatory answer fields, refusing to answer without a source.

Breakdown
Will diverge on a repeat
540
With a made-up fact
240
Total in question
780
Share of the flow
13,0%
Go through review
1 500
Review hours
150 hours
Workload in full-time slots
0,9
Questionable and unchecked
585

This is an order-of-magnitude estimate. Divergence rates depend on the exact wording of requests and on the state of the knowledge base, so a precise figure only comes from measuring your own flow.

How the calculation works

The flow splits into two risk groups. The first is answers that will diverge from one another when the same question is repeated. The second is answers containing a fact that exists neither in the knowledge base nor in reality.

The divergence share starts at 9% and is multiplied by the format coefficient. A template with mandatory fields leaves the model little freedom and halves the share. Free-form wording raises it almost twofold: the same knowledge gets retold in different ways, and some of those retellings change the conclusion.

The fabrication share is counted from 4% and depends on what the answer rests on. When the text is assembled from knowledge base excerpts, fabrication has nowhere to appear — coefficient 0.35. When the model answers from general knowledge, the share grows 2.2 times.

Next comes the work of people. The critical share of the flow is checked in full, and the time is multiplied by the hourly cost. The result is the monthly cost of review in rubles and the reviewer's workload in full-time slots.

What affects the result most

Minutes per check. This is a multiplier across the whole volume: the difference between two minutes and eight changes the bill fourfold. Cutting them is easier than it seems — a templated answer reads faster than free text.

The share checked in full. Every next 10% of the flow costs as much as the previous one. That is why the zone of full review is described as a list of cases, not as a share picked out of thin air: first list where an error costs money, then look at what share that turns out to be.

The answer format and the source are the only two fields that move not the cost of checking but the volume of questionable answers itself. The "questionable and unchecked" line shows the other half of the picture: how many divergences reach the client without a second look.

Under what conditions review pays off

The first condition is measurability. Until a divergence log exists, any improvement remains a matter of faith. The log is started before the pilot: question, answer, reviewer's verdict, reason. After a month it becomes clear which topics produce inconsistency.

The second is the gradual handover of checks to code. Numbers are matched against the knowledge base automatically, mandatory fields are checked for presence, an answer without a reference to a source is flagged. A person handles what is left.

The third is that the zone of full review narrows based on evidence. Topics with not a single divergence over three months move to 10% sampling. The freed-up hours go to topics where divergences did occur.

When one figure is enough without a calculation

At a flow of up to a few hundred answers a month, all the checking fits into a couple of hours a week. There is nothing to calculate here — it is enough to assign someone on duty and keep a log.

The opposite end of the scale is also clear without a calculator. If an error in a single answer costs more than the annual review budget, the share checked in full equals 100%, and the conversation is about architecture: matching against the accounting system, refusing to answer on incomplete data, a mandatory second loop.

The in-between case is where the figure changes the decision. A flow of thousands of answers, a different cost of error by topic, limited expert time. The calculation shows the order of magnitude and suggests which of the six fields to move first.

What is left out of scope

Building the automated checks. Number matching, format control and refusal to answer without a source are written once and then work on their own; they are not part of the monthly cost of review.

Putting the knowledge base in order. Regulations that contradict each other produce divergences that neither format nor review can cure: the model faithfully retells what it found. That work is scoped by the state of the documents.

The cost of a missed error. The "questionable and unchecked" line gives a count but not a damage figure: it depends on who was told what. There is a separate tool for that part of the calculation.

FAQ

Why does AI answer differently every time

The answer is assembled one word at a time, and at each step one option is picked from several with close probabilities. Hence the different wordings. The danger is not different text but a different conclusion: when the same question about a delivery date gets Friday one time and Tuesday the next. A strict answer format and reliance on excerpts from the knowledge base reduce the share of such cases.

How to measure the spread on your own data

Take 50 typical questions and ask each one five times. Compare the answers by meaning: does the number, the date, the conclusion match. The share of questions where at least one run diverged from the rest is the spread. The measurement takes a day and gives a figure more accurate than any default coefficient.

Where do the 9% divergence and 4% fabrication come from

These are reference points for a flow where the instruction is given in words and the answer rests on the knowledge base plus the model's reasoning. The option coefficients shift them: a template with fields halves the spread, an answer from the model's own knowledge without sources multiplies the fabrication share by 2.2. Your own measurement overrides these reference points.

Does everything have to be checked

Full review is needed where an error costs money: prices, deadlines, contract terms, medical and legal wording. The rest is sampled at 5–10% with the trend watched. The "share checked in full" field shows exactly what each expansion of that zone costs.

What lowers the cost of review the most

The answer format and the source. An answer built on a template with mandatory fields can be checked by eye in a minute instead of six. An answer with a reference to a clause of the regulation is checked by matching, and part of that matching is done by code: the number in the answer is compared to the number in the knowledge base, and a mismatch goes to a person.