[UTC+3]

The Price of AI Hallucinations

Six numbers — and the cost of not trusting the model becomes visible: reviewer hours plus the price of the errors that slipped past review. Alongside it, the same flow is costed with document grounding switched on.

Your numbers

120 pcs

How many answers reach people or documents — drafts and debugging do not count.

8 %

Answers where a fact, a figure or a reference is not confirmed by the source.

3 min
120 000 RUB

Fully loaded cost of an employee — taxes and contributions included.

Checking answers costs per year
2 990 000 RUB

The sum of two parts: the paid hours of reviewers and the consequences of the errors that bypassed review.

Grounding pays back within a year

A typical range where tying answers to sources returns the investment in 6–12 months. The usual order of steps: first grounding on one class of questions with the highest error price, then expansion guided by the log.

Breakdown
Review hours per month
63 hours
Manual review per month
46 000 RUB
Errors travelling onward per month
116
Price of the errors that got through
203 000 RUB
Grounding: running cost
58 000 RUB
Grounded mode per month
151 000 RUB
Difference per month
98 000 RUB
Grounding payback
7 months

This is an order-of-magnitude estimate. The error rate and the price of a single error depend heavily on the type of tasks and the quality of the source documents — exact values come from a measurement on the actual flow.

How the calculation works

The chain starts with volume. Answers per day are multiplied by 21 working days — that gives the monthly flow. From there the flow splits into two cost items, and both are paid for with real money.

The first item shows up on the timesheet. Minutes per answer are multiplied by the number of answers reviewed and converted into hours. An hour is priced from the salary: the monthly amount divided by 164 working hours. The salary is taken fully loaded, with taxes and contributions.

The second item hides in other budgets. The error rate gives the number of wrong answers per month. Of those that reach review, a person catches around 85%. The rest go out to the recipient and surface later — in correspondence, in a return, in a document revision.

Not every miss turns into a loss. Some errors are caught by the recipient, some stay harmless. In the calculation 35% of misses reach consequences — a conservative reference point that is checked against a quarter of log data.

What moves the result the most

The strongest lever is the price of a single error. Between internal correspondence and a document with commitments the difference is twentyfold, and the yearly total shifts by an order of magnitude. This choice deserves an honest answer.

The second strongest lever is the share of answers reviewed. Full review catches more misses, but hours grow in proportion to the flow. Sampled review saves payroll and leaves errors unattended.

The error rate matters less than it seems. Moving from 8% to 12% raises the total by roughly one and a half times — noticeable, but no jump. Flow volume and error price outweigh it almost always.

When grounding pays off

Grounding against sources means the answer is assembled from retrieved document fragments and comes with a link to the paragraph. The reviewer only has to compare two texts side by side. In the calculation this cuts the error rate fourfold and makes it possible to move to spot checks on 15% of answers.

Running cost consists of two parts. Around RUB 9 per answer goes on database search and the longer context. Another RUB 35,000 a month covers the index, the storage and re-indexing documents when they change.

One-off investment is taken as RUB 700,000: connecting the sources, markup, acceptance on the live flow. Payback is obtained by dividing this amount by the monthly difference between the two modes. Values up to 12 months count as workable, up to 18 as acceptable when the error price is high.

The condition for success is simple. Documents must be current, one of a kind, and available to machine search. When a policy exists in three versions, grounding will honestly reveal the contradiction — resolving it is left to people.

How to measure the actual numbers

The error rate comes from a week-long slice. Fifty consecutive answers are taken, with no cherry-picking of convenient ones, and each is checked against the primary source. Only one thing is recorded: whether the fact is confirmed or not.

Minutes per check are measured by the clock, not by feel. Five answers in a row with a stopwatch give an average of sufficient accuracy. Experience shows a spread — a two-paragraph answer and a memo with references to a policy are checked in very different ways.

The price of an error is assembled from consequences, not from penalties. Rework, a repeat letter, a delayed deal, a manager's time spent sorting it out all count. Three recent cases are enough to get the order of magnitude.

What stays outside the brackets

The calculation gives an order of magnitude and a starting point for a conversation. Left outside the brackets are reputational losses, which cannot be measured in money, and the time needed to put documents in order before grounding is connected. That second item can be larger than the model bill, and it can only be assessed against a specific archive.

The reverse effect is not counted either. A source link next to an answer shortens review time — reading the whole thing is no longer necessary. In the formula this gain enters only through the sampling share, so the grounded total comes out on the cautious side.

FAQ

What is a model hallucination in plain words

A confident answer containing a fact that is not in the sources: an invented clause of a contract, a discount that does not exist, a plausible-looking figure. The tone is exactly the same as in a correct answer — it cannot be told apart by eye, it takes a check against the document.

Why do missed errors cost more in this calculation than the review itself

Review is paid by the hour and grows linearly, while errors cost in jumps. One wrong price in a letter to a client outweighs a week of control hours. That is why in the formula the price of misses usually outweighs the payroll of reviewers.

Where do the 85% and 35% come from

These are two conservative reference points. The first: a person catches roughly 85% of errors among the answers actually read — some slip through during monotonous scanning. The second: around 35% of misses turn into an actual loss, the rest get noticed by the recipient or stay harmless. Both are worth refining against a month of error-log data.

Why does grounding against sources cut the error rate fourfold

The answer is assembled from retrieved document fragments, and a link to the paragraph stays next to it. The model answers from the text in front of it, not from training memory. The 0.25 coefficient in the calculation is a cautious estimate: the result depends on how current and how consistent the documents are.

What to do if the error rate is unknown

Measure it over a week. Take 50 consecutive answers, check each against the primary source, mark it either "confirmed" or "not". The ratio of the second column to the total is the rate in question. As a bonus, it shows which questions produce errors more often than the rest.