[UTC+3]

Error Rate at Volume

Six numbers — and it becomes clear what stream of errors reaches the customer at the chosen review depth, and what the control itself costs.

Your numbers

6 000 pcs

Final replies to the customer are counted, not every call to the model.

4 %

Measured on a sample of real requests, not on invented examples.

40 %
2 min
90 000 RUB

Full cost of an employee — including taxes and contributions.

Wrong replies will reach the customer per month
158

Only the stream of errors is accounted for. The severity of each is counted separately — an error in a price and an error in a greeting cost different amounts.

Working mode with sampling

The flow is manageable, and there are two levers. Sampling depth reduces the number of errors that get through linearly, while analysing recurring cases reduces the share of wrong replies itself. The second is usually cheaper: rules and examples drawn from that analysis work across the whole volume at once.

Breakdown
Wrong replies in total
240
Replies under review
2 400
Removed by review
82
Share of wrong replies at the customer
2,6%
Review hours per month
80 hours
Reviewer headcount
0,5
Cost of review
44 000 RUB

The calculation gives an order of magnitude on a steady request flow. The real picture depends on how closely the test sample resembles live traffic.

How the calculation works

The starting point is volume. The number of replies per month is multiplied by the share of wrong ones measured in testing. That yields the stream of errors at the system's current quality.

Then the stream splits in two. Part of the replies goes under review, the rest go to the customer directly. In the unreviewed part, every error gets through without exception.

Inside the sample, not everything is caught either. Review depth is set by a coefficient: quick scan — 0.7, checklist — 0.85, cross-check against the source — 0.95. These are reference points for a rough estimate, not values measured on a specific project.

The result is how many wrong replies reached the customer over a month. Alongside it, the cost of control is calculated: review hours, headcount and their cost at the employee's rate.

What moves the result most

The first lever is the error share. It acts on the whole volume at once and without people. Dropping from 4% to 2% on a flow of six thousand replies removes 120 errors from the monthly stream, regardless of who reviews and how.

The second lever is sampling depth. Going from 40% to 80% with a checklist roughly halves the number of errors that get through. The price is hours: review doubles too.

The third lever is volume. It is a multiplier for both of the previous ones. A 2% share looks calm on a thousand replies and turns into a stream at forty thousand.

The order of the levers is not accidental. Sampling costs money every month, while work on reply quality pays off once and holds afterwards.

Under what conditions an assistant is released to people

The first condition is reversibility. An error is visible and can be fixed before it becomes an obligation: a wrong date is corrected by a manager, a wrong price is caught by a price-list check on the way out.

The second is a narrow topic. On a single scenario with clear sources, the error share falls fastest, because rules can be stated briefly and checked on a sample within a day.

The third is working escalation. The assistant must be able to say that it does not know and hand the conversation to a human. Declining to answer is cheaper than a confident invention.

The fourth is sampling on an ongoing basis. Even after reaching a low error share, a few percent of replies are looked at by eye: the request flow changes, and the picture of errors changes with it.

How to measure your own error share

Real requests from the past two weeks are taken — 150–200 of them, with no cherry-picking of convenient ones. The assistant answers them separately from customers. The replies are labelled by a specialist who knows the subject area.

Labelling goes into three buckets: correct, wrong, declined to answer. Declines are counted separately and do not go into the error share — they are costly in terms of convenience but safe in terms of consequences.

The resulting share is entered into the calculator. On a sample of a hundred requests the margin of error is around two percentage points, so the number should be read as a range.

What stays outside the brackets

In the calculation, errors are equal to one another; in life they are not. A mixed-up delivery date and a mixed-up warranty condition differ in cost by orders of magnitude. Estimating the damage requires analysis by type.

Errors come in clusters. One gap in the knowledge base produces a run of similar misses back to back, and a flat monthly average does not show such a run.

The lag of control is not accounted for either. Spot-checking is often done after the fact, while the reply goes to the customer immediately — in that case the number of errors that get through does not depend on review depth at all. For the calculation to be honest, review must sit before sending, at least on risky request types.

Finally, the work on the errors themselves stays outside the brackets. Analysing cases, correcting instructions and expanding the knowledge base take hours on top of review. Usually that is between a quarter and a half of the time spent on the sampling itself.

FAQ

Where to get the share of wrong replies if the assistant has not launched yet

A sample of 150–200 real requests from the past two weeks is collected, the assistant answers them offline, and a subject-matter specialist labels the replies. The error share on such a sample is a workable starting point. On a hundred requests the margin of error is about 2 percentage points, which is enough to estimate the order of magnitude.

Why does review not catch every error

A reviewer sees the reply but not always the source. A quick scan filters out obvious nonsense and lets plausible errors through — that is 0.7 in the calculation. A checklist with mandatory fields gives about 0.85, a cross-check against the source document about 0.95, but takes three times longer.

Which gives more — reviewing more deeply or reducing the error share itself

Reducing the error share acts on the whole volume and does not require people every month. Review acts only on the selected share and costs hours continuously. The comparison can be made right in the calculator: move the share of wrong replies down by one percentage point and look at the number that gets through.

Is the damage from errors counted here

No, only their quantity is counted. The cost depends on what exactly the system got wrong: a delivery date, a warranty condition or a price. Estimating the damage requires a separate analysis by error type.

Why is the reviewer's hour divided by 164

164 hours is an average working month on a forty-hour week. The salary including taxes is divided by that number and gives the cost of an hour. The headcount figure in the breakdown shows how many people would need to be kept on control at the chosen sampling depth.