The 90-Day Review: How to Tell If Your AI Receptionist Is Actually Working
Most shops decide whether an AI receptionist is working based on a feeling, usually driven by the one call that went badly. That is not a review. Here is a concrete three-checkpoint protocol — coverage at day 30, quality at day 60, money at day 90 — including the attribution problem nobody wants to state plainly and the specific fixes for each way the numbers can come out bad.

The 90-Day Review: How to Tell If Your AI Receptionist Is Actually Working
Three months after switching the phone over, most owners still answer "is it working?" with a feeling. And the feeling is almost always driven by the worst single call — the one where a customer said something unusual, the bot handled it awkwardly, and somebody forwarded you the recording. That one call becomes the verdict, while ninety days of calls that got answered at 11 PM and turned into jobs never enter the evaluation at all.
As of August 2026 there is no excuse for that, because every meaningful number is already sitting in the call log. What is missing is a protocol: a fixed set of questions asked at fixed points, in a fixed order, so that the answer at the end is a decision rather than a mood.
This is that protocol. Three checkpoints — day 30, day 60, day 90 — each asking one category of question. Coverage first, because if the calls are not being answered nothing else matters. Quality second, because a call answered badly is worse than one answered well and better than one not answered at all. Money last, because it is the hardest to attribute honestly and you need the first two to interpret it.
Before day one: capture the baseline, or the whole exercise is guesswork
The single most valuable thirty minutes in this process happens before the AI answers a single call, and almost nobody spends it.
Write down what your phone was doing beforehand. Pull the last 60 to 90 days from your carrier or call-tracking records and record five numbers:
- Total inbound calls in the period.
- How many were answered by a human, by any measure your system gives you.
- How many rang out, hit voicemail, or got a busy signal.
- How many arrived outside your stated business hours.
- Roughly how many turned into jobs — even if this is your best reconstruction from the schedule rather than a clean number.
If you missed this window and are already live, you can still partially reconstruct it. Carrier call detail records usually go back far enough, and they will at least give you total inbound and answered-versus-not. Do it now rather than at day 90, when the temptation to reconstruct a flattering baseline is strongest.
Why this matters more than any other step: at day 90 you will want to say "the AI generated X." You cannot say that. What you can say is "the shop's answered-call count went from A to B and booked jobs went from C to D over comparable periods." That is a real, defensible statement. Without the baseline, you have the second number and nothing to compare it to, and you are back to a feeling — just a more expensive one.
The seven numbers worth tracking on an ongoing basis, and what each one actually tells you about pricing and staffing, are laid out in the locksmith call log. Set that up in week one and the checkpoints below become a matter of reading rather than gathering.
Day 30 — coverage: is the phone actually being answered?
Thirty days is too early to judge quality and far too early to judge money. It is exactly right for coverage, because coverage problems show up immediately and are usually configuration errors you want to catch in week five rather than month three.
Four questions, in order.
1. What is the answer rate?
Answered calls divided by total inbound calls. This should be very close to complete. The specific thing to look for is not the headline percentage but the shape of the misses: are they clustered at a particular hour, on a particular day, or on a particular inbound number? Clustering points at a configuration problem. Scattered singles usually point at callers who hung up in the first two rings, which is a different thing entirely and not a system failure.
2. How many calls came in outside business hours?
This is the number that most often surprises owners, and it is the clearest single justification for the whole exercise. Count the calls that landed between your closing time and your opening time, plus weekends and any holiday in the window.
Whatever that count is, it is the population that previously reached voicemail. Not all of them were jobs — some were price shoppers, some were wrong numbers, some were the same person calling four times. But you now have the actual number, in your own shop, for the first time. If you want to model what that population is worth before you have booking data, the missed-call cost calculator will do the arithmetic with your own inputs.
3. How many abandoned or no-answer calls remain?
There should be very few. The ones that exist fall into two groups and you must separate them:
- Caller hung up in the first few seconds. Normal. Misdials, people who changed their mind, robocallers. Not actionable.
- The call reached the system and the caller dropped mid-conversation. Actionable. If this is happening repeatedly, the usual culprits are a greeting that is too long, an intake that asks too many questions before acknowledging the emergency, or an audio problem. The greeting is the first thing to shorten.
4. Did concurrent calls collide?
Look for the minutes where two or more calls arrived simultaneously and confirm all of them were answered. This is the specific structural advantage over a human answering serially, and it is worth verifying rather than assuming — the mechanics of why it matters are in the busy-signal problem.
What "good" looks like at day 30: near-total answer rate, misses that are scattered rather than clustered, a concrete after-hours call count you did not previously have, and zero concurrency collisions. If all four hold, stop looking at coverage and move on. If any one fails, fix it now — it is almost always hours configuration, a forwarding rule, or a greeting that is too long, and all three are ten-minute changes.
Day 60 — quality: is it answering well?
Sixty days gives you enough call volume for a real sample. This checkpoint is the one owners skip, because it takes actual time — you have to listen — and because they are afraid of what they will hear.
Listen anyway, and listen to the right sample.
Sample properly, not emotionally
The mistake is listening only to the complaints. Complaints are a biased sample by construction: they over-represent the unusual call and tell you nothing about the ordinary ones, which are the overwhelming majority of your volume and your revenue.
Pull a stratified sample instead — roughly 20 to 30 calls, deliberately spread across:
- Calls that booked (did it close well, or did the customer book despite the call?)
- Calls that did not book (was it a genuine price shopper, or did we lose one?)
- Calls that were escalated or transferred to a human
- Calls after 9 PM — where behavior differs and nobody is watching
- Spanish-language calls, if you take them
- The longest calls in the period, which are where confusion lives
- Two or three of the shortest answered calls, which is where premature endings hide
Block ninety minutes. Listen at normal speed to at least ten of them and skim the rest. The general discipline of running recording review as a repeatable process — including the consent considerations — is covered in call recording and quality review.
Score four things, not everything
Do not build a rubric with fourteen dimensions. Four things matter, and each maps to a specific fix.
1. Quote accuracy against the price book. For every call where a price was stated, check it against your actual price sheet. This is binary and it is the highest-stakes item on the list, because a wrong quote either loses the job or creates a dispute at the door. If quotes are wrong, the price book is nearly always stale rather than the system being creative — the discipline for keeping it current is in the price book for AI phone quoting. Note that a message-taking configuration does not quote at all, so this item is only in scope if you are on a plan that quotes.
2. Intake completeness. For each call, did you get: name, callback number, service needed, location, vehicle or hardware detail, and urgency? Score it out of six. Anything that is systematically missing is a configuration gap, not a judgment failure — if the location is missing on a third of calls, the intake is not asking hard enough for it.
3. Escalation appropriateness. Two failure directions, and they need different fixes:
- Too tight — everything gets escalated to a human, which means you bought an expensive call-forwarder. Symptom: transfers on routine, easily-handled calls.
- Too loose — calls that genuinely needed a person got handled by the bot. Symptom: an angry customer, a warranty complaint, or a legal-sounding question that got a scripted answer.
What should always escalate is worth deciding explicitly; there is a defensible list in what an AI receptionist cannot do.
4. Handoffs that failed. This is the highest-value item in the entire review and it is invisible unless you look for it deliberately. Find every call where a transfer or callback was promised, then check whether it happened. A promise made on the call and not kept by the shop is worse than never making the promise — and it is almost always a shop-side ops failure rather than a software one. The transfer mechanics and their failure modes are in warm versus blind transfers.
What "good" looks like at day 60: quotes match the price book; intake is complete on the large majority of calls; escalations happen on the calls that warrant them and not on routine ones; and every promised callback was actually made. Write down every defect you find with the call it came from — you will need the list for day 90.
Day 90 — money: what did it produce?
Now, and only now, the financial question. Three metrics, and one honest caveat you should read before the metrics.
The attribution problem, stated plainly
You cannot know what a missed call would have been worth. That is not a limitation of any particular product; it is a fact about counterfactuals. The person who reached your voicemail at 10 PM in March and never called back left no record of whether they would have booked, what vehicle they had, or what they would have paid.
Any calculation that assigns a specific dollar value to calls you did not answer is a model, not a measurement. Models are useful — they are how you decide whether to spend money before you have data — but they are not evidence, and you should not present a modeled number to yourself as a result.
What you can measure honestly is the delta against your own baseline. Answered calls before versus after. Booked jobs before versus after. Revenue in comparable periods.
And you must be equally honest about what that comparison cannot prove:
- It is not a controlled experiment. Your ad spend, your seasonality, your reviews, your competitors, and the weather all moved during the same 90 days.
- Seasonality is real and large in this trade. Comparing August to May is comparing two different demand environments. Compare against the same period last year where you can.
- Some of the lift is a change in your own behavior. Having structured call records tends to make owners follow up more, and follow-up produces jobs. That is a real gain, but it is not attributable to the answering layer alone.
State the delta, state the confounders, and make the decision anyway. That is a more rigorous position than either "it definitely made me $40k" or "I don't really know."
The three numbers
1. Booked jobs traceable to answered calls. Match your booked jobs in the period against the call records by phone number. Some will not match — walk-ins, repeat customers texting a tech directly, jobs from a partner. That is fine. You want the count that does match, because it is a floor rather than an estimate.
2. Cost per answered call. Total monthly cost divided by answered calls. On Lite at $149 a month with 100 calls included, a shop using its full allotment is at roughly $1.49 per answered call before any per-minute overage; usage above the included calls adds 50 cents a minute. On Core at $500 a month for 500 AI minutes, the per-call figure depends entirely on your average call length — which is why you should compute it from your own log rather than a table. Whatever it comes to, compare it against the alternative you were considering: a part-time person, an answering service's per-minute rate, or voicemail costing you nothing and producing nothing.
3. Cost per booked job. Total monthly cost divided by traceable booked jobs. This is the number that actually decides the question, because it is directly comparable to your customer acquisition cost from every other channel. If your cost per booked job through the phone layer is below what you pay per booked job through paid search, the phone layer is your cheapest acquisition channel and the conversation is over. Conversion benchmarking context — how many answered calls should be turning into jobs in the first place — is in call-to-booking conversion rates.
4. After-hours revenue specifically. Isolate the jobs that came from calls received outside your business hours. This is the cleanest segment in the whole analysis, because before the switch that population reached voicemail almost by definition. It is the closest thing to a clean before-and-after you will get.
Here is the checkpoint structure in one view:
| Checkpoint | Question it answers | What to measure | What "bad" looks like | First fix |
|---|---|---|---|---|
| Day 30 — Coverage | Is the phone being answered at all? | Answer rate, after-hours call count, abandoned calls, concurrency collisions | Misses clustered at one hour or one number; mid-conversation drops | Hours configuration, forwarding rules, shorten the greeting |
| Day 60 — Quality | Is it answering well? | Stratified sample of 20-30 recordings; quote accuracy, intake completeness, escalation fit, failed handoffs | Wrong quotes; missing location or callback; everything escalating; promised callbacks never made | Refresh the price book; tighten intake prompts; retune escalation rules; assign an owner to callbacks |
| Day 90 — Money | What did it produce? | Traceable booked jobs, cost per answered call, cost per booked job, after-hours revenue | Cost per booked job above your paid-search CAC; no measurable booking lift | Fix day-60 defects first, then re-measure; reconsider plan tier |
| Ongoing — Plan fit | Am I on the right tier? | Minutes used vs included; overage as a share of the bill | Paying more in overage than the next tier's step-up | Move a tier up, or trim call length |
What "bad" looks like, and what to actually do about it
Four failure patterns cover almost everything. Each has a specific fix, and in every case the fix should be tried before the platform is blamed.
Pattern 1: quotes are wrong. Almost always a stale price book rather than an invented number. The bot states what it was given; if what it was given is six months old, it will state a six-month-old price confidently. Fix: re-upload the current price sheet, confirm it on screen, and spot-check five vehicles or five hardware types against the live behavior. Put a recurring calendar item on it — quarterly at minimum, and immediately after any price change.
Pattern 2: escalation rules are wrong in one direction or the other.
- Too tight (everything transfers): you are paying for an answering layer and getting a forwarding service. Identify which call types are transferring unnecessarily and let the bot handle them.
- Too loose (things that should have reached a person did not): usually complaints, warranty callbacks, anything sounding legal, and existing-customer disputes. Add those categories explicitly.
The tell for the first is a high transfer rate on routine calls; the tell for the second is an angry customer who says nobody would talk to them.
Pattern 3: the greeting is too long. This one is underrated and easy. Every extra second before the caller can state their problem raises the chance they hang up — and an emergency caller is the least patient person who will call you all week. If mid-conversation drops cluster in the first fifteen seconds, cut the greeting to a sentence. The first-90-seconds discipline is broken down in the call handling checklist.
Pattern 4: the hours are configured wrong. Genuinely common and genuinely embarrassing. A shop sets 8-to-5 during onboarding, actually runs emergency work until 9 PM, and then wonders why the after-hours behavior seems off. Check the configured hours against how you actually operate, including weekends and how you want holidays handled.
And the ambiguous outcome, which is the most common one of all: coverage improved, quality is fine, and the money is unclear because too many things moved at once. The correct response is not to cancel and not to declare victory. It is to (a) fix every day-60 defect, (b) run one more clean 30-day period with no other changes, and (c) measure again. One clean month with a fixed configuration is worth more than three noisy ones.
Upgrading or downgrading: the plan decision
The review should end in a plan decision, and it is mostly arithmetic.
Look at two things: minutes used against minutes included, and overage as a share of your bill.
- Consistently paying meaningful overage? Do the arithmetic on the next tier up. Core is $500 a month for 500 AI minutes with 45 cents a minute after; Pro is $750 for 1,000 minutes at 40 cents after; Elite is $1,200 for 2,500 minutes at 35 cents after. There is a crossover point where the higher tier is cheaper than your current tier plus overage, and it is worth computing rather than estimating.
- Well under your included minutes, month after month? You may be over-tiered. Look at whether the features you are paying for — live quoting, booking on the call, dispatch — are actually being used. If the honest answer is that you mostly need messages taken reliably, KeyBot Lite at $149 a month does exactly that: message-taking only, no quoting and no booking, 100 calls included, 50 cents a minute after, delivered by Telegram. The specific decision framework for moving between them is in Lite versus Core: when to upgrade, and a worked hypothetical for a small mobile operation is in the two-truck ROI walkthrough.
- Call length trending up? Before you buy minutes, find out why. Long calls are usually a symptom — a confusing intake, a missing price entry causing back-and-forth, or an escalation rule that makes the bot work too hard before handing off. Fixing the cause is cheaper than buying more minutes.
All current tiers are on pricing, every Core, Pro and Elite plan includes a 14-day free trial, and there are no per-seat fees at any tier. If you want to hear the current behavior rather than reason about it, have it call you as your own shop — it rings back in about 30 seconds and the first demo is free.
The one-page version to keep on the wall
Day 0. Write down: inbound calls, answered, missed, after-hours, booked. Five numbers. Thirty minutes.
Day 30. Answer rate. After-hours count. Abandoned calls, split into early-hangups and mid-conversation drops. Concurrency collisions. Fix any clustering immediately.
Day 60. Ninety minutes with 20-30 stratified recordings. Score quote accuracy, intake completeness, escalation fit, and failed handoffs. Write every defect down with its call.
Day 90. Traceable booked jobs. Cost per answered call. Cost per booked job. After-hours revenue. Compare to day 0 and say out loud what the comparison can and cannot prove.
Then decide. Fix the defects, choose the tier, and put the next review on the calendar for six months out. This is not a one-time exercise — a price change, a new service line, or a seasonal shift all merit a fresh look at the quality sample.
The bottom line
The reason most shops cannot tell whether their AI receptionist is working is that they never wrote down what the phone was doing before, and never looked at anything except the one call that went badly. Fix both. Take the baseline before you switch, check coverage at 30 days because coverage failures are configuration errors you want to catch in week five, listen to a properly stratified sample at 60 days because quality defects are invisible from a complaint list, and only at 90 days go to the money — where the honest framing is a delta against your own baseline with the confounders named out loud, not a fabricated value for calls you never answered. When the numbers look bad, the fix is nearly always a stale price book, an escalation rule tuned wrong in one direction, a greeting that is too long, or hours that do not match how you actually operate — all of them ten-minute changes. And when they look ambiguous, run one more clean month rather than deciding on noise. The whole protocol costs about three hours spread over three months, which is less than the cost of guessing wrong in either direction.
Frequently asked questions
How long before I can tell if an AI receptionist is working?
Coverage tells you within 30 days, quality within 60, and money honestly needs the full 90. Coverage is immediate because answer rate and after-hours call counts show up in the first week and any failure there is a configuration error worth catching early. Money takes longest because booked jobs lag the call that produced them, seasonality moves underneath you, and you need enough volume for the comparison against your baseline to mean anything at all.
What metrics actually matter for an AI receptionist?
Four matter and the rest are decoration: answer rate including concurrent calls, intake completeness on the calls that were answered, cost per booked job, and after-hours revenue that previously went to voicemail. Answer rate proves the coverage, intake completeness proves the call was worth answering, cost per booked job makes the phone comparable to every other acquisition channel you buy, and after-hours revenue is the cleanest before-and-after segment available since that population previously reached voicemail by definition.
Can I prove how much revenue an AI receptionist generated?
No, and you should be suspicious of anyone who says otherwise, because you cannot know what a call you never answered would have been worth. What you can measure is the delta between your own pre-switch baseline and the current period — answered calls, booked jobs traceable by phone number, and revenue in comparable windows — while naming the confounders out loud: ad spend, seasonality, reviews, and your own changed follow-up behavior all moved at the same time. State the delta, state the caveats, and make the decision on that basis.
My AI receptionist quoted the wrong price. What went wrong?
Almost always a stale price book rather than an invented number, because the bot states the prices it was given and will state a six-month-old price with complete confidence. Re-upload your current price sheet, confirm it on screen, and spot-check five vehicles or hardware types against the live behavior before you consider it fixed. Then put a recurring quarterly reminder on it and update it immediately after any price change, since this is the highest-stakes defect on the entire quality checklist.
How do I know whether to upgrade my plan?
Compare your minutes used against minutes included, and check whether overage is becoming a meaningful share of your bill. Core is $500 per month for 500 AI minutes with 45 cents per minute after, Pro is $750 for 1,000 minutes at 40 cents after, and Elite is $1,200 for 2,500 minutes at 35 cents after — so there is a crossover point where the higher tier costs less than your current tier plus overage. If you are consistently well under your included minutes and only need reliable message-taking, KeyBot Lite at $149 per month with 100 calls included is the smaller-fit option. All tiers are at https://www.thekeybot.com/pricing.
What if the 90-day numbers are ambiguous?
Fix every quality defect you found at day 60, then run one more 30-day period with no other changes and measure again. Ambiguity almost always comes from too many variables moving at once — a price change, a new ad campaign, a seasonal swing, and a configuration tweak all inside the same window — which makes the result uninterpretable rather than negative. One clean month with a frozen configuration produces a more usable answer than three noisy ones, and it costs you nothing but patience.
About the Author
TheKeyBot Team is dedicated to helping locksmiths grow their businesses through AI automation and smart technology. With years of experience in the locksmith industry, our team provides actionable insights and proven strategies.
