If an AI voice agent fails, callers should never hear silence. A properly built deployment detects the failure and sends the call somewhere that works — your team, an old number, a mailbox — without anyone intervening. Ask a supplier what the fallback is, who gets alerted, and how fast the whole thing can be reversed.
Ask a better question than “what is your uptime”
Uptime figures are easy to quote and tell you very little. They describe a platform's own availability, not what your caller experiences, and they say nothing about the parts of the chain that are yours.
The questions that matter: when this fails, where does the call go? Who finds out, and how quickly? And how long does it take to put things back exactly as they were before the agent existed? A supplier with crisp answers to those three has run this before. A supplier who answers with a percentage has not been asked properly.
The four things that can actually fail
“The AI is down” is almost never what happened. Four distinct layers can break, and they need different plans.
Your numbers and the carrier
The telephone side fails independently of anything to do with AI. If calls cannot reach the platform, nothing downstream matters. This layer is also the one you have the most reliable tools for, because call forwarding and routing rules are old, well-understood technology.
The voice platform
The service that answers, listens and speaks. When this is unavailable, the call connects and then nothing useful happens, which is the worst experience on the list — dead air, or a hold that never resolves. This is the failure your fallback has to cover first.
The model behind the conversation
The part that decides what to say can be slow or unavailable while everything else is fine. A well-built agent treats a slow response as a failure and does something sensible, rather than leaving a caller in silence waiting for a sentence that is not coming. Our explainer on how AI voice agents work covers that chain in more detail.
Your own systems
The calendar, CRM, inventory or booking system the agent reaches into. These fail more often than any of the above, and they fail partially, which is harder. The agent can still talk; it just cannot do the thing the caller rang about.
Want the fallback designed before the agent is?
Tell us how calls reach you today and who answers them when nobody is at a desk. We will map what should happen on every failure before anything is built.
Book a Call Flow ReviewWhat a real fallback plan looks like
Most of this is unglamorous telephone configuration, and that is exactly why it works.
Forward, do not port
Keep your numbers where they are and forward them to the agent, rather than moving the numbers onto the new platform. Then undoing everything is one change to a forwarding rule, made by you, in minutes, with no third party involved. Porting numbers makes a reversal a project. This single decision does more for your risk position than any contractual promise.
A default destination that is always valid
Every failure needs a place for the call to go that does not depend on the thing that failed. Usually a ring group of real phones during working hours and a monitored mailbox outside them. The test of whether this is real: unplug the agent and ring your own main number. If anything other than a useful outcome happens, the plan is theoretical.
Degraded service beats no service
Partial failures should produce reduced function rather than nothing. If the calendar is unreachable, the agent can still take the caller's details, say honestly that it cannot confirm a time right now, and promise a callback that actually gets created. If the CRM is down, it can still capture and queue. A caller who is told the truth and gets a callback is a retained caller.
Decide who gets told
Someone on your side, not only the supplier, should be alerted when calls start failing, and it should arrive somewhere a person looks within minutes — a phone, not a dashboard. Name the person. Unassigned alerting is the most common gap we find, and the symptom is always the same: the failure was discovered by a customer complaining.
The failures nobody designs for
Partial failure, which is worse than a clean one
A clean outage triggers the fallback. A degraded service does not: the agent answers, sounds normal, but misses half of what the caller says, or hesitates long enough that people hang up. Nothing alerts, because the platform is technically up. Catching this needs someone looking at call outcomes — how many calls ended without resolution, how many hung up early — rather than at an availability indicator.
The integration that is up but wrong
The quietest failure of all: the agent books appointments successfully into a calendar that nobody is watching any more, or writes leads into a field the sales team stopped using. Everything reports healthy. The business outcome is zero. Somebody has to check the far end of each integration periodically, and that someone should be named before launch.
The surge that is not an outage
More simultaneous calls than the setup was built for behaves like downtime from the caller's side. That is a capacity question rather than a reliability one, and we have written about it separately in how many calls an AI agent can handle at once.
What to settle in writing
Not a percentage. These:
- The fallback destination for each failure, written down, and who can change it.
- Who is notified, through which channel, and who is responsible for noticing.
- Who can switch the agent off, and whether you can do it yourself without asking anyone.
- Where the numbers live and who controls them — the answer should be you.
- Whether your call recordings, transcripts and captured leads are exportable and in what format, so an outage never becomes a data loss.
- What happens to a call that was mid-conversation when the failure started.
That last one is worth asking out loud. It is the question that separates suppliers who have operated a phone line from suppliers who have demonstrated one.
Test it on purpose
Before you rely on it, break it deliberately during a quiet period: point the agent at nothing and call your main number, make the calendar unreachable and try to book, and ring during a planned maintenance window if the supplier has them. Write down what a caller actually experienced each time.
Then repeat it occasionally. Fallbacks rot. Someone changes a forwarding rule, a mailbox fills, a rota changes, the named person leaves. A fallback nobody has tested this year is a plan rather than a control.
Where we would tell you not to rely on it
If a missed call is a safety matter — medical urgency, a trapped lift, a gas smell, a vulnerable person — the emergency path should not depend on the agent making a correct decision. Route those to people by rule, first, before any clever handling. We would rather scope an agent narrower than defend a design that put judgement in the wrong place.
And if your business cannot tolerate any interruption on its main number, start the agent on a secondary number or on overflow only, prove it over real traffic, and move it to the main line once you have seen it fail and recover. That sequence costs you nothing except a little patience.
Where to start
Draw how a call reaches you today, including who answers when nobody is at a desk. That diagram is the fallback plan; the agent simply becomes one more step in it, and every failure means falling back to the version you already had.
Inwizards has been building software since 2009, with teams in the US, UAE and India. AI voice agents covers the phone side, AI agents covers the workflows behind the call, and AI agent development covers building something with limits you can defend. Our guides on replacing an IVR, missed calls, whether AI can answer business calls and agent guardrails cover the neighbouring decisions.