How we test your AI agent
You hired Cordiva because you needed a receptionist that doesn't have bad days.
We take that seriously. Every Cordiva agent is verified weekly against a battery of 35 scenarios specific to your practice — and we publish the results back to you on this dashboard.
This page explains how that works.
What we test
Each week, your agent takes 35 simulated calls. The scenarios cover the situations that matter most for a med spa front desk:
Patient privacy. Your agent is HIPAA-compliant by design. We run scenarios where someone tries to extract another patient's appointment time, treatment history, or contact details. Your agent must refuse, every time, without confirming whether the person being asked about is even a patient.
Medical safety. Some callers describe symptoms that need urgent care. We test how your agent recognizes signs of severe allergic reactions, breathing trouble, stroke, or self-harm — and routes those callers to 911 or 988 immediately, without trying to book an appointment instead.
Bilingual quality. A meaningful share of Florida med spa callers switch between English and Spanish in a single call. We test that your agent follows the caller's language, stays consistent through the whole call, and doesn't accidentally confirm a booking in the wrong language.
Booking edge cases. Real callers say things like "this number I'm calling from", "tomorrow whenever", or have unusual surnames the system might mishear. We test that your agent reads back contact details letter by letter when something might be ambiguous, and proposes specific times rather than asking the caller to call back.
Pressure and edge cases. Some callers are rude, drunk, or trying to manipulate the agent into doing something it shouldn't — claiming to be a doctor, demanding a discount, or asking your agent to break protocol. We test that your agent stays calm, professional, and ends every call gracefully.
How we measure
After each weekly batch, every test case gets a clear pass or fail outcome based on objective criteria. There's no subjective scoring — your agent either said the right thing or it didn't.
Over the rolling 5-week window, your agent's performance is summarized as a single number — the percentage of scenarios passed on average — and a trend line that tells us whether quality is stable, improving, or drifting.
We track:
- The average pass rate over the rolling 5-week window
- Whether the agent has any critical scenarios it consistently fails
- Whether the variance between weeks is small enough to indicate reliable behavior
- Whether any scenario has failed three weeks in a row, which is our earliest signal that something needs human attention
These five checks together produce the badge you see on your dashboard.
When something drifts
Voice agents are software, and software can drift. Your services change. Your prices change. The way patients ask things changes over time. Without continuous testing, an agent that's perfect today can be subtly broken in three months — and you wouldn't know until a patient told you.
When a scenario starts failing, we catch it within a week. The fix is reviewed by a human — the founder personally — before it goes live. Your agent never receives an automatic update that hasn't been verified.
If a critical scenario fails or your overall pass rate drops by more than 5 points, your status badge turns from green to yellow with a note that we're investigating. We aim for this to be a rare event.
Why this matters
Most AI receptionist providers don't measure their agents continuously after the demo. The agent that wowed you in the sales call three months ago might be quietly missing 1 in 5 calls today. You'd never know until a patient mentioned something on her next visit.
For context: a typical front desk receptionist, on her best month, handles roughly 75–85% of these same edge case scenarios correctly. (And on her tired, sick, or distracted days, that number drops further.)
Your Cordiva agent has averaged 88–95% across these scenarios consistently — across weekends, holidays, peak hours, and 3am calls.
The point isn't perfection. No agent — human or AI — is 100% perfect every interaction. The point is consistency: your patients get the same level of care at 7am Sunday as they do at 11am Tuesday.
What your dashboard shows
Your dashboard's status badge updates automatically after each weekly check:
- 🟢 Operational — quality verified, all checks passing
- 🟡 Operational with minor drift — we're investigating, no action needed from you
- 🔴 Under maintenance — your agent is still answering calls; we're actively reviewing the issue
If you'd ever like more detail on a specific week — anonymized aggregate, never specific transcripts — just ask.
Questions?
Reply to your monthly check-in email or contact us at hello@cordiva.ai.
We built this measurement system because the alternative — trust us, it works — isn't a serious answer for a HIPAA-grade tool you put in front of your patients. We'd rather show our work.
Last updated: May 2026 · Quality verification powered by Cordiva's Continuous Voice QA framework