See us at Cisco GSX · August 24 – 27 · Las Vegas. Book a meeting at the show →
Cisco PartnerSenior engineers in North America & Europe · 24/7 escalation
SupportContact Sales
Home/Use Cases/Clinical Voicemail Triage

Turning a blind voicemail backlog into prioritized clinical action — a proof of concept.

A Québec health network needed to know whether AI could read clinical urgency reliably — in Québec French as readily as in English — before committing to a build. Ngenium engineers answered that question on fully synthetic voicemails, with no patient information involved.

Phase-0 proof of concept · synthetic data · no patient information
SCOPE & DATA PROVENANCE — READ THIS FIRST

Everything below was measured on fabricated voicemails.

This was a Phase-0 proof of concept: a go/no-go evaluation built and scored on a purpose-built dataset of synthetic bilingual voicemails. No real patient left any of these messages, no protected health information was processed, and no part of this ran against a live clinical queue.

It demonstrates the solution Ngenium proposes to deliver. It is not a completed production deployment and not a delivered customer outcome. Every number on this page carries that label beside it, because a proof-of-concept measurement and a production result are different claims and should never look alike.

55 min
sooner a true Critical is reached under the ranked queue than under first-in-first-out, in a modelled single-agent backlog
Synthetic POC dataset — not a delivered clinical outcome
100%
of synthetic emergencies caught as Critical by the deterministic safety net — every model, every evaluation run
Synthetic POC dataset — not a delivered clinical outcome
90%
tier agreement with clinician labels; transcription word error rate 4% fr-CA and 3% en-CA
Synthetic POC dataset — not a delivered clinical outcome
The challenge

Three priorities behind a single blind backlog.

When call volume peaks, patients roll to voicemail — and those messages get worked first-in-first-out, triaged only after a human listens to each one. A patient describing chest pain can sit behind dozens of routine scheduling requests, and nobody can see it happening.

01 — BLIND, FIRST-IN-FIRST-OUT

Every voicemail looked identical.

Triage happened only after someone spent the time listening end to end. Until then an urgent clinical concern was indistinguishable from a routine scheduling question, so staff worked messages in the order they arrived rather than the order clinical urgency demanded.

02 — SLAS MISSED SILENTLY

The failure mode was invisible by design.

A time-critical message could sit behind dozens of routine ones and quietly blow past its response target, with nothing surfacing the breach until it was too late. The most urgent callbacks needed to rise automatically, with the clock made visible.

03 — NO SUPERVISOR VISIBILITY

No view of the backlog at all.

Supervisors could not see how many urgent messages were waiting, whether the team was keeping pace, or where the backlog was concentrating — so they could not staff to demand or catch a developing pile-up.

THE COMMON THREAD

All three came from one root cause: voicemail is opaque. Urgency is locked inside an audio file and stays invisible until a person listens — so prioritization, SLA protection and supervisor visibility are all impossible at the exact moment they matter. The evaluation demanded something that could make urgency legible the instant a message arrived, in either language, while keeping patient information inside the hospital network.

The solution

One AI layer, from voicemail to ranked callback.

CADENCE sits behind the existing Cisco voicemail system rather than replacing it. The same code runs hosted for a demonstration or fully on-premises in production, so patient data never has to leave the hospital.

TRANSCRIBE
Bilingual speech-to-text
Whisper transcribes each voicemail in its original language — authentic Québec French first, Canadian English supported. On phone-quality audio the proof of concept measured a 4% word error rate (fr-CA) and 3% (en-CA).
UNDERSTAND
Hybrid urgency classifier
A language model reads clinical keywords and time-sensitivity cues and assigns one of four tiers with a 0–100 score and a plain-language reason. Both the hosted and the recommended on-premises model scored 90% tier agreement with clinician labels.
SAFEGUARD
Deterministic safety net
A rule-based scan of emergency red-flag phrases, in French and English, force-escalates a message to Critical regardless of what the model says. Nothing is ever auto-dismissed, and low-confidence calls are flagged for human review.
PRIORITIZE
Ranked queue & dashboard
Messages surface Critical-first with live SLA countdowns, highlighted trigger phrases and inline audio. A supervisor dashboard shows queue health, SLA breaches and the clinician-agreement rate at a glance.
The engagement

Prove the hardest question first, then build.

The work was structured so the riskiest assumption was tested on synthetic data before anyone committed budget to a clinical deployment. Only Phase 00 has been delivered.

PHASE 00 — DELIVERED
Proof of concept
Built a 32-message synthetic bilingual dataset, the full transcribe-classify-rank pipeline, a branded login-gated application and a formal accuracy scorecard.
PHASE 01 — SCOPED
Discovery & design
Confirm the network’s real SLA tiers and thresholds, the Cisco Unity integration path, and the privacy and data-residency requirements for an on-premises deployment.
PHASE 02 — SCOPED
Pilot build
Wire CADENCE to live Cisco Unity voicemail, add PHI redaction and role-based access, and run a supervised pilot on a bounded queue with clinicians reviewing every tier.
PHASE 03 — SCOPED
Rollout & tuning
Expand across the call centre, tune red-flag lists and thresholds with clinician feedback, and operate the on-premises model with monitoring and audit.
In the Phase-0 evaluation, every synthetic emergency was surfaced as Critical, and the most urgent callbacks were reached far sooner than first-in-first-out — enough evidence to move forward to a clinical pilot.
Outcome of the Phase-0 proof of concept · measured on synthetic data
The outcome

Evidence to fund the build — and a de-risked hardware decision.

Right-sized the hardware before spending

Running Phase 0 in the cloud let our engineers compare hosted and open-weight models side by side, test accuracy across languages and accents, and size the Cisco AI POD and GPUs — so the customer can choose the on-premises model and hardware on evidence rather than guesswork, before committing budget.

PHI never has to touch the cloud

In the proposed production design the entire pipeline runs on the on-premises Cisco AI POD, so no audio, transcript or patient information leaves the hospital data center. The proof of concept confirmed an open-weight model matches the hosted benchmark closely enough to make that realistic.

Human-in-the-loop by design

CADENCE only ranks — it never closes, dismisses or medically judges a message. Every ranking shows the words that drove it, and an agent can re-tier any message with one click. The deterministic safety net does not depend on which model is chosen.

Extends the Cisco estate rather than replacing it

CADENCE sits behind the existing Cisco Communications Manager and Unity Connection, adding an AI prioritization layer to voice infrastructure the network already owns, and landing the new workload on a Cisco AI POD.

Questions

About this proof of concept.

Are these figures from real patient data?

No. Every figure on this page was measured on a fabricated dataset of synthetic voicemails containing no patient information and no protected health information. This was a Phase-0 proof of concept run to answer a go/no-go question. It is a demonstration of capability, not a delivered clinical outcome.

Who was the customer?

We do not name customers without explicit written permission. This proof of concept was built for a Québec health network operating a bilingual clinical call centre with patient voicemail.

Does patient data leave the hospital network?

Not in the proposed production design. The pipeline is built to run entirely on an on-premises Cisco AI POD, so no audio, transcript or patient information leaves the hospital data center. The proof of concept confirmed that an open-weight on-premises model matches the hosted benchmark closely enough to make that realistic.

Does the AI make clinical decisions?

No. CADENCE ranks messages; it never closes, dismisses or medically judges one. A deterministic rule-based safety net force-escalates emergency phrases regardless of what the model says, low-confidence calls are flagged for human review, and an agent can re-tier any message with one click. A person is always in the loop.

Want to see CADENCE prioritize a live queue?

The demonstration is login-gated. Ask us for a walkthrough and we will show you the ranked queue, the safety net and the supervisor view.

Request a Demo