An AI verdict fast enough to change the outcome.
Post-hoc detection produces a report. Proveniro produces a decision — inside the call, inside the onboarding session, inside the approval window — with the model reasons attached so a human can act on it.
- verdict latency
- < 400 ms
- audio required
- 3 s
- ai models per verdict
- 5
- languages
- 24
- codec_fingerprint—
- prosody_continuity—
- spectral_artefacts—
- replay_liveness—
- registry_match—
Not ‘does this sound like them’. Does this sound like a person at all.
Voice biometrics asks an identity question, and a generative AI clone is trained to pass it. Proveniro asks a physics question first: could a human vocal tract, in a real room, on a real line, have produced this signal? Identity is one signal among five, weighted last.
Generative AI speech leaves residue. Neural vocoders smooth spectral detail in ways lungs do not. Synthesised prosody drifts from natural micro-timing over long utterances. Re-encoding a generated waveform through a telephony codec leaves a chain that does not match a live handset. Individually these are weak signals; together, calibrated to your traffic, they are a decision.
Every verdict ships with reason codes — the specific AI models that flagged, the thresholds in force, and the confidence. Analysts get an explanation, not a number they have to trust.
- acousticvocal tract plausibility
- spectralvocoder & upsampling residue
- temporalprosody, breath, micro-timing
- channelcodec chain & re-encode history
- identityregistry voiceprint match
Relative contribution shown for a single scored sample. No model votes alone; the ensemble is calibrated per customer during shadow evaluation.
Voice, face, agent credential, and the pipe the media arrived through.
Attacks do not stay in one modality. A vishing call is followed by a video call; an onboarding selfie is injected straight into the capture stream, never passing a camera at all.
Voice clone detection
Zero-shot TTS, neural voice conversion and cloned speech over PSTN, VoIP and WebRTC — scored on as little as three seconds of audio.
AI agent verification
Authorised agents present a signed credential with a declared synthesis class and scope, so your own automation is allowed through while impersonations are not.
Replay & liveness
Recorded-audio replay, splice attacks and pre-rendered responses, with an optional active challenge for high-value approvals.
Face swap & morphing
Frame-level and temporal artefacts across live video calls and stored onboarding captures, including morphed document portraits.
Injection detection
Virtual cameras and emulators that bypass the sensor entirely — the attack most liveness checks were never designed to see.
Channel forensics
Codec chains, re-encode history and network path signals that reveal media which never travelled a real call leg.
Registry cross-check
Optional match against an enrolled voiceprint, faceprint or agent credential held in the registry, with explicit consent on file.
Inline, passive, and reversible.
Nothing about the caller or customer experience changes. Detection runs on a media fork and returns a signal to the system that already owns the decision.
- 01
Fork the media
SIPREC, WebRTC, or a REST upload from your recording pipeline. No changes to the call path.
day 1
- 02
Score in shadow
Every call scored, no decisions taken. Base rates and thresholds established on your traffic.
weeks 1–2
- 03
Surface to humans
Risk badge and reason codes in the agent desktop, case tool or approval queue.
week 3
- 04
Automate the edges
Auto-hold the highest-confidence synthetic verdicts; route the ambiguous middle to review.
week 4+
- input
- 8 kHz / 16 kHz PCM, Opus, G.711, G.722 · H.264 / VP8 video · image capture
- transport
- SIPREC, WebRTC media fork, RTP mirror, REST batch, S3-compatible drop
- latency
- Under 400 ms p95 for a 3-second window; streaming partial verdicts available
- output
- Decision, label, confidence, per-model reason codes, agent credential status, evidence identifier
- integrations
- Genesys, Amazon Connect, Twilio, NICE, Five9, ServiceNow, Salesforce, generic webhook
- residency
- EU (Frankfurt, Dublin), UK (London), US (Virginia, Oregon)
- retention
- Configurable 0–90 days for media; evidence records retained on your schedule
- model training
- Customer media is never used to train or fine-tune any model, under contract
What a fraud team asks in the first ten minutes.
The honest answer is that it depends on your traffic, your codecs and your customer base, and any vendor quoting a single number without seeing your data is quoting a benchmark, not a forecast. We run a shadow evaluation, publish precision and recall against your own base rates, and let you pick the operating point.
See the models run on your own calls.
A shadow evaluation takes days to stand up and answers the only question that matters: what would this have caught, and what would it have got wrong?
soc 2 type ii in progress · eu & uk data residency · no ai training on your data