newCallback verification is a control that voice cloning quietly broke.
Proveniro
Detect

An AI verdict fast enough to change the outcome.

Post-hoc detection produces a report. Proveniro produces a decision — inside the call, inside the onboarding session, inside the approval window — with the model reasons attached so a human can act on it.

verdict latency
< 400 ms
audio required
3 s
ai models per verdict
5
languages
24
proveniro · verify.livelive
+44 20 7946 0812
Claims to be: Group CFO
Payment authorisation · callback
ref PA-40871
capturing 3.2 s sample16 khz · 5 ai models
  • codec_fingerprint
  • prosody_continuity
  • spectral_artefacts
  • replay_liveness
  • registry_match
establishing session…
What we measure

Not ‘does this sound like them’. Does this sound like a person at all.

Voice biometrics asks an identity question, and a generative AI clone is trained to pass it. Proveniro asks a physics question first: could a human vocal tract, in a real room, on a real line, have produced this signal? Identity is one signal among five, weighted last.

Generative AI speech leaves residue. Neural vocoders smooth spectral detail in ways lungs do not. Synthesised prosody drifts from natural micro-timing over long utterances. Re-encoding a generated waveform through a telephony codec leaves a chain that does not match a live handset. Individually these are weak signals; together, calibrated to your traffic, they are a decision.

Every verdict ships with reason codes — the specific AI models that flagged, the thresholds in force, and the confidence. Analysts get an explanation, not a number they have to trust.

detection ensemble
  • acousticvocal tract plausibility
  • spectralvocoder & upsampling residue
  • temporalprosody, breath, micro-timing
  • channelcodec chain & re-encode history
  • identityregistry voiceprint match

Relative contribution shown for a single scored sample. No model votes alone; the ensemble is calibrated per customer during shadow evaluation.

Coverage

Voice, face, agent credential, and the pipe the media arrived through.

Attacks do not stay in one modality. A vishing call is followed by a video call; an onboarding selfie is injected straight into the capture stream, never passing a camera at all.

Voice clone detection

Zero-shot TTS, neural voice conversion and cloned speech over PSTN, VoIP and WebRTC — scored on as little as three seconds of audio.

AI agent verification

Authorised agents present a signed credential with a declared synthesis class and scope, so your own automation is allowed through while impersonations are not.

Replay & liveness

Recorded-audio replay, splice attacks and pre-rendered responses, with an optional active challenge for high-value approvals.

Face swap & morphing

Frame-level and temporal artefacts across live video calls and stored onboarding captures, including morphed document portraits.

Injection detection

Virtual cameras and emulators that bypass the sensor entirely — the attack most liveness checks were never designed to see.

Channel forensics

Codec chains, re-encode history and network path signals that reveal media which never travelled a real call leg.

Registry cross-check

Optional match against an enrolled voiceprint, faceprint or agent credential held in the registry, with explicit consent on file.

Deployment

Inline, passive, and reversible.

Nothing about the caller or customer experience changes. Detection runs on a media fork and returns a signal to the system that already owns the decision.

  1. 01

    Fork the media

    SIPREC, WebRTC, or a REST upload from your recording pipeline. No changes to the call path.

    day 1

  2. 02

    Score in shadow

    Every call scored, no decisions taken. Base rates and thresholds established on your traffic.

    weeks 1–2

  3. 03

    Surface to humans

    Risk badge and reason codes in the agent desktop, case tool or approval queue.

    week 3

  4. 04

    Automate the edges

    Auto-hold the highest-confidence synthetic verdicts; route the ambiguous middle to review.

    week 4+

technical specification
input
8 kHz / 16 kHz PCM, Opus, G.711, G.722 · H.264 / VP8 video · image capture
transport
SIPREC, WebRTC media fork, RTP mirror, REST batch, S3-compatible drop
latency
Under 400 ms p95 for a 3-second window; streaming partial verdicts available
output
Decision, label, confidence, per-model reason codes, agent credential status, evidence identifier
integrations
Genesys, Amazon Connect, Twilio, NICE, Five9, ServiceNow, Salesforce, generic webhook
residency
EU (Frankfurt, Dublin), UK (London), US (Virginia, Oregon)
retention
Configurable 0–90 days for media; evidence records retained on your schedule
model training
Customer media is never used to train or fine-tune any model, under contract
Objections

What a fraud team asks in the first ten minutes.

The honest answer is that it depends on your traffic, your codecs and your customer base, and any vendor quoting a single number without seeing your data is quoting a benchmark, not a forecast. We run a shadow evaluation, publish precision and recall against your own base rates, and let you pick the operating point.

Get started

See the models run on your own calls.

A shadow evaluation takes days to stand up and answers the only question that matters: what would this have caught, and what would it have got wrong?

soc 2 type ii in progress · eu & uk data residency · no ai training on your data