You bring the audio
Real recordings from your own queue beat any curated sample set. Anything from a phone call to a WhatsApp voice note — MP3, WAV, M4A, OGG/Opus, FLAC, WebM or AMR.
AppCraft’s voice-forensics stack for banks, insurers and contact centres — flags synthetic speech from every major TTS / voice-clone engine in production today.
This detector runs inside a client’s perimeter, so we show it live on a call rather than as a public sandbox — on your own call recordings, with your own thresholds.
Live demo
Bring a handful of your own recordings — a suspected vishing call, a claim statement, a voice note that felt off — and we score them live while you watch the read-outs move. Forty minutes is usually enough to know whether this belongs in your flow.
Real recordings from your own queue beat any curated sample set. Anything from a phone call to a WhatsApp voice note — MP3, WAV, M4A, OGG/Opus, FLAC, WebM or AMR.
Verdict, per-channel contributions and the full heuristics panel, clip by clip. Where the stack is unsure, you see that too — the borderline cases are the interesting part.
Where the check sits in your call flow, what latency it adds, how thresholds are tuned to your risk appetite, and what on-prem or VPC deployment looks like on your infrastructure.
The detector runs inside a client’s perimeter, so we show it live on a call rather than as a public page. That also means the session is worth more than a sandbox would be: your audio, your channel conditions, your thresholds.
Technical screener, not a verdict. This is a forensics signal — not proof a voice is real or cloned. A high-end clone laundered through a phone codec can pass, and real speech can trip a flag, so the output routes a case to verification rather than deciding it.
Data handling: in a deployment the audio and every artefact it produces stay inside your perimeter, under your retention policy. We never train models on client audio. See our privacy & data policy.
The report
Every clip comes back with a verdict, a confidence figure, the contribution of each detection channel and the raw heuristics read-outs. Nothing is a black box you have to take on faith — your analysts see the same numbers the ensemble saw.
Multiple detection channels fired. The ensemble reads the speech as synthesised by a TTS or voice-clone engine. Treat the speaker as untrusted for any flow where voice authenticity matters.
No synthetic-speech signals fired. The audio is consistent with a real human recording. Absence of evidence is not proof of authenticity — a high-end clone laundered through a phone codec is harder — but for everyday workflows this is a green light.
Signals were mixed. The clip may be borderline — heavily compressed real speech, a non-speech sample, or a TTS family the calibration does not cover yet. In production these route to manual review rather than to a decision.
Each channel votes independently and the report shows what each one contributed, so a verdict can be argued with rather than merely accepted. The ensemble weighting itself stays internal.
Wav2Vec2-XLS-R, fine-tuned on commercial 2025 TTS (ElevenLabs, Kokoro, Hume, Polly, Speechify, Luvvoice). Reports a calibrated P(AI-generated) per window.
XLS-R 300M trained on ASVspoof-flavour data — a different training lineage, deliberately biased conservative, so it disagrees with channel 1 where it should.
Phase chaos, F0 stability, silence-floor cleanliness, HNR, MFCC micro-jitter — physics rather than training data, which is what keeps the stack useful against an engine nobody has trained on yet.
A lightweight on-device model (Whisper-tiny) confirms the clip contains continuous speech and detects its language, and the report carries a rough auto-draft of what was heard. It is a sanity check, not a transcription service — the wording has no effect on the authenticity verdict, it only tells you whether the pipeline had speech to judge in the first place.
Eight read-outs from the heuristics layer travel with every report. Surfaced so you can see the signal — the actual ensemble weighting stays internal.
These eight are a slice of the full pipeline. A production deployment runs additional channels (speaker identity, prosody modelling, mid-call drift, channel fingerprint on PSTN/VoIP) before issuing a verdict.
Coverage
Our primary classifier is calibrated for the 2024–2026 commercial neural-vocoder TTS landscape — the actual attack surface for vishing and claim-fraud calls today.
Listen + compare
Two ten-second English clips — one a public-domain reading by a real person, the other generated by OpenAI’s commercial TTS in May 2026. To the ear they read like the same kind of polished, mid-tempo English narration. Our stack tells them apart in under three seconds.
Human reading
Live recording · English
A short clip from a public-domain English-language reading. Natural micro-jitter in pitch and timing, irregular phase, expressive prosody — all the things modern vocoders smooth out.
AI-generated
Modern neural TTS · English · 2026
Generated by a current-generation commercial voice engine. Reads as fluent English to most listeners. Our primary classifier catches the vocoder fingerprint within the first 3 seconds.
We were unable to clone a specific speaker’s timbre for this demo without a paid voice-cloning provider, so the two samples are different voices reading different sentences. The point is the verdict: real human → Human; commercial TTS → AI-generated, both with high confidence. In production we run this same stack against bank-grade voice-clone calls (matching a real account-holder’s timbre) and still surface them.
Why this matters now
Five seconds of leaked WhatsApp voice note is enough to clone an executive. The attacks moved from theory to weekly incidents on African finance teams in 2024–2025.
Year-on-year increase in deepfake-related fraud incidents across Africa over the past year.
Sumsub Identity Fraud Report 2025
Global increase in deepfake incidents targeting fintech onboarding flows in 2023 alone.
Sumsub / Identity Fraud Report
Loss in one engineered video-and-voice deepfake call against an MNC finance officer in Hong Kong (2024).
CNN / Hong Kong Police
Share of biometric fraud cases tracked across African fintechs that involve generative AI in 2026.
Smile ID 2026
Deloitte estimate of US generative-AI-driven fraud losses, much of it voice-clone-enabled vishing.
Deloitte Center for Financial Services
How little reference audio modern voice-cloning systems require to produce a convincing impersonation.
McAfee Beware the Artificial Impostor 2023
Sources: Sumsub Identity Fraud Report 2025, Smile ID Digital Identity Fraud Report 2026, CNN coverage of the Arup HK deepfake heist (2024), Deloitte Center for Financial Services, McAfee "Beware the Artificial Impostor" 2023.
Built for
We deploy AppCraft voice-forensics as an on-prem service, in your VPC, or as a managed REST / WebSocket API behind your own ingress. It plugs into the workflows your team already uses.
Score inbound calls and voice-message requests for synthetic speech before the call lands on a human agent. Sub-call latency, audit-trail logging your compliance team will accept.
First-notice-of-loss calls and recorded statements run through the voice forensics pipeline. Flag AI-fabricated witness recordings before they enter the adjuster’s queue.
Layer voice-deepfake detection on top of voice-biometric auth, or screen citizen-portal call-back requests for synthetic speech at scale.
Need this in production?
Tell us about your call volume and your verification flow. We book a walkthrough on your own recordings and come back with an integration plan and a price.