AppCraft.Africa
AppCraft voice-forensics · proprietary stack · ~3 s

Is this voice real?
Our stack answers in seconds.

AppCraft’s voice-forensics stack for banks, insurers and contact centres — flags synthetic speech from every major TTS / voice-clone engine in production today.

This detector runs inside a client’s perimeter, so we show it live on a call rather than as a public sandbox — on your own call recordings, with your own thresholds.

Audio stays inside your perimeter On-prem, your VPC, or a managed API ~3 s per 10-second clip

Live demo

We run it on your audio, on a call.

Bring a handful of your own recordings — a suspected vishing call, a claim statement, a voice note that felt off — and we score them live while you watch the read-outs move. Forty minutes is usually enough to know whether this belongs in your flow.

Step 1

You bring the audio

Real recordings from your own queue beat any curated sample set. Anything from a phone call to a WhatsApp voice note — MP3, WAV, M4A, OGG/Opus, FLAC, WebM or AMR.

Step 2

We score it in front of you

Verdict, per-channel contributions and the full heuristics panel, clip by clip. Where the stack is unsure, you see that too — the borderline cases are the interesting part.

Step 3

You get an integration plan

Where the check sits in your call flow, what latency it adds, how thresholds are tuned to your risk appetite, and what on-prem or VPC deployment looks like on your infrastructure.

No public sandbox — a real walkthrough instead.

The detector runs inside a client’s perimeter, so we show it live on a call rather than as a public page. That also means the session is worth more than a sandbox would be: your audio, your channel conditions, your thresholds.

Technical screener, not a verdict. This is a forensics signal — not proof a voice is real or cloned. A high-end clone laundered through a phone codec can pass, and real speech can trip a flag, so the output routes a case to verification rather than deciding it.

Data handling: in a deployment the audio and every artefact it produces stay inside your perimeter, under your retention policy. We never train models on client audio. See our privacy & data policy.

The report

One verdict, and everything that produced it.

Every clip comes back with a verdict, a confidence figure, the contribution of each detection channel and the raw heuristics read-outs. Nothing is a black box you have to take on faith — your analysts see the same numbers the ensemble saw.

AI-generated

Multiple detection channels fired. The ensemble reads the speech as synthesised by a TTS or voice-clone engine. Treat the speaker as untrusted for any flow where voice authenticity matters.

Human

No synthetic-speech signals fired. The audio is consistent with a real human recording. Absence of evidence is not proof of authenticity — a high-end clone laundered through a phone codec is harder — but for everyday workflows this is a green light.

Inconclusive

Signals were mixed. The clip may be borderline — heavily compressed real speech, a non-speech sample, or a TTS family the calibration does not cover yet. In production these route to manual review rather than to a decision.

Three channels, fused

Each channel votes independently and the report shows what each one contributed, so a verdict can be argued with rather than merely accepted. The ensemble weighting itself stays internal.

Channel 1

Primary classifier

Wav2Vec2-XLS-R, fine-tuned on commercial 2025 TTS (ElevenLabs, Kokoro, Hume, Polly, Speechify, Luvvoice). Reports a calibrated P(AI-generated) per window.

Channel 2

Secondary classifier

XLS-R 300M trained on ASVspoof-flavour data — a different training lineage, deliberately biased conservative, so it disagrees with channel 1 where it should.

Channel 3

Vocoder heuristics

Phase chaos, F0 stability, silence-floor cleanliness, HNR, MFCC micro-jitter — physics rather than training data, which is what keeps the stack useful against an engine nobody has trained on yet.

Speech-presence check

Is there actually speech?

A lightweight on-device model (Whisper-tiny) confirms the clip contains continuous speech and detects its language, and the report carries a rough auto-draft of what was heard. It is a sanity check, not a transcription service — the wording has no effect on the authenticity verdict, it only tells you whether the pipeline had speech to judge in the first place.

Signal panel ~3 s per clip

What the stack measures

Eight read-outs from the heuristics layer travel with every report. Surfaced so you can see the signal — the actual ensemble weighting stays internal.

Phase chaos
Vocoders reconstruct phase; real rooms and real larynxes do not. Synthetic speech sits unnaturally ordered.
Pitch variability
Standard deviation of F0. Human pitch wanders; generated pitch contours are smoother than a person can hold.
Spectral smoothness
Mean spectral flatness. Over-smoothed spectra are a classic neural-vocoder fingerprint.
Silence floor
What the quiet parts sound like. Real recordings carry room tone; synthesis often carries a suspiciously clean floor.
HNR
Harmonic-to-noise ratio — how much breath and turbulence rides along with the voiced part.
MFCC jitter
Frame-to-frame micro-variation in timbre. Humans jitter; generators drift.
Voiced ratio
Share of the window that is actually voiced speech, used to sanity-check the rest of the read-outs.
RMS loudness
Level and dynamics of the clip — context for the channel it arrived through.

These eight are a slice of the full pipeline. A production deployment runs additional channels (speaker identity, prosody modelling, mid-call drift, channel fingerprint on PSTN/VoIP) before issuing a verdict.

Coverage

Detects speech from every major voice-clone engine.

Our primary classifier is calibrated for the 2024–2026 commercial neural-vocoder TTS landscape — the actual attack surface for vishing and claim-fraud calls today.

  • ElevenLabs
  • OpenAI TTS (tts-1 / tts-1-hd / gpt-4o)
  • Hume AI
  • Amazon Polly Neural
  • Microsoft Azure Neural Voice
  • Google WaveNet / Chirp
  • Speechify
  • Kokoro
  • Luvvoice
  • Resemble AI
  • PlayHT
  • Bark / Coqui XTTS
  • and many others

Listen + compare

One human voice. One AI clone. Two seconds of analysis.

Two ten-second English clips — one a public-domain reading by a real person, the other generated by OpenAI’s commercial TTS in May 2026. To the ear they read like the same kind of polished, mid-tempo English narration. Our stack tells them apart in under three seconds.

Human reading

Live recording · English

Ground truth: Human

Real voice, recorded in the wild

A short clip from a public-domain English-language reading. Natural micro-jitter in pitch and timing, irregular phase, expressive prosody — all the things modern vocoders smooth out.

Human · 74%

AI-generated

Modern neural TTS · English · 2026

Ground truth: AI

Commercial neural TTS

Generated by a current-generation commercial voice engine. Reads as fluent English to most listeners. Our primary classifier catches the vocoder fingerprint within the first 3 seconds.

AI-generated · 70%

We were unable to clone a specific speaker’s timbre for this demo without a paid voice-cloning provider, so the two samples are different voices reading different sentences. The point is the verdict: real human → Human; commercial TTS → AI-generated, both with high confidence. In production we run this same stack against bank-grade voice-clone calls (matching a real account-holder’s timbre) and still surface them.

Why this matters now

The voice-clone fraud wave hit Africa first.

Five seconds of leaked WhatsApp voice note is enough to clone an executive. The attacks moved from theory to weekly incidents on African finance teams in 2024–2025.

1 410%

Surge in deepfake fraud, Africa

Year-on-year increase in deepfake-related fraud incidents across Africa over the past year.

Sumsub Identity Fraud Report 2025

700%

Growth in deepfake attempts, fintech

Global increase in deepfake incidents targeting fintech onboarding flows in 2023 alone.

Sumsub / Identity Fraud Report

$25M

Single-call voice-deepfake heist

Loss in one engineered video-and-voice deepfake call against an MNC finance officer in Hong Kong (2024).

CNN / Hong Kong Police

69%

African fintech fraud is AI-generated

Share of biometric fraud cases tracked across African fintechs that involve generative AI in 2026.

Smile ID 2026

$10B

Projected AI-fraud losses by 2027

Deloitte estimate of US generative-AI-driven fraud losses, much of it voice-clone-enabled vishing.

Deloitte Center for Financial Services

5 sec

Voice needed to clone a target

How little reference audio modern voice-cloning systems require to produce a convincing impersonation.

McAfee Beware the Artificial Impostor 2023

Sources: Sumsub Identity Fraud Report 2025, Smile ID Digital Identity Fraud Report 2026, CNN coverage of the Arup HK deepfake heist (2024), Deloitte Center for Financial Services, McAfee "Beware the Artificial Impostor" 2023.

Built for

Where this lives in your stack.

We deploy AppCraft voice-forensics as an on-prem service, in your VPC, or as a managed REST / WebSocket API behind your own ingress. It plugs into the workflows your team already uses.

Banks & fintechs

Vishing & voice-OTP fraud

Score inbound calls and voice-message requests for synthetic speech before the call lands on a human agent. Sub-call latency, audit-trail logging your compliance team will accept.

Insurance

Claim audio & FNOL triage

First-notice-of-loss calls and recorded statements run through the voice forensics pipeline. Flag AI-fabricated witness recordings before they enter the adjuster’s queue.

Contact centres & gov

Authentication & e-services

Layer voice-deepfake detection on top of voice-biometric auth, or screen citizen-portal call-back requests for synthetic speech at scale.

Need this in production?

API access, bulk batch and on-prem deployment for medium and large businesses.

Tell us about your call volume and your verification flow. We book a walkthrough on your own recordings and come back with an integration plan and a price.