Satya-Vani: Catching a Cloned Voice While the Call Is Still On

Seven pluggable signals, one explainable 0–100 risk score that updates every half second: how I built live voice-spoof detection.

Share
Satya-Vani logo beside a phone showing an incoming call from an unknown number at 23:40, a waveform turning red, and a high-risk meter
💡
TL;DR
• I built Satya-Vani ("true voice"). It listens to a call and estimates, while the call is still going, how likely it is that the voice is a clone or the caller is running a scam.
• Seven independent, pluggable signals score every 2-second slice of audio. One small YAML file combines them into a single 0–100 risk score that updates every half second.
• Nothing is a black box. Every score comes with a breakdown of what each signal said, how much it counted, and one plain-English sentence explaining why.
• What's built today is the Live Call Path: the runtime detection pipeline plus a browser dashboard. The banking "block the transfer" screen is still on the roadmap.

When a familiar voice isn't who you think

A late-night call from an unknown number. The voice on the line sounds exactly like someone you love, and they're in trouble. They need money, right now, and you mustn't tell anyone. Sometimes a second voice joins in, claiming to be from the police or the "cyber crime cell", warning of legal action if you don't transfer the money immediately.

Your brain does what any brain would do: I know that voice, so this is real. That instinct is exactly what the attack relies on.

Modern voice cloning can copy a voice convincingly from a few seconds of recorded audio: a voice note, a video, a talk. Caller ID is easy to fake. Calling back only helps if you think of it while someone is shouting at you to hurry. The same trick works on a bank employee who gets a call that sounds just like their CFO.

The Satya-Vani blueprint frames the key idea in one line, and I keep coming back to it:

🔑
Key idea: don't ask "who is speaking?", which is exactly what a clone fools. Ask "how was this sound made?" A neural voice generator rebuilds a waveform rather than recording a human throat, and the rebuild leaves fingerprints.

What I built

Satya-Vani is a self-hosted service with a browser dashboard. You give it a call, either by speaking into your microphone or by uploading a recording. It cuts the audio into overlapping 2-second slices and sends each slice past seven independent "judges". It then combines their opinions into one risk score from 0 to 100, which moves as the call goes on. Each score falls into a band: Low (monitor silently), Elevated (show a quiet advisory) or High (block the sensitive action and call back on a known number). Every score also tells you why.

Four steps: a phone call with call details, Satya-Vani slicing audio and scoring it with seven signals, one 0-100 risk score with an explanation, and three action bands: low, elevated, high
The whole idea in one picture: listen, score, explain, act.

How it works

The whole stack comes up with one docker compose up. There is one main Python service, two small speech-to-text helpers, and a log viewer.

Architecture diagram: browser dashboard talks over HTTP and WebSocket to the live-call-api FastAPI service on port 8020, which windows audio, runs seven detectors, fuses them and saves to SQLite; sidecars whisper-live on 9090 and vexyl-stt on 8091 handle transcription; dozzle on 8888 shows logs
Everything that starts on docker compose up. The dashed box is a single FastAPI service.

The dashboard

The dashboard is a plain HTML and JavaScript page with no build step and no framework. It has four tabs: Upload a file, Microphone (live), Voiceprints (enroll someone's voice once) and History (every past session, replayable). A "How this score is calculated" page draws the whole pipeline and fills in the real weights from the running config, so it can't drift out of date.

Dashboard result: risk score 74 HIGH with per-signal bars and explanations
A synthetic test call (computer-generated voice reading a scam script) scores 74, HIGH, with every signal explained.

The service: live-call-api

This is a FastAPI app on port 8020 with thin routes. POST /v1/score/file scores an upload. WS /v1/stream/{session_id} takes live audio one window at a time. The other routes cover history, enrollment and the current formula.

Inside, the code follows a "ports and adapters" layout. Think of a power strip: the engine only knows the shape of the socket (an interface called DetectorPort), and every detector is a plug that fits it.

The seven signals

SignalWhat it listens for, in plain words
Acoustic (AASIST)A published neural network trained to tell real speech from synthetic speech. Its final layer was retrained on Hindi data.
Prosodic (Praat)Tiny wobbles in pitch and loudness (jitter, shimmer) and breathiness (harmonics-to-noise ratio). Human throats have them; synthesis only approximates them.
Voiceprint / context (the "third signal")Compares the voice against an enrolled voiceprint of the person the caller claims to be. If nobody is enrolled, it falls back to common-sense context rules (more below).
IntentA zero-shot classifier that reads the transcript and rates how well it fits fraud-like labels.
Semantic riskA small local LLM (Phi-3-mini, CPU-only, no cloud call) that reads the transcript for urgency, money requests, authority claims and isolation tactics like "don't tell anyone".
Perth watermarkChecks for the invisible watermark that some cloning tools, such as Chatterbox, embed in every clip.
Phase incoherenceA cheap bit of maths on the waveform. Voice generators tend to rebuild "too tidy" phase compared with a real vocal tract.

Transcripts come from two sidecars. whisper-live handles the live mic. vexyl-stt covers 14 Indian languages, and a quick language-ID check decides which engine gets the audio.

Memory and logs

Every score is saved, with its full breakdown, to a SQLite file (sessions.db). Enrolled voiceprints live in voiceprints.db as numeric embeddings only, never raw audio. Every score is also written as one JSON log line, and Dozzle shows those logs in a browser on port 8888.

One number from many signals

Here's the part I'm proudest of, because it's simple. Think of the seven signals as a panel of judges. Some are more experienced, so their vote counts for more. The weights live in config/risk_formula.yaml:

Concept diagram of the risk formula: bar chart of the seven configured weights (acoustic 0.60, prosodic 0.20, third signal 0.20, intent 0.15, semantic 0.15, Perth watermark 0.15, phase incoherence 0.10), steps for abstain-aware renormalisation and per-signal smoothing with alpha 0.35, the weighted-sum formula, and the three bands
The real weights, smoothing factor and band thresholds from the config file.

Three rules turn the votes into one score.

  1. A judge who has nothing to say doesn't vote zero. If there's no transcript yet, or nobody is enrolled, that signal abstains. The remaining weights are rescaled so they still add up to 1, so a silent judge never drags the score toward "safe".
  2. Each judge's opinion is smoothed. The new value is 35% "this slice" and 65% "what you thought before". One odd half-second can't swing the score wildly, and a judge who abstains for a moment keeps its last opinion instead of vanishing.
  3. The final number is banded. 0–34 is Low, 35–69 is Elevated and 70–100 is High, and each band comes with a recommended action.

The weight sizes are deliberate. The newer, less-proven signals (intent, the LLM, the watermark and phase checks) carry small weights, so none of them can push a genuine call into High on its own.

"Tells you exactly why"

A score of 72 with no explanation is useless to the person on the call, and worse for a bank clerk who has to justify blocking a transfer. So every result carries a components list, one entry per signal, with:

  • its raw score, or null if it abstained, along with the reason;
  • its configured weight and the effective weight it actually got this slice;
  • its exact contribution to the final number;
  • the measured values behind it, plus one plain-English sentence, such as "the call was placed at an unusual hour (23:00)".

The neural acoustic model is the honest exception. It can't point at a specific feature, so its sentence says what it actually knows, the probability it assigned to "synthetic", rather than inventing a reason.

Scores are also reproducible. The fusion step is a pure function: the same audio, context and config always give the same score. Every result is stamped with a formula_version, so an old score can always be traced back to the exact formula that produced it.

How this score is calculated page with the pipeline and detector cards
The "How this score is calculated" page shows the real pipeline and the live weights.

Context clues: the scam script, as rules

Back to that late-night call. Even without any voice analysis, it has red flags a careful friend would notice. The contextual rules engine turns those flags into points:

RuleFires when…Points
Unknown numberthe number isn't on file0.30
Unusual hourbetween 22:00 and 07:000.20
Financial requestwords like "transfer", "OTP", "account" (or Hindi/Marathi equivalents)0.30
Urgency language"urgent", "right now", "don't tell anyone", "तुरंत"…0.20
Authority claim"cyber crime", "police", "this is your bank"…0.25
Authority + money togetherboth of the above: the classic scam script+0.25

The total is capped at 1.0. A classic scam call like the one above trips every rule. Each match points to an actual word in the actual transcript, so you can see exactly what triggered it.

Two subtleties made this work on real calls. First, on the live path the transcript is a rolling window of about 8 seconds, so an authority claim made 20 seconds ago would normally "age out". I made those flags sticky: once a session has heard "cyber crime cell", it remembers that for the rest of the call. Second, in the default auto mode, this rule engine is the fallback. If the real person's voice had been enrolled beforehand, the third signal would compare voiceprints instead.

Follow one slice of a live call

Here's what happens every half second while the microphone tab is on:

Sequence diagram with five lanes: browser, ws_router, live transcription, engine, fusion and history. Nine numbered steps from sending a 2-second window to receiving the FusedScore back
One 2-second window's round trip. Transcription runs on its own clock so it never slows the score.
  1. The browser captures mic audio with an AudioWorklet, cuts a 2-second window, and sends it over the WebSocket with the call details (known number, hour, claimed identity).
  2. The router appends only the newest half-second to a rolling transcription buffer, since windows overlap.
  3. Every 4 seconds of new audio, a background task transcribes the last 8 seconds. Separate tasks then run the intent classifier and the local LLM, so a slow one never blocks the others.
  4. The latest finished transcript and flags are merged into this window's context.
  5. The router hands the window and its context to the engine (off the main event loop, so other calls stay responsive).
  6. The engine applies the sticky flags and runs all seven detectors. Each one scores or abstains.
  7. The detector results go to fusion, in config order.
  8. Fusion rescales, smooths and bands the result, saves it to history and logs it.
  9. The full score and its explanation go back over the same WebSocket, and the meter moves.

Uploads follow the same engine, but transcription, intent and the LLM run once for the whole file instead of every few seconds.

Design decisions worth stealing

1. A new risk factor is one class plus one YAML entry

The choice: only one file, the registry, ever names a concrete detector class. To add a signal, you write a class with a score(window, context) method, add it to risk_formula.yaml with a weight, and restart.
Why: it let me add four signals beyond the original three without touching the engine or API.
Trade-off: wiring by dotted path in YAML means a typo shows up at startup, not in your editor.

2. Abstain, don't zero

The choice: a detector returns "no opinion" rather than 0 when it can't judge.
Why: a zero would quietly mean "looks safe", which is the most dangerous possible default in a fraud detector.
Trade-off: the effective weights change from slice to slice, so the response reports both configured and effective weights to avoid confusion.

3. Expensive models: compute once, read many times

The choice: transcription, intent and the LLM run once per call (or in background tasks on the live path). The detectors only read the precomputed results from context.
Why: running an LLM for every 2-second window would melt a CPU.
Trade-off: live transcript signals lag real speech by a few seconds, and the first seconds of a call have no transcript at all.

4. Write the honesty into the code

The choice: every module carries an "honesty note" saying what it does and doesn't prove, and the dashboard shows every intent label's score, not just the winner.
Why: a security tool that oversells itself is its own kind of risk.
Trade-off: the docs are long. I think that's a price worth paying.

What was hard (and what I learned)

  • The model only knew English. AASIST was trained on a 2019 English dataset. I built a licence-checked Hindi corpus (1,423 genuine and 340 spoofed clips) and retrained just its final layer. Equal error rate (the point where false alarms equal misses, so lower is better) dropped from 36.1% to 19.5% on a test that held out both the speakers and the cloning tool. That's real progress, but it's measured on 40 held-out spoof examples, which is a calibration-scale sample.
  • My hand-tuned prosody rule was a coin toss. When I finally measured it on real data, it scored exactly 50% balanced accuracy. A small logistic regression on the same three features reached 82.5% on held-out data.
  • A fast model that lies quietly. The Indian-language speech engine was more accurate and much faster on Hindi and Marathi. Given English, though, it returned Malayalam-script gibberish with no error. Hence the language-ID gate in front of it.
  • Hindi into Urdu. Whisper's "base" model wrote Hindi speech in Urdu script, so the default became "small".
  • Dinner plans are not fraud. The zero-shot intent model rated "are we still on for dinner tonight?" as 38.9% "creating urgency". That's why it carries a low weight, and why the local LLM now runs alongside it. In hand checks, the LLM scored that sentence as zero urgency.
  • One caller, a hundred speakers. An early speaker-splitting bug counted over 100 phantom speakers in a single call. I fixed it by cutting at real pauses and capping the count at 8.
  • Real time is unforgiving. Routing live-mic audio through the Indian-language engine created a growing backlog of 9–13 seconds. I reverted it, then fixed it properly: the language is decided once per session, and each session keeps one persistent connection.

Try it yourself

The quickest way is the live dashboard at play.sanyal.net/satya-vani. Upload a sample recording, add some call details, and watch every signal explain its part of the score.

To run your own copy, you need Docker:

docker compose up --build
# then open the dashboard
open http://localhost:8020/

One heads-up: the Indian-language sidecar wraps a gated HuggingFace model. Your account needs access to it, and you export HF_TOKEN before building. The other models download anonymously.

Prefer the command line? Score a file with some call context:

curl -F "file=@services/live-call-api/samples/genuine_tone.wav" \
     -F 'context={"known_number": true, "hour_of_day": 14}' \
     http://localhost:8020/v1/score/file | python3 -m json.tool

# the exact formula the running instance loaded
curl http://localhost:8020/v1/config | python3 -m json.tool
  • Interactive API docs: http://localhost:8020/docs
  • Logs (Dozzle): http://localhost:8888
  • Try a new formula: edit risk_formula.yaml, then docker compose restart live-call-api. No rebuild needed.

Note that the bundled sample is a synthetic tone, not a voice, so the speech-trained acoustic model will flag it as not human. For a meaningful demo, use real recorded speech.

History tab listing past sessions with one replayed
History keeps every scored session, so any past call can be replayed with its full breakdown.

What's built, and what's next

Built and runningNot built yet
The Live Call Path: windowing, seven detectors, fusion, explanationsThe mock banking approval screen that actually blocks a transfer on High
Dashboard: upload, live mic, voiceprints, historyCalibrated thresholds for voiceprint matching, speaker splitting and the LLM's risk mapping
Speaker splitting for uploads and live callsAn independently measured English baseline for the acoustic model
Persistent history, Docker Compose, JSON logsMore held-out spoof data to make the Hindi numbers statistically solid

Further out are multi-instance plumbing (Redis for in-progress call state, Postgres for history) and automatic retention for stored sessions, which today are deleted by hand. Each of these is a new adapter behind an existing port, so the core doesn't change.

Closing

The instinct "I know that voice" is exactly what voice cloning exploits. Satya-Vani's answer is to stop trusting that one instinct and gather many small, explainable clues instead, while the call is still on. You can try the live dashboard and see it for yourself.

If you've worked on voice fraud, deepfake detection or Indian-language speech, I'd love to hear what you'd add as an eighth signal. Leave a comment, and subscribe if you'd like the next build log in your inbox.