Voice cloning costs 30 seconds of audio and $5. The new threat model.
Short answer
Cloning a recognizable version of someone’s voice now takes thirty seconds of clean audio and roughly five dollars on a commercial service. The threat is no longer theoretical: voice-cloned scam calls targeting families and law firms have produced documented losses through 2025 and 2026. The defense is procedural rather than technical, and the changes that actually work take about ten minutes to put in place.
What changed in 2025-2026
Until roughly 2023, voice cloning was a research curiosity. The output sounded uncanny but recognizable as synthetic. Producing a clone required minutes of clean training audio and significant compute. Detection tools worked reasonably well.
The current generation of consumer voice cloning services (ElevenLabs Professional Voice Cloning, Resemble AI, OpenAI Voice Engine, several smaller competitors) crossed three thresholds together. The training audio requirement dropped from minutes to seconds. The output quality crossed the threshold where untrained listeners can no longer tell the difference in short phone calls. The price dropped to a few dollars per voice. The combination is what creates the operational threat.
The training data is the easy part. Thirty seconds of clean voice from a podcast, a YouTube video, a voicemail, or a video posted on social media is enough. Most adults with a public profile have hours of qualifying audio available without consent. The protective assumption that your voice is private has not survived the year.
The three attack patterns we are seeing
The grandparent scam, upgraded
The traditional version: someone calls an older relative claiming to be a grandchild in trouble, asks for emergency money. Detection rate, decent. The upgraded version uses a cloned voice trained on the grandchild’s social media. Detection rate, low. Documented losses through 2025 ranged from a few thousand dollars per incident to over six figures in cases where the cloned voice asked for a wire transfer to a “lawyer.”
The CEO fraud variant
An accounts-payable employee receives a call from someone whose voice matches the CEO. The caller asks for an urgent wire transfer to a new vendor for a deal that is closing this hour. The employee verifies the voice. The voice matches. The wire goes out. The actual CEO learns about it the next day. Variants of this attack have been documented at companies of every size since 2023, with the audio quality of the cloned voice improving every year.
The legal-services variant
A client receives a call from someone whose voice matches their attorney, asking the client to wire closing funds to an account that is not the firm’s escrow account. The client recognizes the voice and complies. By the time the firm is contacted to confirm, the wire has cleared. This pattern intersects with the broader exposure pattern we covered in your law firm uses ChatGPT from a different angle.
The procedural defense that works
Technical detection of cloned voice in real time is unreliable. The detection tools available to consumers (and to most enterprises) lag the generation tools by twelve to eighteen months. Procedural defense, by contrast, works regardless of the audio quality.
1. Establish a verbal challenge phrase with people who matter
Pick a specific question that only you and a particular family member or colleague know the answer to. The question is not “what is mom’s maiden name.” It is “what did we order at that restaurant in May 2019” or “what is the joke about Uncle Frank.” A caller unable to answer the agreed challenge is not the person they claim to be. Decide the challenge phrase in person or over a verified channel. Use it the first time someone calls asking for money or a sensitive action. The challenge takes five seconds and stops the attack cold.
2. Create a callback rule
Any request for money, credentials, or urgent action made over a phone call is suspended until you call the person back at a number you already have. Not the number that called you. The number stored in your contacts from before the request. The number on the company’s website you reached separately. The number on the back of the credit card. The cloned voice cannot follow you to a different channel under your control.
The rule is the simplest and most effective. It is also the one people abandon under social pressure. Train yourself to slow down. The five minutes of awkwardness while you call back is much shorter than the months of recovery from a cleared wire transfer.
3. Tell the people you work with that you will never call them in a crisis
For executives and business owners specifically: send a memo to your finance team stating that you will never call them with an urgent wire-transfer request. Any voice that calls them claiming to be you with such a request is, by definition, fraudulent. The same memo to your family removes the emotional pressure on the grandparent-scam pattern: any urgent ransom or bail request from a “family member” is checked through the callback rule before any money moves.
The memo is the single highest-leverage document a company or family can produce against this attack class. It is also the one that almost no one has produced. Write it this week.
What about your own audio exposure
You cannot prevent your voice from being cloned if you have any public audio presence. Podcast hosts, journalists who appear on broadcast, lawyers who give CLE presentations, and anyone whose voicemail greeting is the standard “leave a message” all have qualifying training audio in circulation. The realistic framing is no longer about preventing cloning at all, but about reducing what a clone can actually accomplish once it exists. One concrete example of how that audio gets harvested without explicit consent sits in our piece on the Alexa voiceprint class action and what was actually recorded.
Two operational measures help. Replace your voicemail greeting with a non-personalized message (the default carrier greeting, which uses a generic voice rather than yours). This denies the easiest training source for an attacker who only has your phone number. Inform the people who would be targeted by an impersonation that the procedural defenses above are now standing policy. A cloned voice on a real-time call cannot fabricate a callback to a known, verified number.
The framework for deciding whether your specific exposure justifies more aggressive measures is the same one we walk through in how to build a threat model in 20 minutes. For most people the procedural defense is enough. For executives, attorneys handling closing transactions, and high-profile public figures, additional layers are worth the cost.
Frequently asked questions
Can I detect a cloned voice in real time during the call?
Reliably, no. Consumer-grade detection tools lag the generation tools and produce both false positives and false negatives. Some clones have artifacts (slightly off cadence, occasional audio glitch when the model encounters a word it handles poorly) but the better services produce output that crosses the threshold of consumer detection. Procedural defense is the answer because it does not depend on detecting the clone. It depends on not relying on voice recognition for verification.
Are the legitimate voice cloning services legally responsible?
Increasingly, partly. ElevenLabs and Resemble AI both implemented consent-verification flows in 2024 and 2025 that require the voice owner to authorize the clone. The flows have known bypasses. Smaller services and open-source models bypass the flows entirely. The legitimate services are not the primary vector. The open-source and offshore services are.
Should we use voice biometrics to authenticate banking calls?
No. Several major banks rolled back voice-biometric authentication in 2023 and 2024 after voice cloning crossed the quality threshold. The biometric layer is no longer reliable as a sole authentication factor. Banks that still use it generally combine it with other factors. As a customer, prefer banks that have moved off voice-biometric authentication or that explicitly require additional factors.
Does this affect family video calls?
Video deepfakes are improving on a similar trajectory. Quality varies more than voice clones today, with sustained real-time video deepfakes still showing artifacts under careful inspection. The trajectory is the same: in three years, the procedural defenses applied to voice will need to be applied to video. The challenge phrase and the callback rule generalize. They are good habits to build now.
There’s no perfect setup. Anyone selling you perfect is selling fear. The goal is simple: make yourself a harder target than the person next to you.
