
What is AI voice cloning? How to detect AI voices and fake calls
Key takeaways
- Short audio samples are enough to produce convincing real-time voice clones
- Gartner found that 62% of organisations have already experienced a deepfake attack
- Successful vishing incidents cost an average of 1.5 million US dollars in incident response
- Typical patterns are CEO fraud, help desk manipulation, MFA pressure and phishing preparation
- Detection rests mainly on signals such as time pressure, secrecy and process bypass
- AI voice cloning attacks cannot be filtered technically, only trained against, for example through vishing simulations in security awareness training
Real-time voice imitation puts business processes at risk. Find out how AI voice cloning gets past technical controls, and which verification steps hold up when a familiar voice asks for an exception.
Contents
- What is voice cloning, and what is AI voice cloning?
- How does voice cloning work?
- How cybercriminals use AI voice cloning
- AI voice cloning scams
- How to detect AI voices
- Voice cloning authenticity protocols
- How to protect your organisation
- Checklist against AI voice cloning
What is voice cloning, and what is AI voice cloning?
Voice cloning is the artificial reproduction of a human voice. Anyone asking what voice cloning is today almost always means the AI-driven form of it. AI voice cloning relies on generative speech models that derive a voice profile from a recording and then generate any new sentence from it. Unlike earlier speech synthesis, the sentence in question never has to have been spoken. The model carries timbre, speech rhythm, stress patterns and regional colouring over to new text. In technical terms, the meaning of voice cloning has shifted: the results now fall under audio deepfakes, synthetic media that audibly imitate a real person.
For organisations, this stopped being a niche topic some time ago. A Gartner survey of 302 security leaders across North America, EMEA and Asia-Pacific, conducted between March and May 2025, found that 62% of organisations had experienced a deepfake attack that either used social engineering or exploited automated processes. The knock-on costs are considerable. A successful vishing attack costs affected organisations an average of 1.5 million US dollars in direct recovery and incident response.
The technology is not purely an attack tool, however. It powers multilingual dubbing, read-aloud features in accessible applications, and speech restoration for people who have lost their voice to illness. That same dual use is what makes the technology hard to defend against. There is no single provider you could switch off, and no reliable signature for a filter to look for.

Straight to the solution?
Costly vishing incidents start with a single phone call. Train your team to recognise AI voice cloning before it reaches finance or the help desk.
How does voice cloning work?
AI voice cloning runs through three steps, and they are much the same whichever service is used.
- Voice sample: The model needs a recording of the target voice. What matters is the quality of the material rather than its length. Microsoft demonstrated as early as 2023, with its VALL-E research model, that three seconds of audio are enough to reproduce a voice recognisably. Commercial services work in a similar range today. ElevenLabs, for example, recommends one to two minutes of clean audio for its instant voice clone.
- Model training: A pre-trained speech model that has generalised from tens of thousands of hours of speech is adapted to the target voice. It extracts a voice profile, a mathematical description of timbre, frequency behaviour and speaking habits. Depending on the service, this step takes minutes rather than days.
- Generated output: The system then converts written input, or the attacker’s own live speech, into the target voice. Because this happens in real time, the result is not limited to a pre-recorded message. Voice cloning scam calls can be held as genuine conversations, complete with follow-up questions, laughter and hesitation.
One technology, two applications
The tools behind legitimate applications and attacks are identical. Providers such as ElevenLabs supply voice synthesis for audiobooks, localisation and assistive systems, and attackers have access to the same quality. A Consumer Reports investigation published in March 2025 examined six widely used voice cloning services and found that four of them asked for nothing more than a tick box or a comparable self-declaration confirming that the user was entitled to use the voice. With four of the six, a name and an email address were enough to open an account. On the evidence of that study, the barrier to entry is minimal, and it sits in the open market rather than on the dark web.
How cybercriminals use AI voice cloning
AI voice cloning is not a new attack vector in its own right. It amplifies patterns that are already familiar, and it moves the attack onto a channel that technical controls barely cover.
What do cybercriminals need to create a fake voice?
Three things, and none of them is difficult to obtain.
First, a voice sample. It rarely comes from a breach. Public sources supply more than enough material: podcast appearances, webinar recordings, conference videos, interviews in trade media, product clips on LinkedIn, voicemail greetings. Executives, press officers and sales leaders are particularly exposed, because their voices are online as a function of their job.
Second, contextual knowledge. A call only sounds convincing once it names the right people, projects and responsibilities. Organisational charts, press releases about acquisitions, job adverts listing internal system names and out-of-office replies all supply that material. An executive travelling on business is a classic trigger, because it gives attackers a plausible reason why checking back is difficult.
Third, a delivery route. Caller IDs can be spoofed so that a known internal number appears on the display. Attackers also move the conversation to a messenger app or a video call, where the number stops mattering altogether.
How is voice cloning used by cybercriminals?
A cloned voice buys trust that an email cannot, and attackers combine that trust with time pressure and confidentiality. The pattern is consistent: a familiar authority, an urgent exception, a request for discretion.
Recent incidents show how common the approach has become. On 15 May 2025, the FBI warned of an ongoing campaign in which criminals impersonated senior US officials by text message and AI-generated voice message to gain access to accounts and contact networks.
CrowdStrike reported that almost every Scattered Spider incident it investigated up to mid-2025 involved telephone social engineering against the IT help desk, aimed at getting passwords and MFA factors reset for Microsoft Entra ID, single sign-on and virtual desktops. One detail stands out. The callers answered the help desk’s verification questions correctly, because they had researched the answers in advance.
In August 2026, a coordinated wave of vishing calls using AI-generated voices hit several large US financial firms, among them Citadel, Two Sigma and Point72. Two Sigma stated that it had found no indication of any impact on data or systems, and Point72 told investors that no client data had been lost. Citadel made no public comment.
Voice cloning-as-a-service
The criminal market has packaged the technology as a service. In June 2026, Rapid7 described an underground offering that markets voice clones alongside face swapping, synthetic profiles, forged documents and camera injection for video calls, aimed squarely at defeating identity checks. That shifts the picture on the attacker side. Anyone who wants to use a convincing voice clone no longer needs to understand the models, only to hold an account with a service of this kind. Rapid7 describes underground pricing as volatile, ranging from entry-level offers costing a few dozen US dollars to several thousand for professional systems. For defenders, it means planning for a broader and less specialised set of adversaries than two years ago.
AI voice cloning scams: the attack patterns inside organisations
An AI voice cloning scam looks different inside a company from the emergency calls reported in consumer cases. Four patterns come up regularly.
CEO fraud with a cloned voice
In CEO fraud, attackers pose as a member of the executive team and push through a payment or a change of bank details. As long as the attack arrived by email, it often failed at a simple counter-check. Finance called the executive on the internally held number and had the instruction confirmed. AI voice cloning devalues exactly that control, because the attack now arrives as a call in the first place. Someone who has already heard what they believe to be the CEO’s voice treats the check as done, even though nobody called back. The rule that still holds is that the verification has to be outbound. Hang up and dial the number held in the internal directory, not the one shown on the display and not the one given over the phone. Scenarios with built-in secrecy make the pattern harder to break, for example an acquisition that is supposedly under way and must not be discussed with anyone.
MFA approvals and help desk fraud
MFA requests usually only succeed once a person approves them, and that is precisely where the phone call comes in. While push notifications are arriving, a voice from what appears to be internal IT asks the recipient to approve the prompt so that the fault can be fixed. This is how a call using deepfake voice cloning turns an MFA fatigue attack into a successful intrusion. The caller supplies a plausible explanation, and the employee clears the request. Help desks are a target for the same reason. Attackers present themselves as staff locked out of an account and press for a quick reset. Organisations handling a high volume of password, VPN and MFA token requests are especially exposed.
Multi-channel attacks: when the call announces the phishing email
The call is not always the attack itself. It is often the preparation for one. Someone who appears to be a known contact announces a message, which then arrives shortly afterwards. The phishing email that follows looks more credible and gets checked less carefully. The same logic works in reverse, where an unremarkable email makes a later call seem legitimate. The difficulty for organisations is that channels are usually monitored in isolation, so an AI voice cloning call and the email that follows it are never connected to each other.
The call about expired SSO access
A further pattern puts AI voice cloning to work on single sign-on access that has supposedly expired. Callers claim to work in IT support, refer to a system migration or a security update, and guide the person towards a login portal or through setting up a second factor again. Because the process sounds technically plausible and the voice sounds familiar, it is rarely questioned. Voice cloning scam calls of this type are aimed at one thing, which is control over an authentication factor.
How to detect AI voices: what to listen for on the call
Ask how to detect an AI voice and the usual answer is technical: spectral analysis, artefact detection, biometric voice recognition. None of that helps during a live call, because nobody starts a forensic analysis mid-conversation. What counts in practice is what a person can actually notice at the time.

Articulation gives the first clues. AI voice cloning sometimes betrays itself through consistently flat intonation, pauses in unnatural places or missing breath sounds. Sentences can be grammatically perfect while carrying no emotional movement at all. Volume that stays exactly level over several minutes can point the same way. Contradictions in the background are worth noticing too, for instance a caller who claims to be driving with no road noise behind them. A faked voice message is even harder to judge than a live call, because there is no chance to ask a spontaneous question. A voice message on its own never confirms who someone is, however familiar the voice sounds.
During the conversation itself, other signals usually tell you more. Systems built on AI voice cloning often respond with a delay when someone interrupts, sidestep unexpected questions, or falter as soon as the conversation leaves the prepared structure. A spontaneous question about a shared detail that is not publicly known can therefore reveal more than any amount of listening to the voice.
Behavioural signals carry the most weight. An unusual channel, an unusual time of day, explicit time pressure, a request for confidentiality, or an instruction that bypasses an established process. Where several of these occur together, the call is worth verifying, no matter how convincing the voice sounds.
The limits of voice cloning detection
Listening alone will not reliably tell you whether a voice is real. A study in the journal PLOS ONE found that participants identified synthetic speech correctly in only around three out of four cases, and that prior training improved the hit rate by less than four percentage points on average. Signals of this kind are useful for raising suspicion, and that is what they are for. Voice cloning detection by ear is not a decision-making instrument, and against deepfake voice cloning, the decision belongs to a defined process rather than to an individual judgement in the moment.
Voice cloning authenticity protocols: proving a voice is genuine
For organisations, the question shifts from detection to evidence. How do you establish that a voice is genuine, and how do you record that the check took place? Voice cloning authenticity protocols answer that as a matter of process rather than tooling, which is where they differ from detection products.
- Provenance and watermarking: Two complementary approaches are taking shape on the technical side. Provenance metadata under the C2PA standard attaches a signed history to an audio file, documenting how a recording was created and subsequently modified. Watermarking instead embeds a signal in the audio itself that people cannot hear and that is designed to survive conversion. Since 2 August 2026, Article 50 of the EU AI Act (Regulation (EU) 2024/1689) has required providers of generative systems to mark synthetic audio, image, video and text content in machine-readable form. The regulation prescribes no specific technology. It does require implementation that is effective, interoperable, robust and reliable, as far as this is technically feasible. Systems already on the market before that date have a transition period until 2 December 2026. Anyone publishing deepfakes is required to disclose that the content was artificially generated, subject to exemptions and lighter obligations for artistic, satirical and certain law enforcement uses. These rules apply directly across the EU. Organisations elsewhere should expect different labelling requirements, or none at all, and check what applies in their own jurisdiction. For day-to-day defence, this matters less than it might sound. Labelling obligations bind compliant providers, not criminal tools. There is a technical limitation as well. A study on the durability of audio watermarks presented at the Interspeech 2025 conference found that none of the methods examined survived conversion by neural codecs intact, because watermark and codec compete for the same signal space. Data compression and codecs of exactly this kind are standard in voice services, which means provenance evidence tends to fail at the moment it would be needed. It remains valuable for examining a recording after the fact, but not for deciding what to do during a live call.
- Verification procedures: While a conversation is under way, a defined checking process offers more than any signal analysis. That starts with calling back on an extension held in the company directory rather than the number shown on the display. Control questions whose answers cannot be researched publicly are a second option. Management, finance functions and executive assistants can agree a code word in advance. External requests to the help desk should be verified through a second, independent channel, for example a confirmation from the manager in an internal system. None of these steps depends on recognising AI voice cloning by ear, which is precisely their advantage.
- Documentation: For these steps to stand up to an audit, they belong in a written procedure. A log recording the time and outcome of each verification makes the practice traceable. Evidence of this kind can feed into management systems aligned with ISO 27001, NIS2 or DORA, the latter two being EU frameworks. The result is a control whose application and effectiveness can be demonstrated, which is what turns voice cloning authenticity protocols from good intentions into something auditable.
How to protect your organisation against AI voice cloning
AI voice cloning cannot be blocked at the gateway. Protection comes from a small number of precise process rules working together with a workforce that recognises the behaviour behind the call and knows what to do about it.
Voice cloning simulations in security training
Processes only work if people apply them under pressure, and that can be practised. Standard e-learning modules explain that vishing exists. What they do not create is the reflex for the moment when a familiar voice asks politely for an exception. Recognising an AI-generated call takes practical experience rather than a definition, which is why simulations add a lived situation to the knowledge component.
SoSafe has offered the Interactive Vishing Lesson for this since June 2026. At the end of an e-learning unit on social engineering, participants work through a realistic AI-supported phone scenario directly in the browser. Organisations do not need to collect employees’ phone numbers or call personal devices. That removes the familiar obstacles to traditional vishing campaigns, including the absence of company mobiles, data protection concerns and, where such bodies exist, consultation with employee representatives.
The approach is particularly relevant for organisations where large parts of the workforce work by phone. Logistics, manufacturing, healthcare and public administration are typical examples, as are organisations handling a high volume of help desk requests.
The accompanying security awareness training is built around behavioural change rather than one-off knowledge transfer. The instructional design draws on learning psychology, aligns with NIST, CIS, ISO/IEC 27001, NIS2 and DORA, and can be ready to run in up to two days. Ongoing operation takes from 15 minutes a week. Knowledge retention after twelve months is 90 per cent.
Recognising AI voices when it counts
Give your teams realistic practice against voice cloning and social engineering, in the situation where it actually matters.
Checklist against AI voice cloning
The right-hand column names functions rather than job titles. In organisations with smaller security teams of three to fifteen people, several of these will sit with the same person.
| Area | Measure | Owner |
| Verification | Call back on the number held in the internal directory for any phone instruction involving money or access | All employees |
| Verification | Agree a code word for executives, finance functions and executive assistants, and change it annually | Security, executive team |
| Payments | Require two-person approval above a defined amount, with no shortcut for urgency | Finance |
| Payments | Confirm changes to bank details through a second, independent channel only | Finance, procurement |
| Help Desk | Never release password or MFA resets on the strength of a phone call alone | IT service |
| Help Desk | Use verification questions whose answers cannot be researched publicly | IT service |
| Access | Roll out phishing-resistant MFA such as FIDO2 keys for privileged accounts | Security |
| Detection | Train the signals: flat intonation, background that does not fit, delayed responses to interruptions, time pressure, requests for confidentiality | Security |
| Reporting | Provide a low-barrier route for reporting suspicious calls, and welcome false alarms explicitly | Security, communications |
| Training | Add vishing simulation to the training programme, for example through the Interactive Vishing Lesson | Security, HR |
| Evidence | Record verification procedures in writing and log the decisions taken | Compliance |
| Review | Compare attack patterns against current incident reports twice a year and update procedures accordingly | Security |








