Side profile of a woman with digital sound waves coming out of her mouth.

What is AI voice cloning? How to detect AI voices and fake calls

Updated on: 25 September 2026 · 15 min read

Key takeaways

  • Short audio samples are enough to produce convincing real-time voice clones
  • Gartner found that 62% of organisations have already experienced a deepfake attack
  • Successful vishing incidents cost an average of 1.5 million US dollars in incident response
  • Typical patterns are CEO fraud, help desk manipulation, MFA pressure and phishing preparation
  • Detection rests mainly on signals such as time pressure, secrecy and process bypass
  • AI voice cloning attacks cannot be filtered technically, only trained against, for example through vishing simulations in security awareness training

Real-time voice imitation puts business processes at risk. Find out how AI voice cloning gets past technical controls, and which verification steps hold up when a familiar voice asks for an exception.

Contents

  1. What is voice cloning, and what is AI voice cloning?
  2. How does voice cloning work?
  3. How cybercriminals use AI voice cloning
  4. AI voice cloning scams
  5. How to detect AI voices
  6. Voice cloning authenticity protocols
  7. How to protect your organisation
  8. Checklist against AI voice cloning

Voice cloning simulations are available from several security awareness providers, usually as callback simulations that require phone number management. SoSafe instead runs its Interactive Vishing Lesson in the browser as part of security awareness training, with no private numbers collected. Compare data protection effort, consultation requirements and realism.

Protecting your business against AI voice cloning and audio deepfakes is mainly organisational. Use callback verification on known numbers, two-person approval for payments, no password resets by phone alone, phishing-resistant MFA and security awareness training with realistic voice simulations. Technical filters do not catch phone calls.

The best voice cloning AI cannot be narrowed down to one product. Several commercial services now reach a quality that is hard to see through on a phone call, and research demonstrations have worked from three seconds of audio. The verification procedure matters more than the tool.

Cybercriminals need three things to create a fake voice. A short, clean voice sample from public sources such as podcasts or webinars, contextual knowledge of roles, projects and absences, and a delivery route such as a spoofed number or a video call. No special technical skills are required.

Voice cloning is used by cybercriminals mainly for CEO fraud, for calls to the IT help desk to get access restored, to have MFA requests approved under conversational pressure, and to prepare multi-stage attacks in which a call makes a follow-up phishing email look credible.

What is voice cloning, and what is AI voice cloning?

Voice cloning is the artificial reproduction of a human voice. Anyone asking what voice cloning is today almost always means the AI-driven form of it. AI voice cloning relies on generative speech models that derive a voice profile from a recording and then generate any new sentence from it. Unlike earlier speech synthesis, the sentence in question never has to have been spoken. The model carries timbre, speech rhythm, stress patterns and regional colouring over to new text. In technical terms, the meaning of voice cloning has shifted: the results now fall under audio deepfakes, synthetic media that audibly imitate a real person.

For organisations, this stopped being a niche topic some time ago. A Gartner survey of 302 security leaders across North America, EMEA and Asia-Pacific, conducted between March and May 2025, found that 62% of organisations had experienced a deepfake attack that either used social engineering or exploited automated processes. The knock-on costs are considerable. A successful vishing attack costs affected organisations an average of 1.5 million US dollars in direct recovery and incident response.

The technology is not purely an attack tool, however. It powers multilingual dubbing, read-aloud features in accessible applications, and speech restoration for people who have lost their voice to illness. That same dual use is what makes the technology hard to defend against. There is no single provider you could switch off, and no reliable signature for a filter to look for.

Profile of a man with a sound wave coming out of his mouth facing a cloned version of himself.

Straight to the solution?

Costly vishing incidents start with a single phone call. Train your team to recognise AI voice cloning before it reaches finance or the help desk.

Explore Awareness Training

How does voice cloning work?

AI voice cloning runs through three steps, and they are much the same whichever service is used.

  • Voice sample: The model needs a recording of the target voice. What matters is the quality of the material rather than its length. Microsoft demonstrated as early as 2023, with its VALL-E research model, that three seconds of audio are enough to reproduce a voice recognisably. Commercial services work in a similar range today. ElevenLabs, for example, recommends one to two minutes of clean audio for its instant voice clone.
  • Model training: A pre-trained speech model that has generalised from tens of thousands of hours of speech is adapted to the target voice. It extracts a voice profile, a mathematical description of timbre, frequency behaviour and speaking habits. Depending on the service, this step takes minutes rather than days.
  • Generated output: The system then converts written input, or the attacker’s own live speech, into the target voice. Because this happens in real time, the result is not limited to a pre-recorded message. Voice cloning scam calls can be held as genuine conversations, complete with follow-up questions, laughter and hesitation.

One technology, two applications

The tools behind legitimate applications and attacks are identical. Providers such as ElevenLabs supply voice synthesis for audiobooks, localisation and assistive systems, and attackers have access to the same quality. A Consumer Reports investigation published in March 2025 examined six widely used voice cloning services and found that four of them asked for nothing more than a tick box or a comparable self-declaration confirming that the user was entitled to use the voice. With four of the six, a name and an email address were enough to open an account. On the evidence of that study, the barrier to entry is minimal, and it sits in the open market rather than on the dark web.

How cybercriminals use AI voice cloning

AI voice cloning is not a new attack vector in its own right. It amplifies patterns that are already familiar, and it moves the attack onto a channel that technical controls barely cover.

What do cybercriminals need to create a fake voice?

Three things, and none of them is difficult to obtain.

First, a voice sample. It rarely comes from a breach. Public sources supply more than enough material: podcast appearances, webinar recordings, conference videos, interviews in trade media, product clips on LinkedIn, voicemail greetings. Executives, press officers and sales leaders are particularly exposed, because their voices are online as a function of their job.

Second, contextual knowledge. A call only sounds convincing once it names the right people, projects and responsibilities. Organisational charts, press releases about acquisitions, job adverts listing internal system names and out-of-office replies all supply that material. An executive travelling on business is a classic trigger, because it gives attackers a plausible reason why checking back is difficult.

Third, a delivery route. Caller IDs can be spoofed so that a known internal number appears on the display. Attackers also move the conversation to a messenger app or a video call, where the number stops mattering altogether.

How is voice cloning used by cybercriminals?

A cloned voice buys trust that an email cannot, and attackers combine that trust with time pressure and confidentiality. The pattern is consistent: a familiar authority, an urgent exception, a request for discretion.

Recent incidents show how common the approach has become. On 15 May 2025, the FBI warned of an ongoing campaign in which criminals impersonated senior US officials by text message and AI-generated voice message to gain access to accounts and contact networks.

CrowdStrike reported that almost every Scattered Spider incident it investigated up to mid-2025 involved telephone social engineering against the IT help desk, aimed at getting passwords and MFA factors reset for Microsoft Entra ID, single sign-on and virtual desktops. One detail stands out. The callers answered the help desk’s verification questions correctly, because they had researched the answers in advance.

In August 2026, a coordinated wave of vishing calls using AI-generated voices hit several large US financial firms, among them Citadel, Two Sigma and Point72. Two Sigma stated that it had found no indication of any impact on data or systems, and Point72 told investors that no client data had been lost. Citadel made no public comment.

Voice cloning-as-a-service

The criminal market has packaged the technology as a service. In June 2026, Rapid7 described an underground offering that markets voice clones alongside face swapping, synthetic profiles, forged documents and camera injection for video calls, aimed squarely at defeating identity checks. That shifts the picture on the attacker side. Anyone who wants to use a convincing voice clone no longer needs to understand the models, only to hold an account with a service of this kind. Rapid7 describes underground pricing as volatile, ranging from entry-level offers costing a few dozen US dollars to several thousand for professional systems. For defenders, it means planning for a broader and less specialised set of adversaries than two years ago.

AI voice cloning scams: the attack patterns inside organisations

An AI voice cloning scam looks different inside a company from the emergency calls reported in consumer cases. Four patterns come up regularly.

CEO fraud with a cloned voice

In CEO fraud, attackers pose as a member of the executive team and push through a payment or a change of bank details. As long as the attack arrived by email, it often failed at a simple counter-check. Finance called the executive on the internally held number and had the instruction confirmed. AI voice cloning devalues exactly that control, because the attack now arrives as a call in the first place. Someone who has already heard what they believe to be the CEO’s voice treats the check as done, even though nobody called back. The rule that still holds is that the verification has to be outbound. Hang up and dial the number held in the internal directory, not the one shown on the display and not the one given over the phone. Scenarios with built-in secrecy make the pattern harder to break, for example an acquisition that is supposedly under way and must not be discussed with anyone.

MFA approvals and help desk fraud

MFA requests usually only succeed once a person approves them, and that is precisely where the phone call comes in. While push notifications are arriving, a voice from what appears to be internal IT asks the recipient to approve the prompt so that the fault can be fixed. This is how a call using deepfake voice cloning turns an MFA fatigue attack into a successful intrusion. The caller supplies a plausible explanation, and the employee clears the request. Help desks are a target for the same reason. Attackers present themselves as staff locked out of an account and press for a quick reset. Organisations handling a high volume of password, VPN and MFA token requests are especially exposed.

Multi-channel attacks: when the call announces the phishing email

The call is not always the attack itself. It is often the preparation for one. Someone who appears to be a known contact announces a message, which then arrives shortly afterwards. The phishing email that follows looks more credible and gets checked less carefully. The same logic works in reverse, where an unremarkable email makes a later call seem legitimate. The difficulty for organisations is that channels are usually monitored in isolation, so an AI voice cloning call and the email that follows it are never connected to each other.

The call about expired SSO access

A further pattern puts AI voice cloning to work on single sign-on access that has supposedly expired. Callers claim to work in IT support, refer to a system migration or a security update, and guide the person towards a login portal or through setting up a second factor again. Because the process sounds technically plausible and the voice sounds familiar, it is rarely questioned. Voice cloning scam calls of this type are aimed at one thing, which is control over an authentication factor.

How to detect AI voices: what to listen for on the call

Ask how to detect an AI voice and the usual answer is technical: spectral analysis, artefact detection, biometric voice recognition. None of that helps during a live call, because nobody starts a forensic analysis mid-conversation. What counts in practice is what a person can actually notice at the time.

Articulation gives the first clues. AI voice cloning sometimes betrays itself through consistently flat intonation, pauses in unnatural places or missing breath sounds. Sentences can be grammatically perfect while carrying no emotional movement at all. Volume that stays exactly level over several minutes can point the same way. Contradictions in the background are worth noticing too, for instance a caller who claims to be driving with no road noise behind them. A faked voice message is even harder to judge than a live call, because there is no chance to ask a spontaneous question. A voice message on its own never confirms who someone is, however familiar the voice sounds.

During the conversation itself, other signals usually tell you more. Systems built on AI voice cloning often respond with a delay when someone interrupts, sidestep unexpected questions, or falter as soon as the conversation leaves the prepared structure. A spontaneous question about a shared detail that is not publicly known can therefore reveal more than any amount of listening to the voice.

Behavioural signals carry the most weight. An unusual channel, an unusual time of day, explicit time pressure, a request for confidentiality, or an instruction that bypasses an established process. Where several of these occur together, the call is worth verifying, no matter how convincing the voice sounds.

The limits of voice cloning detection

Listening alone will not reliably tell you whether a voice is real. A study in the journal PLOS ONE found that participants identified synthetic speech correctly in only around three out of four cases, and that prior training improved the hit rate by less than four percentage points on average. Signals of this kind are useful for raising suspicion, and that is what they are for. Voice cloning detection by ear is not a decision-making instrument, and against deepfake voice cloning, the decision belongs to a defined process rather than to an individual judgement in the moment.

Voice cloning authenticity protocols: proving a voice is genuine

For organisations, the question shifts from detection to evidence. How do you establish that a voice is genuine, and how do you record that the check took place? Voice cloning authenticity protocols answer that as a matter of process rather than tooling, which is where they differ from detection products.

  • Provenance and watermarking: Two complementary approaches are taking shape on the technical side. Provenance metadata under the C2PA standard attaches a signed history to an audio file, documenting how a recording was created and subsequently modified. Watermarking instead embeds a signal in the audio itself that people cannot hear and that is designed to survive conversion. Since 2 August 2026, Article 50 of the EU AI Act (Regulation (EU) 2024/1689) has required providers of generative systems to mark synthetic audio, image, video and text content in machine-readable form. The regulation prescribes no specific technology. It does require implementation that is effective, interoperable, robust and reliable, as far as this is technically feasible. Systems already on the market before that date have a transition period until 2 December 2026. Anyone publishing deepfakes is required to disclose that the content was artificially generated, subject to exemptions and lighter obligations for artistic, satirical and certain law enforcement uses. These rules apply directly across the EU. Organisations elsewhere should expect different labelling requirements, or none at all, and check what applies in their own jurisdiction. For day-to-day defence, this matters less than it might sound. Labelling obligations bind compliant providers, not criminal tools. There is a technical limitation as well. A study on the durability of audio watermarks presented at the Interspeech 2025 conference found that none of the methods examined survived conversion by neural codecs intact, because watermark and codec compete for the same signal space. Data compression and codecs of exactly this kind are standard in voice services, which means provenance evidence tends to fail at the moment it would be needed. It remains valuable for examining a recording after the fact, but not for deciding what to do during a live call.
  • Verification procedures: While a conversation is under way, a defined checking process offers more than any signal analysis. That starts with calling back on an extension held in the company directory rather than the number shown on the display. Control questions whose answers cannot be researched publicly are a second option. Management, finance functions and executive assistants can agree a code word in advance. External requests to the help desk should be verified through a second, independent channel, for example a confirmation from the manager in an internal system. None of these steps depends on recognising AI voice cloning by ear, which is precisely their advantage.
  • Documentation: For these steps to stand up to an audit, they belong in a written procedure. A log recording the time and outcome of each verification makes the practice traceable. Evidence of this kind can feed into management systems aligned with ISO 27001, NIS2 or DORA, the latter two being EU frameworks. The result is a control whose application and effectiveness can be demonstrated, which is what turns voice cloning authenticity protocols from good intentions into something auditable.

How to protect your organisation against AI voice cloning

AI voice cloning cannot be blocked at the gateway. Protection comes from a small number of precise process rules working together with a workforce that recognises the behaviour behind the call and knows what to do about it.

Against AI voice cloning, the callback is one of the most effective measures available, because it takes the voice out of the authentication chain altogether. Any instruction given over the phone that carries financial weight or affects IT security should be checked against a verified number held in internal records. The callback belongs to the normal process rather than being a sign of distrust. Organisations need to say that openly, so that employees do not hesitate when the caller appears to be their own manager.

Payments above a defined threshold, and any change to bank details or master data, require a second approval through a separate channel. The rule that makes this work is that urgency never shortens the process. That shortcut is what every attack of this kind is looking for.

Help desks are a rewarding target for AI voice cloning. Password and MFA factor resets should never be released on the strength of a phone call alone. An additional confirmation through an independent channel helps, as do verification questions whose answers cannot be found publicly. Phishing-resistant authentication such as FIDO2 security keys adds another layer, because the factors involved cannot be read out or passed on over the phone.

Avoiding recorded voice material entirely is unrealistic for most organisations, and usually undesirable. Every publicly available recording is nonetheless potential raw material for deepfake voice cloning. For executives and people in public-facing roles, it is worth mapping systematically where audio of them can be found. Stricter identity checks can then be applied to that group.

A suspicious call should not end when the receiver goes down. Passing it to the security team is what makes it possible to spot the same approach elsewhere in the organisation early. False alarms have to be explicitly welcome, or nobody reports the borderline cases.

Voice cloning simulations in security training

Processes only work if people apply them under pressure, and that can be practised. Standard e-learning modules explain that vishing exists. What they do not create is the reflex for the moment when a familiar voice asks politely for an exception. Recognising an AI-generated call takes practical experience rather than a definition, which is why simulations add a lived situation to the knowledge component.

SoSafe has offered the Interactive Vishing Lesson for this since June 2026. At the end of an e-learning unit on social engineering, participants work through a realistic AI-supported phone scenario directly in the browser. Organisations do not need to collect employees’ phone numbers or call personal devices. That removes the familiar obstacles to traditional vishing campaigns, including the absence of company mobiles, data protection concerns and, where such bodies exist, consultation with employee representatives.

The approach is particularly relevant for organisations where large parts of the workforce work by phone. Logistics, manufacturing, healthcare and public administration are typical examples, as are organisations handling a high volume of help desk requests.

The accompanying security awareness training is built around behavioural change rather than one-off knowledge transfer. The instructional design draws on learning psychology, aligns with NIST, CIS, ISO/IEC 27001, NIS2 and DORA, and can be ready to run in up to two days. Ongoing operation takes from 15 minutes a week. Knowledge retention after twelve months is 90 per cent.

Recognising AI voices when it counts

Give your teams realistic practice against voice cloning and social engineering, in the situation where it actually matters.

Explore Awareness Training

Checklist against AI voice cloning

The right-hand column names functions rather than job titles. In organisations with smaller security teams of three to fifteen people, several of these will sit with the same person.

AreaMeasureOwner
VerificationCall back on the number held in the internal directory for any phone instruction involving money or accessAll employees
VerificationAgree a code word for executives, finance functions and executive assistants, and change it annuallySecurity, executive team
PaymentsRequire two-person approval above a defined amount, with no shortcut for urgencyFinance
PaymentsConfirm changes to bank details through a second, independent channel onlyFinance, procurement
Help DeskNever release password or MFA resets on the strength of a phone call aloneIT service
Help DeskUse verification questions whose answers cannot be researched publiclyIT service
AccessRoll out phishing-resistant MFA such as FIDO2 keys for privileged accountsSecurity
DetectionTrain the signals: flat intonation, background that does not fit, delayed responses to interruptions, time pressure, requests for confidentialitySecurity
ReportingProvide a low-barrier route for reporting suspicious calls, and welcome false alarms explicitlySecurity, communications
TrainingAdd vishing simulation to the training programme, for example through the Interactive Vishing LessonSecurity, HR
EvidenceRecord verification procedures in writing and log the decisions takenCompliance
ReviewCompare attack patterns against current incident reports twice a year and update procedures accordinglySecurity

Do you want to stay ahead of the cyber game?

Sign up for our newsletter to receive the latest cyber security articles, events, and resources. No spam, only content that truly matters.

Newsletter visual
Hero Background

Experience our products first-hand

Use our online test environment to see how our platform can help you empower your team to continuously avert cyber threats and keep your organization secure.

SoSafe Security Awareness Training Leader Enterprise 2026 Sosafe Cyber security training platform top 50 award 2026 SoSafe Security Awareness Training Leader 2026 SoSafe Security Awareness Training Momentum Leader 2026 SoSafe Security Awareness Training Leader Mid-Market 2026 SoSafe Security Awareness Training Leader Europe 2026

This page is not available in English yet.

Diese Seite ist noch nicht in Ihrer Sprache verfügbar. Sie können auf Englisch fortfahren oder zur deutschen Startseite zurückkehren.

Cette page n’est pas encore disponible dans votre langue. Vous pouvez continuer en anglais ou revenir à la page d’accueil en français.

Deze pagina is nog niet beschikbaar in uw taal. U kunt doorgaan in het Engels of terugkeren naar de Nederlandse startpagina.

Esta página aún no está disponible en español. Puedes continuar en inglés o volver a la página de inicio en español.

Questa pagina non è ancora disponibile nella tua lingua. Puoi continuare in inglese oppure tornare alla home page in italiano.