ElevenLabs currently makes every Professional Voice Clone pass a spoken verification step before fine-tuning can begin. The voice owner reads a generated CAPTCHA aloud, and the service checks that recording against the uploaded training voice.
Runway takes a similar route with a different policy boundary. Someone can let you train a custom voice from their recordings, but they have to record the platform’s consent script themselves, in the app, using the same voice as the sample. No live participation, no verified clone.
Both systems are trying to stop the obvious abuse case of uploading somebody else’s recordings and ticking a box. Useful safeguard. Still, a successful match proves something narrower than people often assume.
The check does not magically read a contract sitting in somebody’s inbox. It cannot know whether a singer approved one campaign, one language, one album cycle, or every future use of the model. It also cannot infer whether permission expires next month, excludes political material, or allows an agency to hand access to a subcontractor.
This is where artist voice cloning and personalized vocals need to stay separate in your head. A platform can verify who supplied a voice without proving that every feature built around the resulting model carries the same permission.
Speaker matching has limits of its own. A 2026 evaluation of speaker verification under synthetic attack found that modern cloned speech could bypass commercial speaker-verification systems, while anti-spoofing tools struggled to generalize across unfamiliar synthesis methods. Voice checks are useful friction, not a magic truth machine.
Runway handles the same problem differently. It allows another person’s voice to become a custom voice when the person records the required consent task directly inside Runway. The service compares the consent voice with the submitted sample, so a producer cannot simply claim permission on somebody else’s behalf and move on.
Resemble AI also uses a recorded consent clip and speaker matching. Descript requires a consenting speaker and uses authorization audio for its cloned voices. Speechify’s current terms take another route by allowing the speaker or an authorized representative to attest that the necessary permission exists.
None of those approaches is automatically absurd. They just protect different points in the workflow. One platform may demand direct participation from the speaker, another may combine voice matching with account controls, while another puts more weight on contractual attestations from the person uploading the material.
The difference matters for artists because “verified” can sound more comprehensive than it is. A green badge may mean the system matched two recordings. It may mean the speaker completed a consent script. It may mean an uploader accepted legal terms. Those are not interchangeable facts.
Nothing about the original voice match proves the advertisement was approved. The verification succeeded because the singer really did participate at creation time. The later problem is scope, access, and permission, not identity.
The same issue appears when a voice changes hands inside a team. An artist may approve one producer, one workspace, or one application while rejecting public sharing. Technical access controls can help enforce those boundaries, but they need the permission rules first. Software cannot enforce a limit nobody bothered to define.
Revocation makes the gap even clearer. A person can genuinely authorize a clone today and withdraw future access later. The historical verification event stays true, while the current permission changes. Treating old verification as permanent consent turns a useful security check into a paper shield.
For a fan-facing music product, the clean model is layered. Verify the speaker at enrollment, record what they actually approved, restrict who can invoke the voice, keep later uses inside those limits, and make withdrawal operational rather than ceremonial. Identity comes first. Permission still has to survive everything that happens after it.
Runway takes a similar route with a different policy boundary. Someone can let you train a custom voice from their recordings, but they have to record the platform’s consent script themselves, in the app, using the same voice as the sample. No live participation, no verified clone.
Both systems are trying to stop the obvious abuse case of uploading somebody else’s recordings and ticking a box. Useful safeguard. Still, a successful match proves something narrower than people often assume.
Verification checks enrollment, not the whole relationship
Voice verification is basically an enrollment gate. It asks whether the person speaking during the check appears to be the same person represented in the material being submitted for cloning. ElevenLabs describes its voice CAPTCHA as a safeguard rather than a technical requirement for synthesis, which is an important distinction.The check does not magically read a contract sitting in somebody’s inbox. It cannot know whether a singer approved one campaign, one language, one album cycle, or every future use of the model. It also cannot infer whether permission expires next month, excludes political material, or allows an agency to hand access to a subcontractor.
This is where artist voice cloning and personalized vocals need to stay separate in your head. A platform can verify who supplied a voice without proving that every feature built around the resulting model carries the same permission.
Speaker matching has limits of its own. A 2026 evaluation of speaker verification under synthetic attack found that modern cloned speech could bypass commercial speaker-verification systems, while anti-spoofing tools struggled to generalize across unfamiliar synthesis methods. Voice checks are useful friction, not a magic truth machine.
Platforms draw the consent line in different places
ElevenLabs uses one of the stricter product rules for its Professional Voice Clone. You can only create your own PVC. Even permission from another person is not enough for you to build it on your account. The owner has to create and verify the clone themselves, then share access afterward.Runway handles the same problem differently. It allows another person’s voice to become a custom voice when the person records the required consent task directly inside Runway. The service compares the consent voice with the submitted sample, so a producer cannot simply claim permission on somebody else’s behalf and move on.
Resemble AI also uses a recorded consent clip and speaker matching. Descript requires a consenting speaker and uses authorization audio for its cloned voices. Speechify’s current terms take another route by allowing the speaker or an authorized representative to attest that the necessary permission exists.
None of those approaches is automatically absurd. They just protect different points in the workflow. One platform may demand direct participation from the speaker, another may combine voice matching with account controls, while another puts more weight on contractual attestations from the person uploading the material.
The difference matters for artists because “verified” can sound more comprehensive than it is. A green badge may mean the system matched two recordings. It may mean the speaker completed a consent script. It may mean an uploader accepted legal terms. Those are not interchangeable facts.
A matched voice can still sit inside the wrong use
Imagine a singer agrees to create a clone for clean demo vocals during a six-month production deal. The enrollment is genuine, the speaker passes every verification check, and the model works exactly as promised. A year later, somebody still having access uses it for an unrelated advertisement.Nothing about the original voice match proves the advertisement was approved. The verification succeeded because the singer really did participate at creation time. The later problem is scope, access, and permission, not identity.
The same issue appears when a voice changes hands inside a team. An artist may approve one producer, one workspace, or one application while rejecting public sharing. Technical access controls can help enforce those boundaries, but they need the permission rules first. Software cannot enforce a limit nobody bothered to define.
Revocation makes the gap even clearer. A person can genuinely authorize a clone today and withdraw future access later. The historical verification event stays true, while the current permission changes. Treating old verification as permanent consent turns a useful security check into a paper shield.
For a fan-facing music product, the clean model is layered. Verify the speaker at enrollment, record what they actually approved, restrict who can invoke the voice, keep later uses inside those limits, and make withdrawal operational rather than ceremonial. Identity comes first. Permission still has to survive everything that happens after it.