ElevenLabs currently says its Professional Voice Clones do not support singing, and the audio used to build them must contain spoken voice only. The restriction is easy to mistake for a product switch waiting to be flipped, but speech cloning and convincing singing ask a model to solve different problems.
A speech clone mainly needs to preserve who you sound like while producing intelligible words with believable cadence, accent, pronunciation, and expression. Singing keeps identity in the picture, then adds melody, sustained notes, tightly controlled timing, wider pitch movement, vibrato, register changes, and vocal technique. The same product boundary sits behind the difference between personalized vocals and voice cloning because recognizable timbre alone does not make a model a singer.
Duration becomes more demanding too. A spoken vowel might last only as long as natural phrasing requires, while a singer can stretch the same vowel across a bar and change its pitch or intensity while holding it. Consonants still need to arrive at musically useful moments, so a system that sounds fluent when reading a sentence can sound stiff or garbled when phonemes have to fit a score.
High notes add another complication. Fundamental frequency and the resonances that shape vowel identity can interact differently as pitch rises, which is one reason a sung vowel is not merely a spoken vowel moved upward with a pitch shifter. The model has to preserve intelligibility and vocal identity while handling an acoustic range its speech behavior may barely cover. At higher fundamental frequencies, wider harmonic spacing also makes the spectral cues behind vowel quality harder to sample cleanly.
More speech data does not automatically fix the mismatch. Three hours of excellent talking can describe a speaker in far greater detail while still containing little evidence about how the same person sings an octave higher or shapes a sustained note. Extra speech may improve identity fidelity without teaching the model the missing singing behaviors.
Speech recordings are not useless for singing systems, though. Cross-domain singing conversion from ordinary speech data has been demonstrated with an architecture built specifically to bridge speech and singing, using a target speaker's normal speech while preserving singing content from a source performance. The important part is the specialized conversion system around the speech data, not an assumption that an ordinary text-to-speech clone can suddenly perform a song.
Even this separation is imperfect. Timbre can leak into content representations, pitch errors can alter the musical line, and loudness or phonation can drag style information along with them. Recent systems devote explicit machinery to vibrato because natural singing contains fast pitch fluctuations that broad note-level control can miss.
Vibrato makes the distinction especially clear. Its rate and extent live inside the pitch contour, so a model can hit the correct note and still produce flat or unstable singing if it mishandles those movements. A 2025 singing voice conversion system used wavelet-based processing to isolate higher-frequency F0 movement for more direct vibrato control.
Speech cloning does not normally need that level of musical control. A clone can sound convincingly like you while reading a new sentence because the synthesis system only has to generate plausible speech inside the behavior it was built to handle. Singing asks the system to preserve identity while obeying a melodic performance, and the architecture has to represent both jobs without letting one corrupt the other.
A speech clone mainly needs to preserve who you sound like while producing intelligible words with believable cadence, accent, pronunciation, and expression. Singing keeps identity in the picture, then adds melody, sustained notes, tightly controlled timing, wider pitch movement, vibrato, register changes, and vocal technique. The same product boundary sits behind the difference between personalized vocals and voice cloning because recognizable timbre alone does not make a model a singer.
Singing asks a model to control more than identity
A convincing clone can capture the spectral traits that make your speaking voice recognizable and still fall apart on a melody. Spoken prosody moves pitch constantly, but it does not normally ask the system to land on an intended musical note, hold it for a precise duration, bend into another note, then add controlled vibrato without smearing the words. Singing turns pitch from a supporting feature into part of the performance itself.Duration becomes more demanding too. A spoken vowel might last only as long as natural phrasing requires, while a singer can stretch the same vowel across a bar and change its pitch or intensity while holding it. Consonants still need to arrive at musically useful moments, so a system that sounds fluent when reading a sentence can sound stiff or garbled when phonemes have to fit a score.
High notes add another complication. Fundamental frequency and the resonances that shape vowel identity can interact differently as pitch rises, which is one reason a sung vowel is not merely a spoken vowel moved upward with a pitch shifter. The model has to preserve intelligibility and vocal identity while handling an acoustic range its speech behavior may barely cover. At higher fundamental frequencies, wider harmonic spacing also makes the spectral cues behind vowel quality harder to sample cleanly.
Speech-trained clones inherit a narrower performance range
Training material teaches a clone more than the basic color of a voice. ElevenLabs tells users that Professional Voice Cloning reproduces characteristics and stylistic tendencies from its samples, and it recommends consistent material that matches the delivery you want. A clean narration dataset therefore gives the model strong evidence about narration, not chest-to-head transitions, belting, falsetto, long sustained vowels, scoops, or vibrato.More speech data does not automatically fix the mismatch. Three hours of excellent talking can describe a speaker in far greater detail while still containing little evidence about how the same person sings an octave higher or shapes a sustained note. Extra speech may improve identity fidelity without teaching the model the missing singing behaviors.
Speech recordings are not useless for singing systems, though. Cross-domain singing conversion from ordinary speech data has been demonstrated with an architecture built specifically to bridge speech and singing, using a target speaker's normal speech while preserving singing content from a source performance. The important part is the specialized conversion system around the speech data, not an assumption that an ordinary text-to-speech clone can suddenly perform a song.
Singing conversion separates pitch from vocal identity
Modern singing voice conversion often treats several parts of a performance separately because they need different control. One stream can represent linguistic content, another can carry fundamental frequency, another can describe loudness, and another can steer the target singer's timbre. Keeping those pieces apart lets the model change who the vocal sounds like without casually rewriting the melody.Even this separation is imperfect. Timbre can leak into content representations, pitch errors can alter the musical line, and loudness or phonation can drag style information along with them. Recent systems devote explicit machinery to vibrato because natural singing contains fast pitch fluctuations that broad note-level control can miss.
Vibrato makes the distinction especially clear. Its rate and extent live inside the pitch contour, so a model can hit the correct note and still produce flat or unstable singing if it mishandles those movements. A 2025 singing voice conversion system used wavelet-based processing to isolate higher-frequency F0 movement for more direct vibrato control.
Speech cloning does not normally need that level of musical control. A clone can sound convincingly like you while reading a new sentence because the synthesis system only has to generate plausible speech inside the behavior it was built to handle. Singing asks the system to preserve identity while obeying a melodic performance, and the architecture has to represent both jobs without letting one corrupt the other.