Menu
Home
Forums
New posts
Search forums
What's new
Featured content
New posts
New media
New media comments
New resources
Latest activity
Media
New media
New comments
Search media
Resources
Latest reviews
Search resources
Nyuuz
Jinaral kantent
Log in
Register
What's new
Search
Search
Search titles only
By:
New posts
Search forums
Menu
Log in
Register
Install the app
Install
Home
Forums
Labrish
Nalij
Jinaral kantent
Why does sung training data change a voice model
JavaScript is disabled. For a better experience, please enable JavaScript in your browser before proceeding.
You are using an out of date browser. It may not display this or other websites correctly.
You should upgrade or use an
alternative browser
.
Reply to thread
Message
[QUOTE="Bombastus, post: 92132, member: 2178"] ElevenLabs currently requires spoken recordings for Professional Voice Cloning and says those clones do not support singing. The restriction matters because a clean speech dataset can describe your vocal identity in great detail without showing how you handle a melody. A voice model learns from the behavior present in its material, not from a checklist of abilities a human speaker might possess. Spoken samples contain timbre, accent, cadence, pronunciation, and ordinary pitch movement, while sung recordings add sustained vowels, controlled note targets, larger duration changes, register shifts, vibrato, and performance dynamics. None of this means every singing system must train on songs from the target person. The more useful distinction behind [B][URL='https://goldmidi.com/community/threads/personalized-vocals-do-not-mean-artist-voice-cloning.77087/']why spoken voice cloning differs from singing models[/URL][/B] is where the system learns vocal identity and where it learns the mechanics of singing. [HEADING=2]Sung recordings expose behavior that speech barely visits[/HEADING] Speech and singing use the same physical voice, but they occupy that voice differently. Parallel recordings of people speaking and singing the same lyrics show substantial differences in pitch, energy, duration, and spectral behavior. Sung vowels also tend to stretch much longer because they carry notes instead of simply connecting consonants inside ordinary speech. Duration alone changes the training problem. A model hearing spoken “me” may encounter a short vowel followed almost immediately by the next word, while a singer can hold the vowel for several seconds and alter pitch, loudness, and tone throughout it. Training on sung material gives the system direct examples of what a voice does during those long, controlled states. Pitch coverage changes too. Normal conversation wanders through fundamental frequency according to language and emotion, whereas singing asks for specific note centers, transitions, bends, and sustained contours. A dataset containing actual singing therefore gives a model evidence about how the target voice behaves across musical pitch ranges instead of forcing it to extrapolate from conversation. Register is where the gap becomes especially obvious. Someone can speak comfortably inside a narrow part of their range but sing far above or below it using chest voice, head voice, falsetto, or mixed production. Hours of speech can still leave those parts of the voice completely undocumented. [HEADING=2]Singing knowledge can come from somebody else[/HEADING] A target singer does not always need to provide sung training material for a system to produce singing in their timbre. The model can learn the mechanics of singing from one group of voices and learn a new person's identity from speech, provided its architecture is designed to keep those jobs apart. The 2019 [B][URL='https://arxiv.org/abs/1912.10128']Learning Singing From Speech[/URL][/B] system used normal speech samples to represent a target speaker while learning singing behavior inside a unified speech-and-singing framework. Its speaker representation could transfer between the two domains, letting the system apply a speech-derived identity to singing learned elsewhere. A later approach called Learn2Sing pushed the same idea further with a singing teacher and speech from target speakers. It added explicit pitch and duration prediction because target speech alone does not provide the musical timing needed at generation time. The clever part was not making speech secretly contain singing. It was supplying the missing singing structure from another source. This distinction changes how you should read claims about training data. A company could build a singing voice for somebody without training the entire singing behavior on that person's songs, while another system could rely heavily on their sung recordings. Knowing that a model sounds like an artist does not tell you which route produced it. [HEADING=2]Dataset design decides what the model can reproduce[/HEADING] More audio is useful only when it expands the behavior the system needs. Adding another hour of calm narration can improve a speech clone while contributing almost nothing about belting, breathy high notes, vibrato, melisma, or sustained dynamics. Variety has to be relevant, not merely abundant. Sung material also needs coverage rather than random song fragments. A dataset concentrated in one comfortable octave may reproduce that register well and struggle elsewhere. Material spanning different vowels, note lengths, intensities, registers, consonant attacks, and pitch transitions gives the model more evidence about how the same identity behaves under changing musical demands. Recording conditions can quietly become part of the problem. Reverb, backing vocals, doubles, instrumental bleed, heavy tuning, and aggressive processing can make it harder to isolate the singer's actual characteristics. A polished master may sound better to a listener while being less informative than a dry vocal recording for learning a specific voice. Speech data still has a valuable role because identity and singing technique are not the same information. Clean speech can provide abundant evidence about timbre and pronunciation, while dedicated singing data or a separate singing model supplies musical behavior. The strongest architecture depends on which parts are learned globally and which parts must remain specific to the individual voice. [/QUOTE]
Insert quotes…
Name
Post reply
Home
Forums
Labrish
Nalij
Jinaral kantent
Why does sung training data change a voice model
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.
Accept
Learn more…
Top