Menu
Home
Forums
New posts
Search forums
What's new
Featured content
New posts
New media
New media comments
New resources
Latest activity
Media
New media
New comments
Search media
Resources
Latest reviews
Search resources
Nyuuz
Jinaral kantent
Log in
Register
What's new
Search
Search
Search titles only
By:
New posts
Search forums
Menu
Log in
Register
Install the app
Install
Home
Forums
Labrish
Nalij
Jinaral kantent
AI singing is harder to detect than cloned speech
JavaScript is disabled. For a better experience, please enable JavaScript in your browser before proceeding.
You are using an out of date browser. It may not display this or other websites correctly.
You should upgrade or use an
alternative browser
.
Reply to thread
Message
[QUOTE="Shamiso, post: 92455, member: 160"] Four speech deepfake detectors tested on the SingFake dataset performed markedly worse on singing than on the speech data they were built for. Singing stretches vowels, locks pitch to melody, changes timing, and usually arrives buried under drums, bass, effects, and mastering. A detector trained mainly on spoken sentences can lose its footing fast. The problem gets messier because an altered vocal is not automatically a cloned singer. Pitch correction, time stretching, formant shifting, stem separation, doubling, denoising, and ordinary production can all leave traces that resemble the artifacts a detector might associate with synthetic audio. One suspicious patch in a spectrogram does not settle who sang the line. Detection is therefore less about finding one robot giveaway and more about testing the recording in the right domain. The same distinction matters when people talk about [B][URL='https://goldmidi.com/community/threads/personalized-vocals-do-not-mean-artist-voice-cloning.77087/']synthetic artist vocal identity[/URL][/B] because a generated voice, a converted performance, and a heavily edited real take can produce different forensic clues. [HEADING=2]Singing breaks the assumptions speech detectors rely on[/HEADING] Speech detectors learn patterns from conversational audio where words move through relatively short vowels, ordinary pitch ranges, and familiar pauses. Singing does almost the opposite. Notes can hang for seconds, vibrato creates regular pitch movement, consonants get squeezed around beats, and the singer may jump registers within one phrase. Background music makes those differences harder to inspect. Cymbals, guitars, synths, harmonies, reverb, and limiting can mask tiny synthesis artifacts, while lossy streaming codecs can smear high-frequency information a detector relied on during training. A model that looks sharp on clean laboratory vocals can become much less convincing after the song has been mastered, uploaded, transcoded, downloaded, and clipped for social media. Domain-specific training changes the picture dramatically. A 2026 [B][URL='https://aclanthology.org/2026.findings-acl.1245/']speech and singing deepfake detector[/URL][/B] reached a 1.82 percent equal error rate on a controlled singing benchmark, while speech-trained systems tested in the same work landed between 37 and 62 percent. The gap is ugly enough to kill the idea that any decent speech detector can simply be pointed at a song. Unseen generators remain another problem. A detector can learn fingerprints from the synthesis systems represented in its training set and then stumble when a newer model produces artifacts in different places. Strong performance against yesterday's generator is useful evidence, not a lifetime warranty. [HEADING=2]A clean vocal stem can still mislead the detector[/HEADING] Pulling the singer out of the mix sounds like the obvious first move, and sometimes it helps. Less accompaniment gives a detector more direct access to breath noise, consonants, pitch transitions, harmonic structure, and other vocal details that the instrumental was covering. Singing-specific datasets have shown clear gains when systems actually train on vocal material instead of relying on speech alone. Stem separation also changes the evidence. Separation models can leave watery tails, missing harmonics, phasing, doubled consonants, or bits of accompaniment inside the extracted vocal. A detector may then react to artifacts created by the separator rather than the original singer or generator. For serious checking, the full mix and the isolated stem should be treated as two views of the same recording, not as interchangeable evidence. Agreement between both is more interesting than one dramatic score from a processed stem, especially when the source file has already been compressed several times. Production history matters too. Auto-Tune, Melodyne-style correction, aggressive timing edits, vocal resynthesis, and formant processing can create locally unnatural regions inside an otherwise human performance. A binary real-or-fake label throws all of those possibilities into one bucket and can make normal studio work look more suspicious than it is. [HEADING=2]Better systems point to the altered section[/HEADING] Newer singing-forensics systems are moving beyond a single score for the whole file. They can identify where a manipulation occurs and separate categories such as pitch correction, pitch shifting, time stretching, and deepfake generation. This is far more useful when only one chorus, word, harmony, or replacement line has been altered. Localization also changes how a questionable result can be checked. If the system repeatedly flags the same two-second phrase, you can compare that section with alternate masters, raw vocal takes, stems, live recordings, or earlier exports instead of arguing about an opaque 73 percent probability attached to the entire song. A detector score still cannot establish consent or ownership. A genuine AI vocal may be fully authorized, while an unauthorized imitation may use ordinary human singing and voice conversion rather than end-to-end synthesis. Audio forensics can help establish how a performance was made, but permission has to come from records outside the waveform. The strongest evidence usually comes from several independent signals agreeing. A singing-aware detector, the untouched source file, known generation or editing history, provenance metadata, and consistent localization all answer different parts of the problem. When those records disagree, the disagreement itself matters because a clean vocal score cannot explain a missing source session or a suspiciously localized replacement line. [/QUOTE]
Insert quotes…
Name
Post reply
Home
Forums
Labrish
Nalij
Jinaral kantent
AI singing is harder to detect than cloned speech
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.
Accept
Learn more…
Top