Menu
Home
Forums
New posts
Search forums
What's new
Featured content
New posts
New media
New media comments
New resources
Latest activity
Media
New media
New comments
Search media
Resources
Latest reviews
Search resources
Nyuuz
Jinaral kantent
Log in
Register
What's new
Search
Search
Search titles only
By:
New posts
Search forums
Menu
Log in
Register
Install the app
Install
Home
Forums
Labrish
Nalij
Jinaral kantent
Reference audio can steer AI music without training
JavaScript is disabled. For a better experience, please enable JavaScript in your browser before proceeding.
You are using an out of date browser. It may not display this or other websites correctly.
You should upgrade or use an
alternative browser
.
Reply to thread
Message
[QUOTE="Bombastus, post: 92398, member: 2178"] ElevenLabs currently lets Music v2 and v2.5 use a short uploaded track to guide sound, instrumentation, tempo, mood, and production style. The company describes the result as a new composition rather than a remix of the uploaded recording. A reference track can therefore matter to one generation without becoming part of the model’s original training set. People often collapse those two events because both involve feeding audio into an AI system, but the technical jobs are different. Training changes a model through repeated optimization over data. Generation uses an already trained model to produce an output from whatever conditions it receives at that moment, which is why [B][URL='https://goldmidi.com/community/threads/umg-has-not-disclosed-an-elevenlabs-training-license.77086/']public claims about AI music training[/URL][/B] need more precision than a statement saying licensed music is involved. [HEADING=2]Conditioning happens during generation[/HEADING] Audio conditioning gives the model another source of direction alongside text. Instead of describing a hazy drum texture, loose tempo, dark synth palette, or dry vocal treatment in words, you can supply audio that already carries some of those qualities. The system does not need to rewrite all its learned weights every time you upload a reference. A conditioning pathway can encode features from the supplied audio, pass those features into the generation process, and let the existing model respond while it creates a fresh result. In [B][URL='https://arxiv.org/abs/2407.12563']audio conditioning for music generation[/URL][/B], text and audio controls can be combined at inference time, including a setup where audio is converted into conditioning information rather than treated as new material for retraining the generator. The important distinction is timing. The reference arrives when the output is being made, while the model’s underlying learned parameters can remain fixed. Current ElevenLabs documentation makes the same product-level separation visible. Audio Reference sits beside prompts and Music Finetunes as a generation control, while custom finetuning is presented as a separate workflow that trains a private adaptation on supplied music. [HEADING=2]Training and generation leave different footprints[/HEADING] A training dataset is used to alter a model over many optimization steps. The effect gets folded into parameters that survive after the individual training job ends, which means later users can benefit from whatever patterns the system learned without supplying the original files again. A reference input has a shorter job. It can shape one request, one session, or one generated track, depending on the product design, then disappear from the immediate generation path once the output is finished. No weight update is required for the reference to influence the result. This distinction also explains why hearing resemblance in an output does not prove how a particular recording entered the system. Similarity could come from broad patterns learned during training, a text prompt, a reference supplied during generation, a dedicated fine-tune, or several controls working together. One more wrinkle matters. A service can use an upload as inference-time conditioning and still have separate contractual rules about retaining submitted data or using future submissions to improve models. ElevenLabs currently lets users change a data-use setting that controls whether new submitted data may be used for model improvement, so the generation mechanism and the provider’s later data policy should be checked separately. [HEADING=2]Reference control has real limits[/HEADING] Reference audio is guidance, not a perfect transfer switch. ElevenLabs says its feature is meant to influence overall sound, production style, instrumentation, tempo, and mood, and it warns that large genre jumps may produce less reliable results. Academic work points in the same practical direction. Audio conditioning can improve control over characteristics that plain text struggles to specify, but different architectures decide which features are extracted, how strongly they are applied, and how they interact with the prompt. You can hear a reference strongly reflected in groove or palette while melody moves somewhere else. Another system may preserve timing more closely but loosen timbre, and a third may treat the same reference as a broad stylistic hint. None of those behaviors require the reference recording to become permanent training material. They only show that the generator received audio as part of the instruction set for the current creation process. For licensing discussions, the cleanest language is also the least dramatic. A track used to condition generation has influenced an output. A track used to train or fine-tune a model has helped alter the model itself, and proving one does not automatically prove the other. A platform can support both workflows under one brand, so the product name alone cannot settle which process used a particular recording. [/QUOTE]
Insert quotes…
Name
Post reply
Home
Forums
Labrish
Nalij
Jinaral kantent
Reference audio can steer AI music without training
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.
Accept
Learn more…
Top