Menu
Home
Forums
New posts
Search forums
What's new
Featured content
New posts
New media
New media comments
New resources
Latest activity
Media
New media
New comments
Search media
Resources
Latest reviews
Search resources
Nyuuz
Jinaral kantent
Log in
Register
What's new
Search
Search
Search titles only
By:
New posts
Search forums
Menu
Log in
Register
Install the app
Install
Home
Forums
Labrish
Nalij
Jinaral kantent
Suno disclosed user data in post-training datasets
JavaScript is disabled. For a better experience, please enable JavaScript in your browser before proceeding.
You are using an out of date browser. It may not display this or other websites correctly.
You should upgrade or use an
alternative browser
.
Reply to thread
Message
[QUOTE="Shamiso, post: 92751, member: 160"] Suno’s California AI disclosure states that certain post-training datasets may contain user Content and User Activity Information. The document has carried that disclosure since January 2026, so user-derived data was already part of Suno’s public training story months before v6 arrived. The wording matters because it is narrower than saying every Suno song was dumped into the main training corpus. Suno separates a large collection of public music used for training from certain post-training datasets that may contain material and activity generated through the service. The distinction makes the January disclosure useful, but easy to overread. It confirms user information can enter model development while leaving open which creations were selected, how often they were used, and what role each dataset played in a particular model. [HEADING=2]Suno separates base training from post-training data[/HEADING] Suno describes its music-model training data as tens of millions of public music audio files with related textual metadata. It says those files help teach models what genres and types of music sound like, while the collection has been running since spring 2023 and remains ongoing as new versions are developed. User data appears in a different sentence. Certain post-training datasets may include Content under Suno’s Terms of Service and User Activity Information under its Privacy Notice. Suno also says it uses both categories to improve its models, with available volume tied to the number of outputs generated through the service at a given time. Post-training can cover several ways of steering a model after its broad capabilities already exist. Preference information is one obvious example because a system can learn which outputs people favor without treating each preference signal as another raw recording in a base audio corpus. [B][URL='https://arxiv.org/abs/2305.18290']Direct Preference Optimization[/URL][/B] shows the broader technical idea clearly, using preference comparisons to fine-tune model behavior after pretraining. Suno has not said its music models use that exact method. The important point is simpler. “Post-training” is not a synonym for “we retrained the base model on every song users made,” and the disclosure does not provide enough detail to make that leap. [HEADING=2]The disclosure gives useful detail without naming songs[/HEADING] The California document goes further than a vague statement about user feedback. It says Suno’s training datasets contain Content and User Activity Information, and that structured identifying information such as usernames is removed before this information is used for training. It also says Suno uses synthetic data in developing its generative AI systems. None of those statements tells you whether your specific track, prompt, like, skip, preference, or generation entered a particular v6 dataset. The disclosure operates at dataset-category level rather than as a searchable record of individual contributions. For anyone trying to pin down whether Suno uses user songs for training, this distinction matters. A category-level disclosure establishes that user-derived material can be in the pipeline without identifying every song, account, model version, or training run involved. This is where [B][URL='https://goldmidi.com/community/threads/suno-has-said-v6-was-trained-on-user-creations.78133/']Suno v6 learning from user creations[/URL][/B] adds something genuinely new. The September statement ties community creations directly to v6, while the January California disclosure had already established the broader possibility of user Content and activity being used in model development. The two disclosures therefore answer different questions. January told users what categories of data may appear in Suno’s training and post-training pipeline. September gave a more model-specific description of what helped develop v6. [HEADING=2]User data does not prove every v6 training claim[/HEADING] A disclosure saying user Content can be used for model improvement does not prove which individual files trained v6, just as removing usernames does not mean the underlying content becomes irrelevant to training. It tells you what Suno says the pipeline can contain, not the complete lineage of one released model. The previously published [URL='https://goldmidi.com/community/threads/private-suno-songs-are-not-a-training-opt-out.78162/']private-song training opt-out breakdown[/URL] matters here because visibility is another separate layer. A song being Link Only concerns who can discover it through Suno, while the California document discusses categories of information that may be used in model development. Likewise, the [URL='https://goldmidi.com/community/threads/suno-paid-song-ownership-still-leaves-a-training-license.78160/']paid ownership and training-license split[/URL] deals with contractual permission rather than dataset provenance. Suno can disclose that Content appears in post-training datasets without the disclosure itself telling you which contractual clause applied to each item or whether any particular output was selected. The clean reading is fairly narrow. Suno publicly acknowledged before v6 that user Content and User Activity Information can enter its model-development pipeline, including certain post-training datasets. The later v6 statement adds specificity around community creations, but neither document gives users an itemized history showing whether one particular song was used, how heavily it mattered, or exactly which training stage touched it. [/QUOTE]
Insert quotes…
Name
Post reply
Home
Forums
Labrish
Nalij
Jinaral kantent
Suno disclosed user data in post-training datasets
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.
Accept
Learn more…
Top