Menu
Home
Forums
New posts
Search forums
What's new
Featured content
New posts
New media
New media comments
New resources
Latest activity
Media
New media
New comments
Search media
Resources
Latest reviews
Search resources
Nyuuz
Jinaral kantent
Log in
Register
What's new
Search
Search
Search titles only
By:
New posts
Search forums
Menu
Log in
Register
Install the app
Install
Home
Forums
Labrish
Nalij
Jinaral kantent
Training data transparency still leaves blind spots
JavaScript is disabled. For a better experience, please enable JavaScript in your browser before proceeding.
You are using an out of date browser. It may not display this or other websites correctly.
You should upgrade or use an
alternative browser
.
Reply to thread
Message
[QUOTE="Bombastus, post: 92403, member: 2178"] The EU’s public training-content summary requirement has applied to general-purpose AI models placed on its market since August 2, 2025. California followed with a separate rule that requires qualifying generative AI developers to publish training-data documentation by January 1, 2026, and before later covered releases. Neither regime turns a developer’s website into a song-by-song training ledger. They force more information into public view, but the level of detail still depends on the rule, the type of model, the source of the data, and how the developer assembled it. For music, the useful distinction is between disclosure and proof. A public summary can narrow the possibilities around [B][URL='https://goldmidi.com/community/threads/umg-has-not-disclosed-an-elevenlabs-training-license.77086/']what public licensing language actually establishes[/URL][/B], while a rights holder may still need contracts, dataset records, technical logs, or litigation discovery to establish whether one recording entered one model. [HEADING=2]Public summaries reveal categories rather than every track[/HEADING] The EU AI Act requires providers of covered general-purpose AI models to publish a sufficiently detailed summary of the content used for training. The Commission’s mandatory template asks for broad modality information, public and private datasets, scraped sources, user data, synthetic data, and relevant processing details. For web-scraped material, the template goes further than a vague statement about using public internet data. Providers must identify crawlers, collection periods, describe the scraped content, and list the top slice of domains represented in the crawl, while large public datasets also have to be identified. Useful, but still not a catalog dump. A rights holder cannot assume the summary will name every recording, artist, songwriter, ISRC, URL, or file that passed through a large corpus. California’s AB 2013 takes a different route. Its disclosure covers sources or owners of datasets, approximate dataset size, data types, whether protected material appears, whether datasets were purchased or licensed, processing steps, collection periods, first-use dates, personal information, and synthetic data. The law expressly describes the required dataset description as high-level. It can tell you a developer used copyrighted audio from licensed sources during a particular period without necessarily telling you whether your track was one of the files. [HEADING=2]Fine-tunes make disclosure history more complicated[/HEADING] Both rules recognize that training does not end with the first model release. California defines training broadly enough to include testing, validation, and fine-tuning, while a substantial modification can include retraining or fine-tuning that materially changes a system’s functionality or performance. The EU template also expects updates after further training. Providers should revise a summary at six-month intervals or sooner when additional data materially changes what the summary needs to say, and modified models can point back to an earlier base-model summary while documenting the new training content used for the modification. This creates a chain rather than one permanent disclosure page. A music model can begin with one corpus, receive a later adaptation, then spawn another version with a different fine-tune. Anyone tracing provenance has to follow the versions instead of treating the brand name as one frozen model. A 2025 framework for [B][URL='https://arxiv.org/abs/2503.18814']responsible AI music transparency[/URL][/B] treats transparency, data governance, explainability, and accountability as connected design problems rather than one disclosure checkbox. Music systems make the problem especially visible because recordings, compositions, performers, metadata, prompts, references, and model versions can all carry different rights histories. [HEADING=2]Transparency laws still stop short of verification[/HEADING] The EU rule is not a universal disclosure law for every music generator. Article 53 targets providers of general-purpose AI models, so a specialized system can raise a scope question before anyone reaches the contents of its training summary. California reaches generative AI systems and services made available to Californians more broadly, but its required documentation still centers on dataset-level information. Neither framework automatically gives an artist direct access to a private corpus, model weights, optimization logs, or a complete work-level manifest. Trade-secret protections also matter in Europe. The Commission built different disclosure depths into its template depending on the source, explicitly balancing public transparency against confidential business information. For a musician or label, these rules are therefore best read as better starting evidence. A summary can identify a licensed dataset, disclose that copyrighted audio was used, reveal collection dates, or show that a later fine-tune introduced new material. Those details can turn a vague suspicion into a much narrower records request. They still cannot substitute for evidence tying one protected work to one training run. Public transparency can show the shape of the pipeline. Establishing the exact contents of the pipe may still require dataset manifests, matching records, contracts, audit access, or court-ordered discovery. [/QUOTE]
Insert quotes…
Name
Post reply
Home
Forums
Labrish
Nalij
Jinaral kantent
Training data transparency still leaves blind spots
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.
Accept
Learn more…
Top