|
|
Log in / Subscribe / Register

Copyright license? I thought the LLM community didn't believe in those...

Copyright license? I thought the LLM community didn't believe in those...

Posted Jun 7, 2026 16:22 UTC (Sun) by parodper (guest, #179096)
In reply to: Copyright license? I thought the LLM community didn't believe in those... by bluca
Parent article: MOT: a tool to fight openwashing in AI

> your copyright or your license are irrelevant w.r.t. training models, as training is specifically allowed to ingest any publicly accessible dataset of any kind, provided its owners did not explicitly opt out

Are you talking about Article 4 of the Directive 2019/790? That talks about «text and data mining» which as defined, doesn't include text generation:

> ‘text and data mining’ means any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations;

As an article[1] notes:

> Finally, it is to be noted that the [text and data mining (TDM)] output should not infringe any exclusive rights as it merely reports on the results of the TDM quantitative analysis, typically not including parts or extracts of the mined materials.

---

> On the other hand, the work they put out _is_ protected by copyright by default, so of course they do have standing to enforce it.

Source on that? As far as I can tell, weights aren't copyrightable.

---

[1] https://www.europarl.europa.eu/RegData/etudes/IDAN/2018/6...(2018)604941_EN.pdf


to post comments

Copyright license? I thought the LLM community didn't believe in those...

Posted Jun 7, 2026 16:39 UTC (Sun) by bluca (subscriber, #118303) [Link] (1 responses)

> Are you talking about Article 4 of the Directive 2019/790? That talks about «text and data mining» which as defined, doesn't include text generation

Of course it includes text. It's explicitly about "TEXT and data mining", which is mentioned multiple times through the directive. It defines the right to ingest publicly available data for training models, subject to opt outs. This is then referenced specifically and directly in the AI act, on the subject of training large language models.

Obviously existing and copyrightable work cannot be reproduced verbatim from such a system. Reproducing != training

> Source on that? As far as I can tell, weights aren't copyrightable.

Who said anything about weights?

Copyright license? I thought the LLM community didn't believe in those...

Posted Jun 14, 2026 22:11 UTC (Sun) by parodper (guest, #179096) [Link]

> Of course it includes text. It's explicitly about "TEXT and data mining",

MINING, not generation.

> Who said anything about weights?

My bad, misread what you wrote and thought you were talking about AI corporations.


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds