[License-review] For Approval: OpenMDW License Agreement, versions 1.1 (OpenMDW-1.1)
McCoy Smith
mccoy at lexpan.law
Thu Aug 27 00:32:53 UTC 2026
On 8/26/2026 4:59 PM, Pamela Chestek wrote:
>
> We are in a world where probably every single line of code by every
> single FOSS developer is in every single LLM. Many FOSS developers are
> unhappy about it and believe it is unlawful, a question that will not
> have any clear answer for a number of years.
>
I do want to probe on this one a bit (with the caveat that the
litigation over AI training -- now over 100 just in the USA alone:
https://chatgptiseatingtheworld.com/ -- are still in the phase where the
public doesn't know a whole lot about what the parties are disclosing or
finding out about how the various AI systems are actually trained and
what they actually retain).
I think it is true to say that every single line of code by every single
FOSS developer (at least that they did distribute through well-known
repos) was used to train at least some of the extant source code
generation tools. I'm not sure it it true to say that "every single line
of code ... is in every single LLM." There have been various analyses
published of how some of the LLM systems work (alas, not for code, but
for "literary" works) and the level of abstraction of the written
material retained by the actual model seems to be fairly high (and what
it takes to get meaningful matching is quite a lot of prompting,
oftentimes specifically designed to get the match). It seems like it is
quite a bit higher than the authors believe it to be, but not
necessarily as high as the AI companies may have lead the public to
believe. And in the case of code, the authors' claim may be even harder
to make, as there is quite a bit of "idea" vs "expression." In the end,
it appears that the authors may eventually be left with only a claim
against the copies that were made of the training materials as a
precursor to training, not what is actually in the LLM models. Which, if
true, doesn't seem to me to have anything to do with the license that
may be attached to the model.
As Pam noted, we don't know where these cases come down, and TBH I
suspect in the end some of the models are going to present a harder
analysis with regard to (C) than others, but I also don't think we
should assume that the presentation by either side of the disputes as to
what is happening under the hood will actually bear out. We might get
some sense of all this (at least in the USA) when the 3rd Circuit
decides the Ross Intelligence case, but whether that is transferable to
the situation with the code generation systems remains to be seen.
McCoy [in my personal capacity]
More information about the License-review
mailing list