[License-review] For Approval: OpenMDW License Agreement, versions 1.1 (OpenMDW-1.1)

Moming Duan duanmoming at gmail.com
Mon Aug 31 09:54:30 UTC 2026


Hi Maffulli,

Some people believe that the data used to train the model is "archived" in compressed form in the weights, therefore the weights aren't just numbers but a full copyright-based derivative of the training data.

That is about training data, which I agree is debatable (TDM, fair use). My concern is RAG. In a production setup, the database stores my text as chunks of the original sentences plus vectors for retrieval. The chunks are verbatim copies, kept precisely so the model can retrieve and reproduce them.

Under the license definition, that database would likely count as associated data, part of the Model Materials. So claiming my own work back means claiming that the Model Materials infringe my copyright, and Paragraph 5 ends my license. Note that Paragraph 5 says "any person or entity": any compute provider who deploys the model and stores user inputs can use the same clause against a user's claim.

You want to keep the right to use a model that you claim has illegally used your copyrighted material and distributes it?

Let me make it concrete. Very few users can self-host an LLM today; most depend on cloud deployments. Suppose a database cloud service stores your private data to improve the experience of its users. Under this license, that data likely becomes part of the Model Materials. When you sue for copyright infringement, your license terminates, and your entire business, which runs on that model, goes down with it. Does that still look like an open source license to you?


(Forgive me for not feeling entirely comfortable. But NVIDIA has just agreed to buy Hugging Face for 12.9 billion dollars.)

You sue the developer of the model and, as a consequence
of initiating the lawsuit, you lose the license to keep using the model.
Just like you lose the license to use Firefox if you sue Mozilla for
patent or copyright violation.

Maybe. I am not a lawyer. But note the premise: that clause only bites if I hold a patent that could threaten Mozilla (I have none). What I do have is a lot of copyrighted content, and I think most people do too. I am not saying it fails the OSD. I am just saying that, as an “Open” model license, 1.1 looks more aggressive than MIT. As a developer, I would avoid models under it.

This is also why the ModelGo licenses leave data alone. I prefer data to keep its own license, outside the licensed unit.

Best,
Moming


-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.opensource.org/pipermail/license-review_lists.opensource.org/attachments/20260831/69c53027/attachment.htm>


More information about the License-review mailing list