|
|
Log in / Subscribe / Register

Incorrect dataset licenses in MOT

Incorrect dataset licenses in MOT

Posted May 30, 2026 21:41 UTC (Sat) by aphedges (subscriber, #171718)
Parent article: MOT: a tool to fight openwashing in AI

> I asked a two-part question about how the submissions to MOT were audited

Reading this article, I was excited to see there was finally a way to find fully open models. Unfortunately, it seems that most of the class I models don't have training data under open licenses. Both the Pythia and OLMo entries say that their datasets are licensed under Apache 2.0, but they both contain proprietary sources according to their linked documentation.

I've reported this as an issue on GitHub that was acknowledged by Le Hors (https://github.com/lfai/model_openness_tool/issues/228), so hopefully these problems are fixed and more validation is done in the future.


to post comments


Copyright © 2026, Eklektix, Inc.
Comments and public postings are copyrighted by their creators.
Linux is a registered trademark of Linus Torvalds