Complaint: Non-free training data makes SmolLM3-3B a non-free ML application

#52
by JLouisBiz - opened

SmolLM3-3B's training configuration includes facebook/natural_reasoning as part of its pretraining data -3. That dataset is licensed CC-BY-NC-4.0 -2-8 — a non-free license, because its NonCommercial restriction discriminates against fields of endeavor and violates the Open Source Definition -11. and four free software freedom as defined here: https://www.gnu.org/philosophy/free-sw.html

The model weights are Apache-2.0. But under the FSF's criteria for free machine learning applications: "we cannot say a ML application is free unless all its training data and the related scripts for processing it respect all users, following the four freedoms" . A free weight license does not make the application free if the training data is non-free.

This is not a matter of probability about output. It is a matter of the training data itself.

Could you please in your next version do some better research and avoid non-free datasets so that output can be safe for users?

Sign up or log in to comment