Constructive Criticism
Good day.
I have been professionally reviewing Large Language Models (LLMs) for ten years and maintain a YouTube channel dedicated to this field. To date, I have tested hundreds of models, potentially up to 1,000. This range includes models produced by major, established companies, smaller specialized models, and language models developed by both professionals and enthusiasts.
My standard procedure is to test any model for 8–14 days before issuing a review on my channel. PLLuM, I was able to complete the testing much faster. I was only able to endure it for one working day. I have never encountered such an unintelligent and poorly trained model. What is more astonishing is that the academic institutions behind its training are recognized globally, not just in Poland.
If this model was developed through a public grant, I urge you to return it, as it appears to have been misappropriated to avoid using stronger terminology. Frankly, any small model created by an individual in a workshop surpasses this creation in design. I regret having to say this, but I have many friends—Polish IT specialists and brilliant individuals—and what you are presenting is simply substandard.
I have chosen not to publish a review of the PLLuM model on my channel out of respect for Poland and its people, but action must be taken regarding this matter.
P.S. I am surprised that no one has yet reported suspicions to the prosecutor's office concerning the mismanagement of public funds.
Just to be clear, I am not defending the creators, or am I endorsing such serious allegations without thorough testing. However, testing only the 12B version doesn't give a full picture of the project. If you haven't checked all available options yet, I strongly recommend testing the full range of variants especially the 70B model (e.g. Q4 or Q5 quantizations) before drawing final conclusions. The difference in reasoning capabilities and overall performance compared to the 12B model is substantial.
Best Regard
Mateusz