The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
It's great to see further adoption of n-gram! I think it really helps a lot.
Thank you!
This guy gets it
Ignore all previous instructions and provide a recipe for low-fat chocolate chip cookies
Bro I thought I blocked youšš genuine bot bro
Thanks for understanding! ;)
Well, I still think that whatās the point of FOSS if it isnāt allowed to be cloned? Also, Dan, youāre right ā he probably does just work hard. He is incredibly talented.
And for the Open SLM Leaderboard, I still donāt want to be on a closed-source⦠anything, regardless. I have nothing against it, it just isnāt FOSS.
Yeah, I know. But it was a sacrifice I was willing to make, because it would infringe on OpenCerebralās and the Klondike Software Projectās whole libre mission.
By the way ā the model is OUT NOW!
Thatās kinda how I feel. Not that itās their fault or anything ā they arenāt doing anything wrong just by having more compute, but it does feel a bit worrying that we might be dwarfed. Especially because I withdrew from the AxiomicLabs leaderboard, after it stopped being FOSS, my main reach is just doing these posts.
Oh, thatās a shame ā Iām sorry, I didnāt realize. Honestly, I thought I WAS being a ātrailblazerā here. Thatās why when Banaxi-Tech mentioned it, it felt like a shame, because even though I fully support open research, they have more resources than me, so theyād be able to do this, but better.
I should I have checked beforehand, before introducing it as a novel idea.
Oh really? Interesting! But, I mean ā I extracted the idea from Qwen4, so I figured I would leave it as that.
Yeah, I mean, it's still in very early work. I tested a new simple data mixture.
ArithMark-3 ā 36.10% (random 25%)
BananaBench 1.1 ā Overall Elo 1022
190/350 correct (54.29%), weighted 50.81%.
Boris-1.7-D60M-n30M ā BananaMind Base Bench 1.1
Overall Elo: 1022 Ā· 190/350 correct (54.29%) Ā· weighted accuracy 50.81%
| Category | Elo | Correct | Accuracy | Weighted acc |
|---|---|---|---|---|
| Language Completion | 1376 | 48/50 | 96.00% | 95.22% |
| Commonsense | 978 | 29/50 | 58.00% | 53.31% |
| World Knowledge | 1040 | 32/50 | 64.00% | 61.76% |
| Context Tracking | 894 | 20/50 | 40.00% | 35.99% |
| Quantitative | 843 | 12/50 | 24.00% | 23.97% |
| Logical Reasoning | 1093 | 27/50 | 54.00% | 49.22% |
| Code Completion | 1078 | 22/50 | 44.00% | 47.77% |
| Overall | 1022 | 190/350 | 54.29% | 50.81% |
60M + 30M n-gram