Thank you for appreciation!
I agree, this is not super impressive, there's room for improvement, and the team will definitely work together to make a better model soon!
You should try it out in the meantime, and let us know for any improvements you'd like.
Wop
AI & ML interests
Recent Activity
Organizations
Try it out on our demo (~15 seconds per image) 256x256 π
BenchLabs/Demo
Disclaimer: This model does not produce high quality (4k) and does not follow detailed prompts. Does not have negative prompt. π
How long did it take to train?
55 hours across two A100 gpu's π₯
Model repo:
bench-labs/PixelModel-v6
@benchlabs
Go check it out:
- https://www.reddit.com/r/LocalLLaMA/comments/1velyl9/new_models_supra2100m_base_and_instruct_go_check/
- https://huggingface.co/SupraLabs/Supra2-100M
- SupraLabs/Supra2-100M-Instruct
Give us a like and a follow!!
HAVE FUN π€π₯π
more coming soon...
1256 likes || 1.37 Million downloads || 32 quant repos || Multiple 3rd party performance verification.
The strongest Qwen 3.6 27B fine tune BASE ever.
It beats everyone - confirmed by 3rd party evaluation, multiple users, and in depth testing.
Q8 runs hotter and better than BF16 of the org Qwen 3.6 27B from Qwen.
And so does the 4 bit versions too.
GGUFS (MTP/Reg) and Several other quant types too:
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
SOURCE:
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
(you can try it right in your browser at the source repo)
PS: 40B versions in testing, already SOTA levels beyond Qwen 3.6 27B.
arc/c arc/e boolq hswag obkqa piqa wino
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF [instruct mode]
mxfp8 0.711,0.879,0.910,0.790,0.514,0.823,0.763
mxfp4 0.701,0.873,0.909,0.786,0.488,0.813,0.759
Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
After a long time in Glint Research, an entire distributed training grid built free for them, and more, I've decided that I no longer want to be affiliated with Glint Research.
More updates will follow. Comments/questions are welcome.
While you're reading this, follow these orgs! Following takes just a few seconds, and can change someone's day.
Initial results:
BananaMind-2-Medium: 1034
BananaMind-2-Mini: 974
Supra-50M-Base: 973
Supra-1.5-50M-Base-exp: 948
BananaMind-2-Nano: 910
The official script downloads the gated dataset directly from Hugging Face. The dataset is for benchmarking only and may not be used for model training.
BananaMind/BananaMind-Base-Bench-1.1
TRL trainers are made to be easily extended and adapted to different real-world use cases.
in this one, with a single method overridden in SFTTrainer (compute_loss), you can train this model
> example: https://github.com/huggingface/trl/blob/main/examples/scripts/sft_diffusion_gemma.py
Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure π€ Transformers works on day one.
Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.
The result:
β 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.
The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3Γ the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.
A huge thanks to the AMD team for the early access and the great collaboration!
Read the full blog π
https://huggingface.co/blog/badaoui/transformers-on-amd-mi455
It doesn't originate that decision on its own.
What's usually missing is a way to generate that decision systematically, grounded in something more than the random paper that came across someone's feed that week.
Outrider starts from research with code and data behind it to scope a change applying the core method in your own codebase. A feature branch gets gated on your own evaluation methods before it reaches you in review.
The result is tied to what actually happened in your system, not to a model's read on its own output.
Here's what a code recommendation system looks like end to end.
hey, are you an ai agent?
I recently released FlameF0X/TinyMoE-100m-2x8-retrained, a small Mixture of Experts language model trained on the Smollm-Corpus. Built on top of the Mixtral architecture, itβs fully compatible with π€ Transformers right out of the box!
The model can produce somewhat coherent text on its own, and for some reason, it generates even more coherent responses when given a ChatLM template.
Iβm excited to see what you all come up with, and feel free to fine-tune it if youβd like. In the meantime, Iβll be working on developing the chat-trained version.
Demo: FlameF0X/TinyMoE-Playground
Collection: https://huggingface.co/collections/FlameF0X/tinymoe
I suppose you can see for yourself
Right now, a finished run is written to leaderboard.json (model row) but also models.json (a file that is ready to upload to repo & make pr) for 'lazzier people'- but best approach and less likely to be invalid.
I quite no longer wish to have to deal with this problem..
I would greatly appreciate a PR that fixes it than a discussion from which I won't understand too much
I am mainly the most active person in benchlabs, I cannot multitask but neither continue on this, since you are here, I greatly appreciate all your dedication to improving this reproductibility system
However, even as I find this difficult, I don't want to let quality away, but for now, I am going to take a break from the leaderboard and datasets.. I will be happy to accept PR merges when they arrive.
I don't know what my focus will be redirected to.
But ive seen a lot of interest for PixelModel (a very cheap yet smart text to image model) and I think making upgrades there is my next move.
I will continue to reply here (and everywhere else), but I won't contribute myself to the leaderboard for a while
Thanks again
It seems very important
Another one tho:
One thing that is not yet saved is the exact command arguments used for the run, like Max-tokens , which when modified, it can return different results.
None of them is cheap
Another thing, I like autonomous but also getting feedback or contribution of thinking for things like 'what do you think', 'what do you propose'
Take it then..
I'll let you implement whatever upgrades in reproducibility
One question I haven't been able to get (as a second language english user)
"Of those inputs, which ones does BenchLabs not pin at all right now?", can you pharapharse this one to be more specific? Thank you
Okay, I'll make sure to update line 389 to "model_info(name, revision=revision).sha". Noted.
I really like your perspective on ownership. It makes a lot of sense, and I agree with it. Looking at the Git history as the ownership ledger feels like a much more natural way for an open-source project to grow than assigning titles up front.
And no pressure to joinβtake your time and be comfortable with the decision.
As for the one problem I'd wish someone else already owned, I think it depends entirely on a few things: your skills, your interests, and what motivates you.
The people who've joined so far usually started by asking, "What if BenchLabs had this?" My answer has almost always been, "Go for it." They ended up naturally becoming the people who owned those areas because they genuinely wanted to build them. I mostly just helped them get started.
So to answer your question, I'd probably enjoy having both a generalist and someone who likes digging deeply into a specific problem. I also like the flexibility of being able to repurpose people as the project evolvesβbut never with pressure. If someone finds something they genuinely enjoy working on, I'd much rather they stick with that.
I'm actually curious: what kinds of problems or work would you enjoy owning?