Open MFA with a TOTP app that should solve it.
Eren Ekşi
AI & ML interests
Recent Activity
Organizations
Boris-2 is 30B out of 200B tokens in, and it is severely behind its competitors in training.
We have determined the bug to be a configuration error. Boris-2 has been in training for ~1 week, and was projected to finish on November 3rd, 2026.
We are unfortunately going to restart training, with proper configuration.
The new projected finish date is ~15-18th of November.
We apologize for the delay.
We've been building an open AI ecosystem focused on practical models that can run closer to the edge — not only in large datacenters.
Today, I'm introducing T3 Gemstone, a growing family of compact AI models built around edge inference, cybersecurity and computer vision.
💎 T3 Gemstone currently includes:
* NanoSOC Gemstone 2B — GGUF
* NanoSOC Gemstone 4B — GGUF
* Gemstone Person/Object Detector Nano
* Edge-focused AI experiments and deployments
The goal is simple:
Build smaller, practical and open AI systems that can actually run on constrained hardware.
This is part of a much larger open-source effort we're building from Türkiye across LLMs, cybersecurity, computer vision, retrieval, speech and edge AI.
There is much more coming.
🤗 Explore my models, datasets and demos:
@GoktugD
🛡️ NanoSOC:
Werea-co/Werea-NanoSOC-8B
Feedback, benchmarks, collaborations and contributions are very welcome.
If you're interested inopen-source AI, Turkish AI research, edge AI or cybersecurity models, follow the journey.
We're just getting started. 🇹🇷
#AI #OpenSource #HuggingFace #LLM #EdgeAI #Cybersecurity #ComputerVision #TurkishAI #MachineLearning
Absolutely not you can check my own models Ivme-Conversate-S-v1-Base and S-v2-Instruct if you want I hate to do unrelated mentioning but
They literally do what you say
Low training tokens, different architecture (for v1), high quality data and it flopped so bad I didn't put v2 to leaderboards.
Second you said SmolLM2 uses curated data
Well
I guess what do we use
Take a wild guess it's FineWeb-Edu!
Why?
BECAUSE IT'S HUGGINGFACE
THAT'S THE POINT OF HUGGINGFACE
Also you said things about water usage I can tell you BananaMind's model is definitely more efficient in the amount of compute spent on what hardware.
The new flagship from Axiomic Labs takes 3rd on the open SLM leaderboard trailing only the SmolLMs, check it out and follow us:
AxiomicLabs/GPT-X2.5-135M
I recently released FlameF0X/TinyMoE-100m-2x8-retrained, a small Mixture of Experts language model trained on the Smollm-Corpus. Built on top of the Mixtral architecture, it’s fully compatible with 🤗 Transformers right out of the box!
The model can produce somewhat coherent text on its own, and for some reason, it generates even more coherent responses when given a ChatLM template.
I’m excited to see what you all come up with, and feel free to fine-tune it if you’d like. In the meantime, I’ll be working on developing the chat-trained version.
Demo: FlameF0X/TinyMoE-Playground
Collection: https://huggingface.co/collections/FlameF0X/tinymoe
Our top most active users:
🥇#1 @Quazim0t0
🥈#2 @IvmeLabs
🥉#3 @GODELEV
If you are an HF creator and want to try entering the podium, feel free to visit the embed page to add your graph!
https://hfviewer.com/model-card-embed
We would like to clarify that SupraLabs has no affiliation, partnership, or connection whatsoever with "SupraLarps" or its members.
Please avoid interacting with their organization, repositories, or Spaces under the assumption that they are associated with us.
We are currently aware of the situation and have already contacted the appropriate channels to address it.
Thank you to everyone who continues to support SupraLabs. ❤️
Why is Supralabs getting a lot of drama every day recently? There are trolls everywhere its sad looking at it.
I mean, like, we are talking about a 50M model here. That thing, already in full precision, can barely do basic math. On 2 bits, even a single question being right is impressive.
Look at how other hardware generations work: a new iPhone costs roughly what the previous one did. Same with Samsung Galaxy phones, same with AMD Ryzen CPUs. Price increases are modest and the specs are straightforward to compare. Nvidia has broken from this pattern. Each new generation costs meaningfully more than the last, while the benchmarks get harder to interpret, not easier.
Look at what's happening with the numbers. BF16 and FP16 performance figures, the ones most relevant for real workloads, are increasingly buried or absent. Sparsity is baked into headline numbers without always being clearly flagged. The B200 is being benchmarked as an NVLink pair rather than a single card in some comparisons. Each of these is defensible in isolation, but together they make it genuinely hard to do apples-to-apples comparisons across generations.
The inference vs training tension is worth flagging too. Lower precision formats are good for inference but introduce real tradeoffs for training. As Nvidia pushes further into inference optimization, NVFP4 being the latest example, it raises a fair question about where training workloads go. NVFP4 training exists, but documentation and real-world results are thin enough that most teams aren't seriously evaluating it yet.
The "ROCm is immature" narrative is also getting a bit stale. I've personally run workloads on the MI300X and honestly it felt smoother than I expected, even compared to CUDA in some cases. And the price to performance ratio isn't even close. AMD isn't a consolation prize anymore.
At some point, labs and enterprises doing the actual buying will have enough alternatives like AMD MI300X, custom silicon, and inference-optimized ASICs to push back. Nvidia is standing on thin ice, and a few more steps like this might be enough to crack it.
What scares me most are the legal consequences of this, and the possibility that all models will start converging on the same tone because everyone is just fine-tuning on everyone else.
Opus 4.6 released
But idk this thing didnt really looked like a magic.
Though Sonnet 5 could also drop i dunno.