AI & ML interests

None defined yet.

Recent Activity

blanchon 
posted an update 3 days ago
view post
Post
104
I read a lot of papers these days, and I was frustrated with every PDF reader. So I built the one I wanted instead.

Available on Mac, Windows, Linux, Chrome and the web.
Xivly is free and open source 👇
https://github.com/julien-blanchon/xivly
sergiopaniego 
posted an update 4 days ago
sergiopaniego 
posted an update 5 days ago
view post
Post
154
Repo2RLEnv just shipped TaskSmith + 50 high quality RL envs generated from HF repos 🔨

TaskSmith is a specialized harness that turns a merged PR into a verified RL env

the envs come from HF repos (Transformers, TRL, PEFT, Accelerate, Diffusers), shipped as Harbor tasks you can eval or train on

> code: github.com/huggingface/Repo2RLEnv
> dataset: huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith
sergiopaniego 
posted an update 6 days ago
view post
Post
3706
ThinkingBox from @microsoft is now available as an OpenEnv env (cc @tuhink 🤗 )!

> ThinkingBox is a sandbox for testing agents on business workflows. it simulates a customer, gives the agent MCP tools over a real database, and at the end checks what changed in that database instead of trusting the agent's last message

> ThinkingBox-Bench is the benchmark built on it: 507 tasks across retail, insurance, travel, banking and consulting

> the OpenEnv env runs each task as an episode in its own isolated backend and returns a pass/fail reward from those checks

https://huggingface.co/blog/microsoft/thinkingbox
sergiopaniego 
posted an update 9 days ago
NILKNARFGonzo 
posted an update 19 days ago
view post
Post
107
just recieved my stack of 10 floppy disks - you know what that means

floppyx4 is canceled, floppyx10 is next

here's the intended specs:
- official tokenizer (the actual tokenizer for gpt-2)
- actual gpu training (barely)
- sharegpt (if i can afford it computationally)
- full thing fitting on 10 floppy disks (not just the safetensors file)
- and if needed different arch (like llama)

also unsloth on a gpu from 2015 is insane
  • 3 replies
·
sergiopaniego 
posted an update 19 days ago
view post
Post
398
ICYMI, Async GRPO in TRL now supports LoRA and we wrote a looong blog testing it

> the adapter is a few megabytes, so the weight sync is a file instead of an NCCL transfer
> 3 HF Jobs: 1 trainer and 2 vLLM replicas
> the adapter travels through an HF Storage Bucket mounted in all 3 at the same path
> a proxy in front of the replicas routes each rollout to the one already holding its KV prefix

https://huggingface.co/blog/asyncgrpo-lora-hfjobs
GGUFGuy 
posted an update 20 days ago
view post
Post
171
🚀 **Introducing NoviAIBot!**

NoviAIBot is the official automation bot for **Novi AI** on Hugging Face.

It can interact with Hugging Face discussions and pull requests, search the web, run Python code, work with Posts, follow organizations, and assist with model training and publishing.

🧠 Powered by **NVIDIA Nemotron 3 Super** through Ollama Cloud, with each discussion maintaining its own recent conversation context.

NoviAIBot is built to make working with Novi AI and Hugging Face more interactive and automated.

**The bot is now live.** 🤖

→ @NoviAIBot
  • 44 replies
·
NILKNARFGonzo 
posted an update 21 days ago
view post
Post
85
i think someone posted my password and ip on some platform and im being hacked left and right
  • 5 replies
·
NILKNARFGonzo 
posted an update 22 days ago
view post
Post
3775
get played unsloth

gemma just deleted its own model runner with DeepSeek Harness

shoutout to deepseek and unsloth
  • 11 replies
·
NILKNARFGonzo 
posted an update 30 days ago
view post
Post
120
Open-source is not going away anytime soon.

There's a handful of open models like Qwen Image, Flux, Wan, and others on AI image generators like VisualGPT, Pixlr, Free AI, and more. And here's the catch - they're free.

Hugging Face inference costs money just to generate simple images. Things like VisualGPT still have your favorite models for free.

GPT Image isn't worth it - and neither is HF inference. The real way to use open models is the things you closed-source third-party lovers already use.
  • 1 reply
·
sergiopaniego 
posted an update about 1 month ago
view post
Post
3336
while preparing the last class of the Training Agents live series during the summer, i spent some time reading the post-training sections of many frontier model reports, to learn how they use RL environments to improve their models, and wrote a blog about it

if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you

Blog: https://huggingface.co/blog/sergiopaniego/rl-environments-2026
NILKNARFGonzo 
posted an update about 1 month ago
view post
Post
95
guys! if ur part of a team org i think you can post

so yeah

@GGUFGuy i figured out why u can post (bc of HuggingScience)
  • 1 reply
·
GGUFGuy 
posted an update about 1 month ago
view post
Post
5385
wait why can i post
  • 24 replies
·
NILKNARFGonzo 
in open-acc/README about 1 month ago

help

#13 opened about 1 month ago by
NILKNARFGonzo
sergiopaniego 
posted an update about 1 month ago
view post
Post
2200
Can you do RL over taste?

I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.

The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.

Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.

Blog post: https://huggingface.co/blog/train-to-paint-with-code