ZeroGPU Explorers
community
AI & ML interests
None defined yet.
Recent Activity
sergiopaniegoย
posted an update about 14 hours ago
Post
44
super interesting new paper from Microsoft "Agent Lightning v1.0: Towards Harnessed Agentic RL" by Zhiyuan He et al.
same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it
now that recipe has a name โ harnessed agentic RL
paper: huggingface.co/papers/2608.17528
the tricky bit they nail down: one rollout is not one training sample
the harness calls the model many times, so a single episode โ a variable number of (prompt, response) rows
you don't even know the batch size until the episode finishes running
its real contribution is being first to systematically map the four problems that fall out of that:
> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic
and it actually works โ plain RL inside the real harness, no reimplementation
Qwen3.5-9B on SWE-bench Verified 41.8 โ 56.4 (+14.6), with only ~6k examples
the whole thing is ~3,500 lines, any harness, self-hosted k8s
from our side, we've shared some materials on the same line you may want to check out :)
> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary
on a similar line:
https://x.com/SergioPaniego/status/2062911580564496576
same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it
now that recipe has a name โ harnessed agentic RL
paper: huggingface.co/papers/2608.17528
the tricky bit they nail down: one rollout is not one training sample
the harness calls the model many times, so a single episode โ a variable number of (prompt, response) rows
you don't even know the batch size until the episode finishes running
its real contribution is being first to systematically map the four problems that fall out of that:
> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic
and it actually works โ plain RL inside the real harness, no reimplementation
Qwen3.5-9B on SWE-bench Verified 41.8 โ 56.4 (+14.6), with only ~6k examples
the whole thing is ~3,500 lines, any harness, self-hosted k8s
from our side, we've shared some materials on the same line you may want to check out :)
> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary
on a similar line:
https://x.com/SergioPaniego/status/2062911580564496576
Post
103
CaraArchive highlights a broader reality of putting data online: once something is publicly accessible, it becomes extremely difficult to guarantee that it will remain under your control.
If there is information or artwork that you absolutely do not want copied, archived, scraped, downloaded, or used by others, the safest option is still not to publish it publicly in the first place. That may sound obvious, but the internet was fundamentally designed to move and reproduce information, and there are countless ways to retrieve publicly accessible images:from ordinary browser tools and web scraping to automated or agentic systems.
That does not mean artists should simply accept every possible use of their work. Artists deserve meaningful control, attribution, compensation, and reasonable ways to express how their work may be used. But treating the technology and peopple using it itself as the enemy is unlikely to solve the underlying problem.
There probably isn't a technical solution that can make a publicly visible image simultaneously viewable by everyone and impossible to copy. The realistic goal should therefore be to create better norms, incentives, licensing systems, and tools around how that content is used.
Like, we can imagine a future where every artist gets his/her own credentials and some kind of fingerprint done just like blockchain works. But that requires substantial cooperation among organizations, companies and individuals.
Technology and art are not inherently opposing sides though.
If there is information or artwork that you absolutely do not want copied, archived, scraped, downloaded, or used by others, the safest option is still not to publish it publicly in the first place. That may sound obvious, but the internet was fundamentally designed to move and reproduce information, and there are countless ways to retrieve publicly accessible images:from ordinary browser tools and web scraping to automated or agentic systems.
That does not mean artists should simply accept every possible use of their work. Artists deserve meaningful control, attribution, compensation, and reasonable ways to express how their work may be used. But treating the technology and peopple using it itself as the enemy is unlikely to solve the underlying problem.
There probably isn't a technical solution that can make a publicly visible image simultaneously viewable by everyone and impossible to copy. The realistic goal should therefore be to create better norms, incentives, licensing systems, and tools around how that content is used.
Like, we can imagine a future where every artist gets his/her own credentials and some kind of fingerprint done just like blockchain works. But that requires substantial cooperation among organizations, companies and individuals.
Technology and art are not inherently opposing sides though.
All ZeroGPU Spaces stuck on "Starting" with clean logs โ CPU and paid GPU unaffected
๐ 2
3
#183 opened 6 days ago
by
luna0805
Post
161
If you lack ideas for a cool model, here's one.
Train a model from scratch on wikipedia with one twist: the tokenizer changes the actual token ids used on every sample fed. If somehow still learns English, you have made an astonishing discovery.
You would have answered the question: Can a model learn human languages from structure alone?
Train a model from scratch on wikipedia with one twist: the tokenizer changes the actual token ids used on every sample fed. If somehow still learns English, you have made an astonishing discovery.
You would have answered the question: Can a model learn human languages from structure alone?
Post
853
- GLM 5.2
- Flux 3
- New Qwen model
- New small model leaderboards
- Lots of people finetuning smol models.
- Some even under 12 year olds clauders are here (was not on my bingo card this year)
- ChatGPT's Sol became a lot faster this week
- LFM2.5 2.6b
- Kimi K3 (though only a few will run it)
- New Ling 3.0 Tiny
- New video model that is making south park videos?
- Deepseek v4 flash being more honest than bigger models
- The new model from meta
Everything Everywhere All At Once
- Flux 3
- New Qwen model
- New small model leaderboards
- Lots of people finetuning smol models.
- Some even under 12 year olds clauders are here (was not on my bingo card this year)
- ChatGPT's Sol became a lot faster this week
- LFM2.5 2.6b
- Kimi K3 (though only a few will run it)
- New Ling 3.0 Tiny
- New video model that is making south park videos?
- Deepseek v4 flash being more honest than bigger models
- The new model from meta
Everything Everywhere All At Once
sergiopaniegoย
posted an update 10 days ago
Post
652
Something I really like when I study a subject is understanding its history, how it reached the point where it is today
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
Post
2898
If you want small models to be great again you should give a follow to people like @Banaxi-Tech or @Datdanboi25
These guys are rocking it with small models lately.
(They are not paying me to say that)
These guys are rocking it with small models lately.
(They are not paying me to say that)
hiyougaย
authored a
paper 14 days ago
hiyougaย
authored a
paper 15 days ago
hiyougaย
submitted a
paper to Daily Papers 15 days ago
sergiopaniegoย
posted an update 15 days ago
Post
255
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
sergiopaniegoย
posted an update 16 days ago
Post
2593
LFM2.5-2.6B just dropped!
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model โ SFT โ specialized teachers per domain (SFT + RLVR) โ on-policy distillation back into one student โ agentic RL
the two most interesting stages
โ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
โ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
โ model: LiquidAI/LFM2.5-2.6B
โ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
โ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model โ SFT โ specialized teachers per domain (SFT + RLVR) โ on-policy distillation back into one student โ agentic RL
the two most interesting stages
โ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
โ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
โ model: LiquidAI/LFM2.5-2.6B
โ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
โ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
Post
1353
Giving free early access to the gguf for some of you today! Tell me what you think.
CEAMFA/palmer-007-preview
CEAMFA/palmer-007-preview
MyNiuuuย
submitted a
paper to Daily Papers 18 days ago
sergiopaniegoย
posted an update 21 days ago
Post
2631
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now
you look at the drawing and you know. but there is no number, so nothing can train against it, no?
I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL
read the details!๐ค
https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
you look at the drawing and you know. but there is no number, so nothing can train against it, no?
I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL
read the details!๐ค
https://huggingface.co/blog/sergiopaniego/pelican-env-openenv