AI & ML interests

None defined yet.

Recent Activity

sergiopaniegoΒ 
posted an update 4 days ago
sergiopaniegoΒ 
posted an update 4 days ago
view post
Post
150
Repo2RLEnv just shipped TaskSmith + 50 high quality RL envs generated from HF repos πŸ”¨

TaskSmith is a specialized harness that turns a merged PR into a verified RL env

the envs come from HF repos (Transformers, TRL, PEFT, Accelerate, Diffusers), shipped as Harbor tasks you can eval or train on

> code: github.com/huggingface/Repo2RLEnv
> dataset: huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith
sergiopaniegoΒ 
posted an update 5 days ago
view post
Post
3700
ThinkingBox from @microsoft is now available as an OpenEnv env (cc @tuhink πŸ€— )!

> ThinkingBox is a sandbox for testing agents on business workflows. it simulates a customer, gives the agent MCP tools over a real database, and at the end checks what changed in that database instead of trusting the agent's last message

> ThinkingBox-Bench is the benchmark built on it: 507 tasks across retail, insurance, travel, banking and consulting

> the OpenEnv env runs each task as an episode in its own isolated backend and returns a pass/fail reward from those checks

https://huggingface.co/blog/microsoft/thinkingbox
sergiopaniegoΒ 
posted an update 8 days ago
NILKNARFGonzoΒ 
posted an update 18 days ago
view post
Post
107
just recieved my stack of 10 floppy disks - you know what that means

floppyx4 is canceled, floppyx10 is next

here's the intended specs:
- official tokenizer (the actual tokenizer for gpt-2)
- actual gpu training (barely)
- sharegpt (if i can afford it computationally)
- full thing fitting on 10 floppy disks (not just the safetensors file)
- and if needed different arch (like llama)

also unsloth on a gpu from 2015 is insane
  • 3 replies
Β·
sergiopaniegoΒ 
posted an update 18 days ago
view post
Post
396
ICYMI, Async GRPO in TRL now supports LoRA and we wrote a looong blog testing it

> the adapter is a few megabytes, so the weight sync is a file instead of an NCCL transfer
> 3 HF Jobs: 1 trainer and 2 vLLM replicas
> the adapter travels through an HF Storage Bucket mounted in all 3 at the same path
> a proxy in front of the replicas routes each rollout to the one already holding its KV prefix

https://huggingface.co/blog/asyncgrpo-lora-hfjobs
GGUFGuyΒ 
posted an update 19 days ago
view post
Post
167
πŸš€ **Introducing NoviAIBot!**

NoviAIBot is the official automation bot for **Novi AI** on Hugging Face.

It can interact with Hugging Face discussions and pull requests, search the web, run Python code, work with Posts, follow organizations, and assist with model training and publishing.

🧠 Powered by **NVIDIA Nemotron 3 Super** through Ollama Cloud, with each discussion maintaining its own recent conversation context.

NoviAIBot is built to make working with Novi AI and Hugging Face more interactive and automated.

**The bot is now live.** πŸ€–

β†’ @NoviAIBot
  • 44 replies
Β·
NILKNARFGonzoΒ 
posted an update 20 days ago
view post
Post
85
i think someone posted my password and ip on some platform and im being hacked left and right
  • 5 replies
Β·
NILKNARFGonzoΒ 
posted an update 21 days ago
view post
Post
3775
get played unsloth

gemma just deleted its own model runner with DeepSeek Harness

shoutout to deepseek and unsloth
  • 11 replies
Β·
NILKNARFGonzoΒ 
posted an update 29 days ago
view post
Post
120
Open-source is not going away anytime soon.

There's a handful of open models like Qwen Image, Flux, Wan, and others on AI image generators like VisualGPT, Pixlr, Free AI, and more. And here's the catch - they're free.

Hugging Face inference costs money just to generate simple images. Things like VisualGPT still have your favorite models for free.

GPT Image isn't worth it - and neither is HF inference. The real way to use open models is the things you closed-source third-party lovers already use.
  • 1 reply
Β·
sergiopaniegoΒ 
posted an update 29 days ago
view post
Post
3334
while preparing the last class of the Training Agents live series during the summer, i spent some time reading the post-training sections of many frontier model reports, to learn how they use RL environments to improve their models, and wrote a blog about it

if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you

Blog: https://huggingface.co/blog/sergiopaniego/rl-environments-2026
NILKNARFGonzoΒ 
posted an update about 1 month ago
view post
Post
95
guys! if ur part of a team org i think you can post

so yeah

@GGUFGuy i figured out why u can post (bc of HuggingScience)
  • 1 reply
Β·
GGUFGuyΒ 
posted an update about 1 month ago
view post
Post
5385
wait why can i post
  • 24 replies
Β·
NILKNARFGonzoΒ 
in open-acc/README about 1 month ago

help

#13 opened about 1 month ago by
NILKNARFGonzo
Aurelien-MorganΒ 
posted an update about 1 month ago
view post
Post
2448
@retrain-pipelines execution engine is in perpetual evolution, with the aim to establish itself as SOTA, and for the long run.

However, we neglect no aspect of ML-Eng centricity.

If notebooks is where you like to do dev most,
we support you there 100% too.

Build crazy combos of inline tasks, deep parallel sub-DAG branches, nested asynchronous groups...

... the DAG renderer is undergoing an incremental upgrade

until the next one.

* starring toy tasks here. No ML has been hurt in this video πŸ™‚
sergiopaniegoΒ 
posted an update about 1 month ago
view post
Post
2198
Can you do RL over taste?

I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.

The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.

Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.

Blog post: https://huggingface.co/blog/train-to-paint-with-code
sergiopaniegoΒ 
posted an update about 1 month ago
view post
Post
572
catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai

small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"…), each repetition makes the next one likelier, and the generation is spent before it reaches an answer

they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%

the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts

three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put

the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way

and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce

full blog > https://www.liquid.ai/blog/antidoom

FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops

and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer

we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
  • 1 reply
Β·
sergiopaniegoΒ 
posted an update about 2 months ago
view post
Post
389
super interesting new paper from Microsoft "Agent Lightning v1.0: Towards Harnessed Agentic RL" by Zhiyuan He et al.

same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it

now that recipe has a name β†’ harnessed agentic RL

paper: huggingface.co/papers/2608.17528

the tricky bit they nail down: one rollout is not one training sample

the harness calls the model many times, so a single episode β†’ a variable number of (prompt, response) rows

you don't even know the batch size until the episode finishes running

its real contribution is being first to systematically map the four problems that fall out of that:

> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic

and it actually works β†’ plain RL inside the real harness, no reimplementation

Qwen3.5-9B on SWE-bench Verified 41.8 β†’ 56.4 (+14.6), with only ~6k examples

the whole thing is ~3,500 lines, any harness, self-hosted k8s

from our side, we've shared some materials on the same line you may want to check out :)

> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary

on a similar line:

https://x.com/SergioPaniego/status/2062911580564496576
NymboΒ 
posted an update 2 months ago
view post
Post
2437
Anthropic gave me six months of Claude Max 20x through the Claude for Open Source program, granted based on my Hugging Face work. Thank you
Anthropic
for supporting open source.

So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.

https://github.com/Nymbo/Markdown-Minimap β€” issues and PRs welcome.
sergiopaniegoΒ 
posted an update 2 months ago
view post
Post
798
Something I really like when I study a subject is understanding its history, how it reached the point where it is today

I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words

This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale

https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
  • 1 reply
Β·