AI & ML interests

The AI community building the future.

Recent Activity

nielsrย  updated a bucket about 11 hours ago
huggingface/paperswithcode-backups
tarekziadeย  updated a bucket about 14 hours ago
huggingface/transformers-ci-telemetry
alvarobarttย  updated a dataset about 14 hours ago
huggingface/DEH-image-scan-data
View all activity

Articles

nielsrย 
updated a bucket about 11 hours ago
tarekziadeย 
updated a bucket about 14 hours ago
evalstateย 
posted an update 2 days ago
view post
Post
1678
Hugging Face MCP Server v0.4.9
~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Attach images from hf_fs tool.
sergiopaniegoย 
posted an update 6 days ago
view post
Post
587
Something I really like when I study a subject is understanding its history, how it reached the point where it is today

I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words

This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale

https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
sergiopaniegoย 
posted an update 11 days ago
view post
Post
235
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"

you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced

and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine

the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL

blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
  • 3 replies
ยท
sergiopaniegoย 
posted an update 12 days ago
view post
Post
2575
LFM2.5-2.6B just dropped!

and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.

basically, a full agent training pipeline but compressed into 2.6B

base model โ†’ SFT โ†’ specialized teachers per domain (SFT + RLVR) โ†’ on-policy distillation back into one student โ†’ agentic RL

the two most interesting stages

โ†’ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution

โ†’ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box

this makes a 2.6B that beats much larger models on instruction following and tool use

SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)

โ†’ model: LiquidAI/LFM2.5-2.6B
โ†’ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
โ†’ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
  • 3 replies
ยท
sergiopaniegoย 
posted an update 17 days ago
view post
Post
2620
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now

you look at the drawing and you know. but there is no number, so nothing can train against it, no?

I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL

read the details!๐Ÿค“

https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
  • 2 replies
ยท
sergiopaniegoย 
posted an update 18 days ago
sergiopaniegoย 
posted an update 20 days ago
view post
Post
2898
quick reminder! ๐Ÿšจ

tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series

๐Ÿง  what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
๐Ÿ—“๏ธ when: Tuesday, July 28 - ๐Ÿ•” 5:00 PM CEST / 8:30 PM IST
๐Ÿ“ where: Live on @huggingface 's X, YouTube, and LinkedIn

live: https://www.youtube.com/watch?v=ztdTed5egrM

class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
  • 1 reply
ยท
sergiopaniegoย 
posted an update 23 days ago
view post
Post
222
you can now train your own coding agents with trl + openenv, starting with opencode

we just added end-to-end support for training agent harnesses:

> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs

you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced

we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.

> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv

and we're working actively on both sides so expect more ๐Ÿค“
  • 1 reply
ยท
sergiopaniegoย 
posted an update 24 days ago
sergiopaniegoย 
posted an update 25 days ago
view post
Post
248
join us next Tuesday, July 28, for Class 3 of the Training Agents live series!

we'll dive into reinforcement learning for agent training, covering the intuition behind GRPO, how it works, and how to implement it in TRL with practical, e2e examples

see you there ๐Ÿค 

live: https://www.youtube.com/live/ztdTed5egrM

> in case you missed class 1:
https://x.com/SergioPaniego/status/2069382207618379813
> and in case you missed class 2: https://x.com/SergioPaniego/status/2075180665184686187
evalstateย 
posted an update about 1 month ago
view post
Post
1366
Hugging Face MCP Server v0.3.29
~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Included "papers" in the new hf_fs tool. Includes listing of trending/daily.

This is a new tool under observation - disable the "Paper Semantic Search" tool for best results.

hf://papers/
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ daily/
โ”‚   โ”œโ”€โ”€ latest
โ”‚   โ””โ”€โ”€ YYYY/
โ”‚       โ””โ”€โ”€ MM/
โ”‚           โ””โ”€โ”€ DD/
โ”œโ”€โ”€ trending/
โ””โ”€โ”€ ARXIV_ID/
    โ”œโ”€โ”€ metadata.json
    โ”œโ”€โ”€ paper.md
    โ”œโ”€โ”€ models/
    โ”œโ”€โ”€ datasets/
    โ””โ”€โ”€ spaces/

sergiopaniegoย 
posted an update about 1 month ago
view post
Post
7756
Frontier models use distillation as a step of their post-training pipelines.

In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.

I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026

It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL!
  • 3 replies
ยท
albertvillanovaย 
posted an update about 1 month ago
view post
Post
3737
๐ŸŽ‰ KTO is now part of the stable TRL API

As of Promote KTO to stable API, KTOTrainer and KTOConfig have graduated from trl.experimental to the stable trl API. https://github.com/huggingface/trl/pull/6175

This one closes out a long road. Over the past 6+ months, the "Align KTO with DPO" effort landed ~90 PRs methodically bringing KTO up to the standard we hold for stable trainers, one carefully-scoped change at a time:
- Feature parity with DPO: full VLM support (incl. multi-image), sync_ref_model, PEFT + Liger, ZeRO-3 + PEFT dtype fix, pad_to_multiple_of, activation offloading, IterableDataset and dict eval_dataset, remove_unused_columns, and reference-logprob precomputation at init.
- Consistency with DPO: aligned method order and signatures, tokenization, _prepare_dataset, PEFT handling, ref-model preparation for distributed training, and config layout โ€” plus a new DataCollatorForKTO and output format. Metrics moved into _compute_loss and simplified to direct averages via the shared _metrics attribute.
- Removing legacy baggage: dropped encoder-decoder support, BOS/EOS handling, null_ref_context, generate_during_eval, model_init, preprocess_logits_for_metrics, model/ref adapter names, and several dead config knobs.
- Coverage: a full test suite mirroring DPO, text collator tests, VLM tests, and slow tests.
- The promotion itself: the experimental โ†’ stable move (#6175) and shim cleanup (#6287), handled so downstream users get a clean deprecation path.

Honestly, this has been one of the more complex tasks I've taken on since joining the team, not because any single change was hard, but because it demanded sustained consistency across a ~2,000-line trainer, with every branch, comment, and edge case kept in lockstep with DPO.

Huge thanks to everyone who reviewed along the way (especially @qgallouedec ), the incremental review cadence is exactly what kept this maintainable.

KTO now sits on equal footing with our other flagship trainers. ๐Ÿš€
  • 2 replies
ยท