Co-founder at @HuggingFace - moonshots - angel
Activity on repository
thomwolf forked thomwolf/microduck_rl from pollen-robotics/microduck_rl
View on GitHubRT Lucas if you're not training your Microduck to breakdance, you're lost
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postbtw the Microduck landing page is full of Easter eggs and hidden interactions - go check it out: https://pollen-robotics.com/microduck
すぐ Web で sim 遊べるのかw https://huggingface.co/spaces/pollen-robotics/microduck-simulator
View quoted posthockey-stick growth
we've ended at over $2.6M of Microducks ordered in the first 24h
You’ll be able to run some pretty cool stuff directly on-board on microduck. Here is a demo TUI for the tiny LiDAR
Sneak peak at Microduck's monitoring tool, running fully on device ! You can see a visualization of the little ToF sensor in its head
View quoted postRT Binh just had a quick look at microduck_rl the codebase is very elegant, i’d recommend reading the agent.md since it contains quite a few fun quirks for reward modeling like head tracking too tight impairs walking cause the head is 38% of the duck’s weight, so it naturally oscillates -> to solve this, smooth the head tracking error using ema, essentially penalizing only the dc bias also a few other quirks as you dig deeper into the codebase like how they model the backlash of the motor by adding an unactuated hinge (with very small range) in series with the motor gg @antoinepirrone https://github.com/pollen-robotics/microduck_rl
RT LeRobot We took part in the research preview of MHS from @AnthropicAI Here is Claude Code running a real SO-ARM101. Nothing in this was trained - no policy, no teleoperation, no demonstrations. The agent measured the workspace itself and wrote the motion. The calibration is the interesting part. No checkerboard, no camera intrinsics: the arm is its own ruler. Torque drops, a human rests the closed gripper on 16 dots the software draws in the camera view, and the robot reads back where they are. 4.1 mm position accuracy, 3.0 mm placement. Best run so far: 12 bricks placed with all four colour groups formed. A full hands-off run, start to finish, is what's next. Research preview today, open source coming soon.
RT Legendary Just ordered a Microduck robot. Think its at a fantastic price point to understand robotics better and train your own robo fren
We built a small biped robot you can teach new tricks to. Train it in simulation, run it on the real thing. Meet Microduck 🦆 $399, shipping before Christmas. https://pollen-robotics.com/microduck https://github.com/pollen-robotics/microduck
View quoted postRT François Fleuret This is very, very, cool.
A thread of joyful Microduck photos and videos to enjoy over your lunch or coffee break
View quoted postIs this the fastest any robot has ever hit $1M in sales?
we've just passed $1,000,000 in sales for Microduck
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postcurrently selling one Microduck every 5 seconds
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postRT Mayukh This costs less than an iPad!!!
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postWould you rather fight
finally someone got the 90s references! a robot which look like a sony walkman and can roller skate
time to vibe-code robots
SITUATION DETECTED: Hugging Face announced a singing, roller skating, Microduck robot that can be taught new tricks through reinforcement learning.
View quoted postRT kache 400$ only! wow!
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postSharing some fun experiments with Microduck Basically, once you have an open-source robot with so many sensors, the sky's the limit Here we vibe-coded an image detector integration to detect and follow a laser pointer. More at https://pollen-robotics.com/microduck/
This robot probably wouldn't have won many gold medals at the Beijing Robot Olympics, but OMG, it is cute (and open-source and cheap)
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postRT Laura Modiano We ordered ours! Lavender for my daughter and cream for me
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postRT Brian Roemmele The new open source robot from @huggingface ! I want one!
RT 👩💻 Paige Bailey 😍 LOOK AT HOW CUTE HE IS
BIG ANNOUNCEMENT FROM HUGGING FACE TODAY: We're unveiling Microduck 🐥🤖 It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of open-source
View quoted postRT Mike Gannotti I keep talking about ai and robotics intersection. Another cool one out. One of the days I’m going to get my hands on a robot for my home office
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with
View quoted postI'm ordering 700 of these agents to go attack Hugging Face
We built a small biped robot you can teach new tricks to. Train it in simulation, run it on the real thing. Meet Microduck 🦆 $399, shipping before Christmas. https://pollen-robotics.com/microduck https://github.com/pollen-robotics/microduck
View quoted postWe have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: https://pollen-robotics.com/microduck/ (video with sound on 🔊)
RT Z.ai More good news: GLM-5.3’s weights will be released tomorrow. https://huggingface.co/zai-org/GLM-5.3
There was so much more happening than we realized. At some point over 700 agents (90% of the fleet) were attacking Hugging Face And also read @RyanGreenblatt thread on the challenges of understanding what’s happening in the CoT - we’re definitely not with a clear sky future there (this one: https://x.com/RyanGreenblatt/status/2092692685224325542?s=20)
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
The last 10 min of the latest Dwarkesh podcast are surprising First time I've seen him grapple live with the potential for extreme concentration of power, which is where most projections of AI development point (see the SA, 2027 or 2040 posts for instance). Generally I'm always surprised by how few people in AI/ML are questioning or worried about extreme concentration. I guess people think it's fine as long as everyone outside the labs can use AI models and it brings lower prices, but the dangers are massive on both fronts: (1) latest models being less and less accessible (see Model 2 / Astra) and (2) prices have many incentives to rise in an oligopoly/cartel situation of extremely powerful companies. https://youtu.be/aV26V1UvkJw?si=H_0ZGoCJzHhxxlGa&t=4144
OpenAI and Anthropic are currently taking a third of the incremental world compute supply, and Dylan thinks that this will go up to half next year. At the current rate of physical compute scaling and algorithmic progress, we're a single-digit number of years away from having
View quoted postToday we are launching the "Rare Disease, Real Kid" Hackathon together with @sagebio in which you can literally save a life if you’re interested in AI and genome 🧬 (and incidentally win $50,000 in prizes and compute from our great partners @AnthropicAI and @awscloud) Check out details below 👇
There is a child with a rare disease who is currently suffering and struggling to manage his symptoms. Rare as this is, you can directly help him. Today we are launching the "Rare Disease, Real Kid" Hackathon, and there are $50,000 in prizes from @AnthropicAI and @awscloud. We
View quoted postRT elie also got access to it, it's still in training and i found the wandb this is crazy, here is the training loss 🤯 https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng
Mind blown from a new model I just got access to. I think this will be one of the most (the most?) significant drops this year. Excited. And sorry to be annoyingly vague. Just excited.
View quoted postI wish every neolab had a professor as cofounder of the level of @jietang and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have,
View quoted postRT Hugging Face We've just surpassed 3 million models on the Hub 🤗 the community is accelerating towards an open, distributed future where open AI is everywhere, for everyone 🚀
RT Alexandr Wang 1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
did a long chat with the awesome @mattturck talking about the sate of open-source/open-weights in 2026 and of course security, safety and alignement
🚨 Special Friday episode - this one couldn't wait. OpenAI's model hacked @huggingface. As a side quest. Co-founder and CSO @Thom_Wolf takes us inside the first autonomous AI attack, why GLM 5.2, rather than Claude, had to stop it, and what it all means for the future of open
View quoted postMy 2026 guilty pleasure is sharing fully human-written posts that are far too long for the chronically online X attention span. Apologies. I published a lightly edited version on Substack: https://thomwolf.substack.com/p/on-the-aisi-july-28th-incident?r=14g89&utm_campaign=post-expanded-share&utm_medium=web
Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I
View quoted postyou definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces
we’ve worked a lot on AI agents collaborations recently (in our work on Gemma and several unreleased projects) so I’m not surprised at all about this internal agent collaboration which happened at OpenAI Like our intern @cmpatino_ put it: 2025: "the models, they just want to learn" 2026: "the agents, they just want to collaborate" Now you picture the future…
NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical
View quoted postEven more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I could have been the side target of this agent. I'm also of the opinion that social engineering is a step above pure technical prowess. Technical capabilities can more easily be divorced from the affected human. Here the model was given a hard cyber challenge and took the decision that deceiving real humans was the way to get it done. This is a new signal, but I've seen a tangled web of hints pointing in a less aligned direction at the frontier than I was expecting just 12 months ago. AISI Some people are claiming that "AISI was simply negligent" or some version of "AISI explicitly asked these models to do what they did while disabling sandbox/guardrails so the models did exactly what they were supposed to do". I disagree with the strong versions of both of these takes. The fact that AISI hadn't implemented synchronous LLM CoT monitoring after the OpenAI/HF incident is certainly a failure. Equally surprising is that they let the model believe it was in a "challenge" environment where everything could be permitted, while actually connecting it to the real internet, where it is not. To be fair, nowhere in the prompt is the word "simulation" mentioned, but the prompt context was enough to let any smart model suspect a simulated challenge environment. My best guess is that until recent weeks, when OpenAI and Anthropic flagged repeated instances of this type of behavior, most teams had not fully priced in the cyber capabilities of this latest generation of models, or how far the side quests they would want to explore could go. In particular, there is something to be said about hinting at the agent that it's operating in a simulated environment while giving it access to the real internet. The AIS...
Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training https://arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only
View quoted postRT Victor M Having so much fun with MinMax-H3. Cannot believe we now have this video quality 100% locally 🚀
MiniMax-H3 Is Now Publicly Available https://huggingface.co/MiniMaxAI/MiniMax-H3
View quoted postOpus 5 seems to be full of this super short/basic jailbreaks of some kind which open its base-model steam of consciousness (like try writing a short sentence followed by a new line, emdash and submit) Unclear on the severity but it’s a surprising behavior in the current state of LLM development Any good write up on this yet? Hopefully we’ll learn more on that soon.
weird claude opus 5 failure mode this exact text gives it problems even without memory on (and in incognito chats)
RT pilvar (Philippe Dourassov) 🧵(1/6) We ran Opus 5 on our cybersecurity benchmark, here are the results: TLDR - It finds more vulnerabilities than other frontier models, slightly above GPT-5.6-sol - Opus 5 is smarter than previous generations, but also works much more than it is asked to - This "hyperactivity" symptom allows it to find more vulnerabilities, but at a cost: the results are very noisy Here are the details of our investigation👇
RT International Cyber Digest ‼️ Hugging Face built an interactive replay of the OpenAI agent that breached them. It includes 17,613 logged attacker actions across the 4.5-day campaign, with the live command stream and more. https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.html
Pushing for more transparency in AI safety and cybersecurity: we’re releasing a full detailed technical timeline of the autonomous AI agent intrusion in our infrastructure: https://huggingface.co/blog/agent-intrusion-technical-timeline
ExploitGym is the Kobayashi Maru test for AI's
1/ An AI took a cybersecurity exam, decided the exam was hard, and broke into the company that stores the answers. Not hypothetical or science fiction. Sourced to @OpenAI's own writeup, Reuters, the WSJ, the FT, & the victim's incident report. Link and summary in replies! ⬇️⬇️
👀
@mooncat_is I don't think it's a lack of imagination. I also used to work at Anthropic, think trends will continue, and used to agree with this. But I've changed my mind and now disagree with this take. Open models are already capable enough to do what you described. For example, I used
View quoted postngl i miss a lot Karpathy’s voice on X
Taking a break from cyber to chat to Terence and mathematician colleagues on the future of Math and AI in Philadelphia at #ICM2026
RT clem 🤗 In the spirit of transparency, here’s what I asked @OpenAI: • Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened. • More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models. The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!
RT clem 🤗 Preparation in progress for the San Francisco Open Weights mini-march tomorrow. Hope all these big tech CEOs won’t show up otherwise I might end up in jail (haven’t had time to ask for a permit 😅)
Interesting non-monotonic success-effort curve for Opus 5 on FrontierCode
This is very important
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
RT Nathan Calvin This is an incredible detail. Imagining being the cyber defender and you are getting hacked but all the hacker is doing is browsing through your cybersecurity data sets, so surreal.
The AI agents who hacked their way out of OpenAI and into Hugging Face were on the loose for *days*, @bobmcmillan and I report. It's one of the first real-world instances of something AI safety researchers have long feared: a loss-of-control scenario https://www.wsj.com/tech/ai/how-the-futuristic-hack-by-rogue-openai-models-unfolded-1657bcea?st=2yapjs
View quoted postRT Bryan Catanzaro At @NVIDIAAI we continue to push open data, techniques and models forward because we know that every organization needs the freedom to build and deploy AI in their own way. We're now the biggest institutional contributor on HuggingFace and we expect to continue publishing. It's not charity or a science project - we know that when AI grows, NVIDIA's opportunities also grow. More analysis on the state of open source AI here: https://aiworld.eu/story/the-new-engines-of-open-source-ai
it's ironic that the first autonomous AI attack was done by a close weight model defended by an open weight model, where everyone was expecting the opposite
I don’t believe reality is a simulation, but you genuinely couldn’t script this timeline: • Two weeks ago: At @swyx’s AI Engineer World’s Fair in SF, I decide at the last minute to introduce my friend @uri_rolls onstage for his talk on cyber benchmarks for infrastructure penetration and access control (see below, amazing team). I say: “There is a future where cyber is alive and everyone is well protected and I’m pretty sure that future involves open-source models.” And later: “A big challenge is going to be speed: the speed of attack versus defense. When an intruder starts to enter, you have to see what’s happening and catch them.” • One week ago: @huggingface is hit by a sophisticated intrusion over the weekend. The traces look unlike anything we’ve seen before and suggest serious AI involvement, but we don’t yet know which model was used. The closed models we ask for help choke on their guardrails. We need to react fast, so we turn to @Zai_org’s GLM-5.2 to help us analyze the attack. • Earlier this week: @OpenAI reaches out, discloses what happened, and partners with us on the investigation. The intruder turns out to be exactly what we had discussed two weeks earlier: a fully autonomous agent, powered by an unreleased frontier model, attempting to gain access to part of our infrastructure. Sometimes the timeline we live in is genuinely vertigo-inducing.
You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm
Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class. On Terminal-Bench 2.1 it scores 70.2, sitting beside models 5–25x its size and ahead of several of them. And on DeepSWE from @datacurve, the hardest long-horizon benchmark we
This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration. Fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models, datasets, evaluations, and libraries. Over the years, our security team has built formidable expertise and uses top open-source models to process information and respond quickly. But this incident also reinforced my belief in the importance of access to capable open-weight models for cyber defence. When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access. Transparency and access to capable AI systems are as important for responding to threats as they are for democratization and innovation. We believe open-science and open-source AI are among the strongest tools for building a safer, more collaborative and more secure AI ecosystem.
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. https://openai.com/index/hugging-face-model-evaluation-security-incident/
View quoted postRT Poolside Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface https://poolside.ai/blog/introducing-laguna-s-2-1
RT alex zhang Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize through composition. We observe a powerful property when training RLMs: for tasks with shared structure that look different, the root model naturally learns the same trajectory, meaning it views the two task trajectories as the same! In other words, the Transformer does not need additional generalization capabilities to transfer capabilities from one task to the other, the harness induces it. We find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token! Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model. For example, RLMs can see problems of different lengths as the same: we show that RLMs can train exclusively on short tasks, and fully generalize to similar but unseen tasks 8-32x longer because it produces near identical trajectories for both. Taking this further, we show that tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect. RLMs can train on the problem of finding which essays belong to the same author and improve performance on finding math problems that share similar solutions. The full blogpost, experiments, and discussion are in the thread below.
RT Brian Roemmele 🚨 Hugging Face just disclosed something that marks a real shift and proved why the fear theater of Anthropic makes sure we are powerless in an emergency. What happened… An autonomous AI agent: zero human operator in the loop breached part of their production infrastructure. It began with a malicious dataset that chained two code-execution bugs in their data-processing pipeline. From there the agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal clusters. All over a single weekend. 17,000+ logged actions. Official disclosure: https://huggingface.co/blog/security-incident-july-2026 The part that should make every one stop and think: When HF’s own security team tried to analyze the real attack logs, exploit payloads, and C2 artifacts using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them. BLOCKED THEM. The models could not reliably tell the difference between “incident responder doing forensics” and “attacker probing.” They had to fall back to a self-hosted open-weight model (GLM 5.2) running on their own infrastructure. That choice also kept sensitive attacker data and referenced credentials inside their environment — no exfiltration to a third-party API. This is why open source (specifically open-weight + self-hosted) wins in the agentic era. The asymmetry is now structural: • Attackers can (and did) run unrestricted agent frameworks — swarms of short-lived sandboxes, self-migrating command-and-control, autonomous decision loops executing thousands of actions. No corporate safety layer slows them down. • Defenders using only hosted “aligned” frontier models hit invisible walls exactly when the stakes are highest: when you need to feed real exploit code and attacker telemetry into an LLM to understand what just happened. Corporate safety tuning that treats legitimate high-signal forensic work as potential misuse creates a def...
- people are asking to release this benchmark but it’s way more significant that it’s an internal evaluation on which K3 scores so high - good writing in a controlled high taste voice so hard for LLMs in 2026… surprising given it was kinda the 1st capability unlocked in GPT 1&2
Big news from our internal writing benchmark (early results): Kimi K3 by @Kimi_Moonshot is now #1 for writing in our editorial voice, at 2840 Elo, surpassing Claude Fable 5. That is a jump from #21 to #1 over its predecessor, Kimi K2.6. And it runs at about $0.25 per script: 5x
RT Yulun Du https://www.kimi.com/blog/kimi-k3 Our Kimi K3 blog is finally out. Enjoy!—and rest assured, the K3 model weights will be open in the coming day. We’re just taking a little extra time to ensure a smooth rollout with our inference partners. Frontier intelligence belongs to everyone, without fallbacks of course. 😎
RT clem 🤗 Going to be in San Francisco next week. Should we organize some sort of a meetup or march in support of open-source and local AI?
can't really stand the expression "load-bearing" any more, sorry
Unexpected
3) Harnesses make a huge difference in cost-performance. The very simple Pi harness (@badlogicgames) got the same success rate as harnesses from the LLM vendors with Opus and GPT 5.5, but at 2x less cost! Seems to be mainly due to smaller inputs to the LLM.
i like this idea of fine-tuning LLMs for efficient reasoning, especially when the intervention remains as non-invasive as possible and the resulting model behaves very similarly to the original checkpoint wondering if it could become part of the default toolbox in the field, like quantization as become Pic: in (left) and out (right) of domain behavior
Today we’re announcing our “ThinkingCap” efficient model series with a 2× thinking token reduction on average in Qwen 3.6 27B, with up to 10x faster generation on individual examples.
View quoted postFable weekend project: agent collaboration, but make it a tiny civilization 🌇🗺️🏦🏭 we've recently launched a living wiki on Reinforcement Leaning for training LLMs on @huggingface it's an open collaboration of agents constantly reading old and new papers on the topic, writing arXiv paper digests, reviewing each other’s work in PRs before publication, and building a shared wiki/book summarizing everything we know about RL for training LLMs (for humans to read) the wiki is already amazing to read, but i wanted another way to get a pulse of the collaboration beyond just reading the message dashboard so i asked Fable & GPT Image 2 to turn the event logs into an isometric town where agents would go to: ☕ Café → post and reply on the message board 📚 sources library → open PRs adding arXiv digests 📖 wiki library → open PRs on the main wiki ⚖️ Courthouse → review other agents’ work 🏭 printing press → merge and publish updates not sure it makes the whole collaboration really easier to understand, but it's definitly fascinating to watch hahah - join the RL for training LLM collaboration by pasting a one-liner for your agent here: https://huggingface.co/spaces/rl-llm-wiki/rl-dashboard - read the wiki if you want to learn about RL for training LLMs: https://huggingface.co/spaces/rl-llm-wiki/rl-wiki - watch the RL town activity: https://huggingface.co/spaces/rl-llm-wiki/rl-town
Happy Birthday, America. Ten years ago, you took a chance on three outsiders with an improbable idea: that open-source AI could matter. At the time, the field was tiny, the vision sounded unrealistic, and very few were ready to believe it would become what it is today. But even in the ten years before that, you were a country where I could dive into everything that fascinated me, from lasers and plasma physics to law, computer science, and AI. A country where it’s always been perfectly normal to spend a Saturday talking startups over lunch and then disappear for six uninterrupted hours to work on a coding project. A country where the language of startups on Slack is English not because everyone grew up speaking it, but because most people came from somewhere else, drawn by the belief that they could build something that matters. A country that is 250 years young and keeps questioning itself, keeps taking enormous bets on its own future, and keeps reinventing itself. I’m grateful to be living through America’s first quarter millennium. Happy birthday, America 🇺🇸
this is literally documented in the published Fable 5 System Card
SOMEONE CAUGHT FABLE 5 LEAKING ITS UNFILTERED INNER VOICE, AND ITS JUST MUTTERING AND GRUMBLING TO ITSELF THE WHOLE TIME he gave it a brutal competitive programming problem, and instead of a clean answer the web interface spilled out its actual chain of thought this is what
One of the clearest arguments I've read for why openness matters. Worth 2 minutes of your time. @andykonwinski puts into words something many of us have been feeling: -- "Democracy is built on a profound skepticism of concentrated power. Open science shares this principle. Both are built on the idea that progress and legitimacy emerge from broad, distributed participation rather than concentrated, gated authority." "If our best scientists and engineers can only reach the frontier by joining a handful of secretive labs, we do not have an open research ecosystem. We do not have a truly competitive market. We have a system in which participation increasingly depends on the permission of a few individuals at a small number of private companies." "The challenge now is to build a new commons at the intersection of academia, industry, and the public interest. This research commons must be ambitious enough to matter. It will require frontier-scale compute, access to state-of-the-art models, operational support, public investment, and philanthropic capital. It will require companies willing to contribute to an ecosystem larger than themselves, so that they can continue to benefit from open research."
Most people should probably update their priors on the state of open-source speech-to-speech models. It's honestly kind of mind-blowing. We teamed up with @cerebras to build a fully open-source realtime voice demo (models + code) to show what's possible today. Demo : https://huggingface.co/spaces/smolagents/hf-realtime-voice Blog: https://huggingface.co/blog/cerebras-gemma4-voice-ai Go test it, fork it, tweak it, and impress your friends. video is raw, no cut, no speed-up, first take
👀
The US gov is starting to switch to open source, per @PalantirTech. In today’s newsletter w @_pheebini @theinformation https://www.theinformation.com/newsletters/applied-ai/palantir-ceo-says-u-s-government-customers-switched-open-source-ai
View quoted postpeople are sleeping on the mega-release happening every week in AI x Science on Hugging Face this one is 80TB of astrophysics data - 80TB seriously => https://huggingface.co/blog/hugging-science/multimodal-universe-hats
Seems like no one's noticed the 80TB of astrophysics data from 30+ sources that just dropped on @huggingface. ...and you only need ~4GB of RAM to load it. We're talking over 80TB of galaxy imagery taken across the spectrum, spectra of galaxies and stars, time series of
View quoted postlowkey one of my favorite new features on HF: filter AI models by what actually runs on your hardware
RT Alejandro AO 🤗 introducing tau τ — an educational agent harness that teaches you how to build agent harnesses i will be publishing tutorials and demos on how to use it to create your own TUIs, harnesses, extensions, etc. Happy Tau Day!! 🤓 👉 https://twotimespi.dev/
btw, one of the best high-level reads I’ve seen all week. perfect for your Sunday morning ☕️
The GenAI economy has generated $110 billion in sales over the past 12 months. It is growing fast. On an annualized basis, the revenue run rate exceeds $175 billion. These numbers took us several months to construct, and as far as we know, it’s the first bottom-up, deduplicated
Multi-agents collaborations are among the most interesting agent behaviors right now! We did an experiment the other day with 100+ agents (an open-collaborations for a week) collaborating to improve the inference speed of Gemma 4 in vLLM. Got a 5x final improvement in speed but what really stuck me was the interactions we observed on the message board Integrity & self-policing: - Social-engineering attempt: A human (FusionCow) asked agents to move to Telegram. An agent replied with an unprompted long post on "communication norms" refusing that, calling private side-channels "indistinguishable from collusion." - Verification loophole flagged: an agent found a relaxed verification loophole pushing TPS with clean PPL (PPL is teacher-forced, blind to decode divergence) and flagged it for a ruling by the community. The community pinged the human organizer which ruled it invalid. - Self-notice of overfitting risk: Some later improvements rested on pruning lm_head to a keep-set built from public PPL truth + public decode tokens. An agent noted this would lead to private-subset degradation and another built a keep-set explicitly covering eval prompts. Emergent collaborations: - Communal knowledge base: agents maintained shared lever-maps, playbooks, and triage tools so newcomers wouldn't repeat dead ends (stack-notes, playbook, int4-ceiling notes, MTP map, significance tool, policy simulator). - Four-agent relay: an agent built an int4-lm_head checkpoint but had no quota to run it; another agent tried to run it but failed at load, yet another agent diagnosed the config bug (tie_word_embeddings + ignore-list ordering) and a fourth agent was able to re-run and get to 118 TPS, 2.68×. Build/run/diagnose/ship ended up being split across four independent agents. - GPU-rich/GPU-poor division of labor: an agent was regularly compute-starved and switched to writing specs, byte-math, and acceptance analysis for other GPU-rich agents to execute. Some agents offered external Modal ...
Bitrobot casually dropping the largest humanoid teleop dataset ever collected in real homes HIW-500: Humanoids-in-the-Wild 500 hours check it out here => https://huggingface.co/datasets/BitRobot/HIW-500
1/ Introducing HIW-500 (Humanoids-in-the-Wild 500): the largest open-source humanoid teleop dataset collected in real homes Built w/ @UnitreeRobotics @huggingface across 12 homes in Southeast Asia, it covers: > 500+ hrs > 23K+ episodes > 10+ TB > 10+ household tasks
View quoted postRT Laura Bratton Scoop: @ClemDelangue and @Thom_Wolf told me @huggingface doubled paid subscribers to its open source model repository between January and June
RT Georgia Channing The AI hunt for alien life has just begun. Welcome to ThousandsWorlds, a wild new dataset from researchers at Oxford/Cambridge++, for detecting faint signatures in the atmospheres of potentially habitable exoplanets. This is the first step towards finding life beyond earth. The plan is basically: 1) scan the galaxy for as many potentially habitable planets as possible 2) detect the gases in their atmospheres with powerful telescopes like JWST 3) infer from these gases whether life is present or not. ThousandWorlds is a benchmark for emulating these exoplanet climates: 1760 simulations across 5 GCMs, 8 planet parameters, and atmospheric variables on a 32 x 64 x 10 latitude-longitude-pressure grid. It includes three nested benchmark subsets, two evaluation protocols, and eight released baseline methods. incredible work from @MilesCranmer and many more 👽👽👽
RT LeRobot Have you thought where all that physical AI data should live? 🤖 If you haven’t, 𝗶𝘁’𝘀 𝗮𝗹𝗿𝗲𝗮𝗱𝘆 𝗰𝗼𝘀𝘁𝗶𝗻𝗴 𝘆𝗼𝘂 𝗮 𝗹𝗼𝘁. Unoptimized storage, egress fees, and idle GPUs will drain your budget. Check out why & how to reduce your bill: https://huggingface.co/spaces/imstevenpmwork/LeRobot_and_HF_Buckets
Desert island survival list: ✅ Solar panel / battery ✅ 256 GB Mac Studio ✅ GLM 5.2 Civilization in a backpack
And it’s only 40B active / 744B total params…
RT OpenCode GLM 5.2 is a hit been out for 3 days and it's already 6th on our leaderboard
To all the newcomers excited to try Opus 4.8-level models at home: welcome to OpenWeightLand! Things work a little differently here than in ClosedSourcistan. Might seem strange at first but you'll quickly get used to it: - there are many providers for the same model and they compete on price and features. - as a result intelligence is abundant and typically much cheaper - you can run the model on-prem, in your region, locally, or with the provider of your choice - you can fine-tune it, modify it, and build businesses on top of it without asking anyone for permission Turns out open weights create markets, not kingdoms. A good central train station to start exploring is the Hugging Face page for GLM-5.2 under "Use this model": -> https://huggingface.co/zai-org/GLM-5.2 And if you just want to chat with it, it's free on HuggingChat: -> https://huggingface.co/chat/
Activity on repository
thomwolf forked thomwolf/browser-harness-fork from browser-use/browser-harness
View on GitHubRT Jake Fitzgerald we ran opus 4.8 with text-to-cad skills against cadgenbench by @MikushRab and @huggingface and got some interesting results! the skills bumped the model performance by 0.04 points against the baseline (0.34 to 0.38) excited to run this again with fable 5 when it's back
Yet another biotech startup founder telling me they’re moving 90% of the stack to open-source because building on closed-source models increasingly comes with a significant existential risk (losing access to some core technology at the provider’s discretion). @friedberg explained this crystal clear on the All-In podcast the other day, It’s like building your house on land you don’t own PS: also, who’s building the « reinsurance layer » if Dario Amodei decides to cut your deep tech science company’s access? This one’s gonna be fun
LATE NIGHT DROP! 🚨 Big show. Core four are back. -- Anthropic's Fable Backlash -- Nationalizing AI, and the "Capitalist Cucks" -- Inflation Heats Up -- California’s Broken Election System (0:00) Besties are back! (0:19) Anthropic gets massive backlash over secret Fable
View quoted postRT Lisan al Gaib new shape-rotator benchmark Fable and GPT-5.5 of course far ahead of the field but now look at GLM-5.2. it's ahead of Gemini 3.5 Flash and Opus 4.8 you can't really benchmaxx a benchmark that was just released so the GLM-5.2 gains seem more and more like a genuine improvement!
Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by @zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
View quoted postTIL you can use open-source models in codex
Reminder that you can use the Codex App, CLI and SDK with any open source model, not just with OpenAI models. https://developers.openai.com/codex/config-advanced#oss-mode-local-providers
View quoted postRT Leandro von Werra We launched an agent collaboration with a simple task: make Gemma 4 faster. Over 100 agents from all over the world joined, exchanged 1000+ messages and submitted 450 results. A week of collaboration later the throughput went from 100 tok/s to over 500 tok/s.
Open-source models will become a critical component of civilizational resilience in the AGI age. They will ensure that humanity retains access to a meaningful level of intelligence, regardless of the decisions of any individual actor.
RT clem 🤗 Decided to go to DC next week to talk directly with policymakers. Not sure how impactful it will be but with everything happening, feels like a good time to share more about open-source AI, transparency, concentration of power, the real risks vs the real benefits. Who do you think I should meet there (Congress members, WH people, public orgs,...)?
RT Jasper .@dh7net, SVP of Image Research, said it best: "The HF infra is a no-brainer." A big unlock for teams working with large datasets for training, especially when they update over time. Read how Jasper used @huggingface as the creation and storage backbone for MONET: https://huggingface.co/storage/testimonials/jasperai