Bringing data science back to AI - https://t.co/Zrmp6LRd9c About Me: https://t.co/P6WyeKkyTa
RT Sarah Catanzaro Eval developers are the next analytics engineers; people who will set standards that ultimately govern how models are trained and deployed. I expect to see many more companies hire for this role.
WE NOW HAVE A HEAD OF EVALS the work he's doing is honestly gamechanging. cannot wait to show it to you
View quoted postReally great use of tokens. Highly recommend. This is an AGI level task b/c have to click forms and such. Codex computer use flying through it
RT jacky we're hiring cracked evals/benchmarking ppl research-oriented role super tough hairy unsolved problems those that thrive in unknowns preferred pls dm me with your actual work (not resume), will fast track you
100% of the replies to this are AI bots Super disappointing
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed directly through the browser. Read the post for discussion of the tradeoffs. 2) They created a new kind of notebook which works with WebMCP that prioritizes meeting people where they are: you bring your own coding agent and files are just markdown. The author uses it to curate runbooks or high quality examples of how to run foundation model evals on their infrastructure. Notebooks are good for this since they require tinkering with state of long running jobs interactively while taking notes inline. And its open source ✨ Blog: https://learn.chatgpt.com/blog/automating-repetitive-work-at-openai-with-codex P.S. this post is authored by Jeremy Lewi who isn't on X but here is his website https://lewi.us/
RT dex You can’t claim “the models are good enough that I don’t have to read the code”. Because if you’re not reading the code, then you can’t possibly know how much slop is getting in. We go live to @0xblacklight and @vaibcode for the scoop
RT Peter Yang I don’t promote other people’s courses often, but I want to recommend Hamel and Shreya’s top rated AI evals course. Here’s why I think it’s worth taking: 1. A 4.7-star rating across 900 reviews, with students from OpenAI and Google, is almost unheard of on Maven. 2. They’ve completely revamped the course and added a 24/7 AI evals assistant to guide you through the process. 3. They’ve personally helped me improve my evals, and their advice is consistently practical, specific, and grounded in real experience. You can watch my free podcast episode with them to learn more about their eval process before deciding: https://youtu.be/bdMHQLvtVaQ The next cohort of their course starts September 5 and you can get 25% off with this link: https://maven.com/parlance-labs/evals?utm_campaign=peteryang&utm_medium=affiliate&utm_source=maven&promoCode=PETERYANG
RT Florian Brand tired: hill climbing an eval by doing a synth env wired: hill climbing an eval by fixing the eval
RT Peter Yang There are two types of AI evals - tops down and bottoms up. From @sh_reya: “Think of top-down as: If you’re in a vacuum, just given the task description, what would you come up with? Claude does a very good job of helping you with these top-down evals. Bottom-up evals are the other half. When you look at lots and lots of sample outputs, what is your gut feedback that you want externalized into evals? “Claude is very, very bad at coming up with bottom-up evals. That’s all you.” 📌 Watch the full episode here: https://www.youtube.com/watch?v=bdMHQLvtVaQ
“The fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them
View quoted postI was wondering why Google is in its own special toilet tier in this video 🤣 Summary: it’s expensive as hell compared to everything else
😅Now I **have to** try it. Firing up the old model training rig that’s been collecting dust
@HamelHusain This is how the best adventures start. Just ask me and @tobi 😂. Many such stories! Never once have I regretted it.
View quoted postSneak peek of segments of the new evals course ✨😃
we've been iterating a lot on how DocWriter internally represents a user's writing style! the naive approach is to dump all a user's prior writing in context and pray the AI figures out how to sound like them. this doesn't really work. and the user can't just manually specify
RT Peter Yang “The fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them to audit the evals I built for my creator skills live. They then demoed a free skill that you can use in Claude Code or Codex to build reusable evals from your feedback. Some quotes from Shreya and Hamel: “Bottom-up evals come from looking at lots of sample outputs and turning that into eval criteria. AI is very bad at coming up with them. That’s all you.” “The agent’s job is not to invent new feedback. But it can help you group and distill the feedback into actionable rubric criteria.” “All your competitors can point Claude at their product and say, ‘Find all the errors.’ What matters is how much taste you can infuse beyond that.” 📌 Watch now: https://www.youtube.com/watch?v=bdMHQLvtVaQ Thanks to our sponsors: @WisprFlow: 4x faster than typing with your voice https://ref.wisprflow.ai/peteryang @linear: The AI agent platform for modern teams https://linear.app/behind-the-craft
It's incredibly hard not to get nerdsniped by Omarchy Looks amazing
RT Peter Yang In my next episode, I asked @sh_reya and @HamelHusain (taught evals to 4,500+ engineers and PMs) to roast the AI evals that I built for my creator skills live. They also showed me how anyone can use their free Error Discovery skill in Claude Code or Codex to turn real AI failures into reusable evals. 📌 Subscribe to get the full episode tomorrow: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
WTF is "deep tech" its a term I only hear in investor speak and makes no sense to me
RT Peter Yang I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate crossing 100K subs on YouTube. Excited to share a lot more practical interviews soon to answer your most burning AI questions: 1. @HamelHusain and @sh_reya (AI evals experts) on how today's AI models have changed evals completely 2. @amoljain_ (Head of Product Engineering at Replit) on which vibe-coded apps have actually become real businesses 3. @ebloch (Product Lead at OpenAI) on how ChatGPT Finance can help you save both time and money 4. @poteto and Roman (SpaceXAI) on how the Grok Bot team uses Grok @bot 📌 Subscribe to get the episodes soon: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
RT sarah guo .@gabepereyra, the research team at @harvey, and their partners are giving everyone building specialized intelligence a clear blueprint for the huge performance and efficiency gains possible from post training and in-domain data and workflow understanding own your intelligence!
Vendor posts that insult a specific competitor make me seriously consider signing up for the competitor Because they must be pretty damn important to name. Very rookie comms and marketing mistake
RT knut sat down and actually read @rosmine's research paper on this. https://deftwriting.com/research/distribution-fine-tuning first of all, i was guilty of having a "take" based on this post, pointing out how the copy in the screenshot is not a great example of what "good writing" is, despite scoring 100 on "human written" in @pangram. I recommend actually reading the work before commenting on it - because sloppy takes are as lazy as sloppy texts. (shame on me for joining the band wagon) but sleeping on it, i am grateful that @rosmine took the time to dive into this stuff. AI-slop fatigue is real, and we should support efforts to make agents produce communication that is clear and lucid, and doesn't feel overly synthetic but there are some assumptions in this, however, that is interesting to question when it comes to "what makes for good writing" as far as i understand, Deft is trained to have more diversity/variation in tokens, so less repetition of what we are recognizing as "AI-tells" (certain word and stylistic choices). And i totally agree with @rosmine that: "LLMs are not the cause of slop. Lack of effort/care is. If you spend days researching and planning a blog post, and put all the information into a detailed, well-structured outline, and ask ChatGPT to generate the post based on the outline, then the output will be interesting to read, even if the text has a lot of em-dashes." I argued the same in our eng blog announcement post yesterday: https://www.sanity.io/engineering/announcing-the-sanity-engineering-blog BUT! I still feel that this report (at least somewhat), but especially the various takes on it, conflates something sounding "human" with it being "good." Tricking @pangram doesn't make a text well written. Some reflections: - Making writing sound more "human" by means of adding more variation in word/style choices, doesn't make it better - What makes for a "good" text is highly contextual. If you are writing a recipe or instruction...
Announcing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We
RT Jonathan Whitaker I've officially left https://answer.ai I've had a good rest, with plenty of time for travel & tinkering, and now I'm thinking about what to do next. I've got a few ideas to share soon, but I'm also open to suggestions - feel free to reach out :)
RT ¯\_(ツ)_/¯ Re @HamelHusain @pangram @rosmine the iron law of tech twitter is when your employees start douche posting you are hitting some revenue wall and they need an outlet for frustration. pathetic stuff.
RT Alex Strick van Linschoten Re Good question. The main secret sauce is... you :) Basically we believe that evals are all about pairing domain experts with systems that allow them to encode their taste and judgement and put this all together in a workflow where you're improving your agent by seeing what's going wrong, using deterministic (where possible) evaluators to capture those failure points, and then using the replay etc to make sure that you've actually fixed things at the root. We're not really at the point where you can just automate evals fully without humans being involved (see @HamelHusain's recent post https://parlance-labs.com/blog/posts/auto-evals/index.html on some of the ways that can go wrong), but for sure tools (like coding agents, or like Kitaru) can help make this process as painless as possible.
New meme template thx to @BEBischof
A few months ago,@sh_reya and I released eval skills plugin. We iterated on it a bunch and recently made it better The biggest change is a new error-discovery skill. Give your coding agent a file of AI outputs or traces, and it builds a custom review app w/intelligent sampling. As you annotate the sample, the agent groups your notes into failure modes and finds related examples. We also added a start skill, which looks at your situation and routes your agent to the right workflow. It can help you find errors in a set of traces or audit an eval pipeline you already have. Writeup: https://hamel.dev/blog/posts/evals-skills/ GitHub: https://github.com/ai-evals-course/evals-skills
Hire John. I worked with him personally at Airbnb and he’s top 1%
I'm looking for my next role. I'm an AI engineer / PM, currently based in Berlin and open to relocating. Most recently Head of AI & Product at an AI fintech until the startup shut down in June. Before that: data science at Trumid and Airbnb, and Head of Data at Circ. Over the
RT John Enevoldsen I'm looking for my next role. I'm an AI engineer / PM, currently based in Berlin and open to relocating. Most recently Head of AI & Product at an AI fintech until the startup shut down in June. Before that: data science at Trumid and Airbnb, and Head of Data at Circ. Over the past couple of months, I’ve also been getting much closer to the model side: training three LLMs from scratch in PyTorch, post-training them with SFT and GRPO, then building the inference and serving stack, including KV caching, streaming and FastAPI. Everything is open source: models, weights, code, write-ups, and a playground with all nine checkpoints: https://huggingface.co/spaces/JohnEnev/modern-llm-playground I’m looking for an AI engineering or AI PM role, ideally somewhere I can stay close to both the model work and the product. Berlin, remote, or open to relocating. If you know of something, or someone worth speaking to, I’d really appreciate a pointer or repost.
Activity on hamelsmu/evals-skills
hamelsmu opened a pull request in evals-skills
View on GitHubRe: watermark - all that’s gonna happen is there will be tools and APIs to strip the watermark out In the end this will result in more friction instead of visibility
RT Shreya Shankar Interesting to see all the agreement. Check out the DocWriter project's plain-writing skill: https://github.com/docwriter-org/plain-writing-skill We recently added evals for the skill --- comparing a gpt-5.5 agent without the skill to with the skill (using an LLM judge to evaluate each criterion)
New term coined: ✨ Fablish ✨
@HamelHusain I also cannot understand Fablish. I feel crazy because everyone else loves Fable so I’m wondering what I’m doing so wrong
View quoted postI cannot use Opus 5.0, yes it can code but it can no longer explain what it did intelligibly It feels like its opaque internal reasoning dialect is now being used to talk to humans?
RT Shreya Shankar opus 4.6 feels like the last model that spoke english
RT Shreya Shankar Lots of new material in our AI evals course this fall. We put together a revised version of the 200-page course reader. Some of my favorite additions include: how to take an agentic approach to building evals, and how to think about safety, risk, and privacy as someone deploying a bespoke agent for their organization. Next cohort starts September 5th. Check out the full syllabus here: https://maven.com/parlance-labs/evals?promoCode=evals-info-url
Grok 4.6 just ranked #1 on CursorBench 3.2 Outperforming Claude Fable 5, Opus 5 and GPT-5.6 Sol on real-world coding performance And what makes this even crazier is the efficiency....the chart gives CursorBench performance against average cost per task, and Grok 4.6 is sitting
RT Isaac Flath Amp orbs are actually pretty useful. Here's an example. I started a thread in the cli working on my blog. Had it use a portal to host the site to show me the changes. I'm working on a technical post and wanted it to render jupyter notebooks nicer so it did that and I saw it work in the site portal. I moved to to the web UI because it was just kinda nice to have my website preview and coding agent both in chrome in the same app. I went back to work on my post, so I told it to make another portal for jupyter. This was nice because I don't have jupyter on this machine and I don't really like telling agents to just go install stuff on my local computer, but in an orb sandbox it's fine. It's on same machine so save of notebook reloads the preview site. And it all just worked. I could have done all this on localhost like I usually do. But it was nice because it was just a bit less friction than normal
Grok Bot is cool because it is "self-driving" (like Codex desktop). It can create new threads/tasks and communicate b/w them It's also a simple interface and decent computer use + nice iOS app. I found computer use lagging behind Codex slightly but its worth trying it, especially if you have a Cursor Ultra sub https://x.ai/bot
The model "just needs encouragement" is a UX issue. It's not something to be proud of
“overcast” describes cloud cover precisely, while “grey” might describe the light, sky, mood,etc. Yes, I care, and it matters.
There is no blog post that will win over developers re: watermarking AI
RT Nick Dobos Anthropic’s response to watermarks confirmed my worst fears. “Our watermarking method doesn’t have any practical impact on the quality or content” That is a lie. They literally say they change the wording in the first paragraph Anyone with an ounce of common sense about how psychology, writing, and legal documents works in the real world, knows this completely changes the meaning and content. This degrades and changes petulance. Look at the example they picked, it’s garbage. “Overcast” vs “grey” are completely different descriptions. That’s their best example, and it changes the sentence subtly yet dramatically. This is unbelievably dangerous This will kill people I am not exaggerating Anthropic just handed the keys to mind control the government They just gave politicians a fully legal back door that can be exploited to influence claude’s thinking to mind control millions of people without anyone ever knowing This is not okay.
We’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
View quoted post💯 “You didn't read the thing when you generated it, I won't read it when I'm reading it.”
new post: write for people https://vickiboykis.com/2026/08/12/write-for-people/
View quoted postRT Shreya Shankar It was great to present our data agent benchmark in the summer of evals series! Slides courtesy of co first author @ruiyingm1120, a second year PhD student at UC Berkeley!!
New session w/@sh_reya where she goes over a useful new eval called the Data Agent Benchmark (DAB). Data agents answer business questions like “which cohort had the highest churn?” that a data analyst would normally answer. DAB recreates the mess of a real data warehouse.
View quoted postRT vicki new post: write for people https://vickiboykis.com/2026/08/12/write-for-people/
RT Lambda Most teams lose sight of what exists outside the world of frontier APIs when it comes to their everyday work. Lambda’s @TheZachMueller sat down with @HamelHusain on when an open model is the right call, and how to serve it well once you’ve committed. It also builds on @_xjdr’s point that open weights can handle ~90% of tasks for ~90% of people. The rest stays on frontier work. http://youtu.be/Pg-IW5puuv0
Over the summer, @sh_reya and I hosted 13 sessions on AI Engineering topics like retrieval, post-training, inference, and evals. I've summarized all the sessions, organized by theme, with links to the source materials. Warning: I've tried to pull the most important ideas from each talk, so some notes are short (9.5 hours of sessions comes out to about 20 minutes of reading). Enjoy! https://hamel.dev/notes/llm/ai-product-engineering/
RT Alexis Gallagher Developing a robot by chatting with a robot https://x.com/i/broadcasts/1qKDzWzBqmkJV
RT Hugo Bowne-Anderson Fuck your skills. @HamelHusain is the guest I’ve had on Vanishing Gradients more than anyone else. We’ve put out nine episodes together over the years, including panels with @jeremyphoward, @sh_reya, @eugeneyan, @BEBischof, and @charles_irl. We’re doing it again on Friday, so I went back through the archive and made this supercut. There’s Hamel deciding he hates his own skill. There’s Hamel getting frustrated with OpenClaw because the tooling had become more work than the tool. There’s Hamel, repeatedly, telling people to look at the actual data. Hamel has been saying versions of the same thing for years: stop collecting AI shit long enough to look at the thing you built. Look at your data. Read your prompts. Find the stupid failure cases. If you can’t tell why the product gave an answer, you can’t improve it. We’ll get into all of that, and more, on Friday for Stop Shipping AI Nobody Can Verify. Register to join us live, or get the recording afterwards: https://luma.com/7lng145m
RT Joe Barrow An easy but surprisingly useful trick you can learn is napkin math for model training and inference. How much work is done during speculative decoding? How fast should a DETR run at various resolutions? If you can approximate these quickly it will help diagnose slowdowns, reason about new techniques, and maybe even invent a few of your own!
RT ben hylak introducing rd-signal-2: a frontier classification model that is 1600x cheaper than GPT 5.6 Sol. free to try in @raindrop_ai, and available via a new API for training/hosting custom classifiers with Zero Data Retention.
RT Joe Barrow Good engineering is about observability. Speculative decoding accelerates LLM inference, but you're running it blind. Every rejected draft token is wasted compute, but are you looking at the drafts? This weekend I wrote specspecs to solve that!
TIL that /visualize is built in codex skill
99% of people don't know you can tell your chief of staff thread to use `/visualize` I have a pinned travel thread that tells me my travel schedule. @PhilippSpiess has done incredible work here
It’s been a long time since I’ve been excited to work through a technical book @rasbt It’s time to bring more ML back into my life
Wisprflow feels really slow, I'm motivated to uninstall it and try something new. What are ya'll using for STT? Good iOS integration is key for all the other apps
When my wife and I wfh at the same time, we both dictate constantly to AI So we can't work in the same room anymore. How is this working in the open/shared office space?
adding LOC should be a negative reward in many situations
Funny story: an OAI employee gifted me 6 month pro sub a while back (they have an employee perk). It's actually better to pay for it b/c only paid accounts can bank all of Tibo's resets. Paid > Free 🤣 🤣
PSA: Slop infographics do NOT help make your slop post clearer
Its crazy how the consensus shifted away from Claude being the favorite to Codex It’s not just Vibes. Things favoring codex: - Better harness (Codex Desktop) - Better pricing: more things included within subscription (fast mode etc). bonus - constant resets from Tibo 😅 - Less refusals / false positive safety triggers - You can use subscription freely wherever you want (Opencode, Amp, etc) When you add all these up it makes codex significantly more attractive. There is likely an industry lag of people still using Claude but feels like it’s gonna swing back
Seems like the frontier has regressed on writing even though it has advanced on coding. The slop level has noticeably gone up
. @sh_reya is at it again. The main reason our students love our eval course is the depth. She spearheaded a new 200 pg. course-reader, so students can follow along The next cohort is Sep 5. Links to course and discount: https://maven.com/parlance-labs/evals?promoCode=evals-info-url
Activity on repository
hamelsmu forked hamelsmu/dslop from aaazzam/dslop
View on GitHubRT Shreya Shankar Has anyone cracked the code of LLM speak? it's truly atrocious I recently added a "deslopify" command to my plain-writing skill to get the agents to write responses better https://github.com/docwriter-org/plain-writing-skill not perfect but better than nothing h/t @bradenjhancock for the idea
I've developed a physiological response to text like this. Eyes absolutely glazing over. It could be a message to myself sent from the future and the words would pass right through me.
RT Shreya Shankar 💯 Check out our skill https://github.com/shreyashankar/error-discovery-skill
One of the best tips from @HamelHusain and @sh_reya Evals course is to build a data labelling tool for every AI app you make. It takes no time at all for an LLM to code one up.
RT Bryan Bischof fka Dr. Donut Can I say something without everyone getting mad? As a preface: I’m a card carrying pure mathematician. I’ve never thought math was irrelevant for it’s uselessness, nor valuable for its utility. Math is art. There exist questions that humans ponder, some times those questions are answerable via practice like stitching thousands of zip ties onto a mesh and staring at it from 20ft back. Other times the questions require a bunch of weird symbols that have imbued meaning by others. I like math because, when I read questions and answers in that system I feel a connection with others who have asked and answered them. Like other art, sometimes the questions are the result of humans interacting with the world. Mathematical physics and Theoretical computer science fall into this description for me, sorta like photography. It’s not clear to me that art of that kind is better or worse than art that purely explores questions that come from within. Some artists feel best when others appreciate their art for its technique. Some want money. Others want the viewer to feel exactly what they felt. But some don’t really care, and like exploring the questions and attempting to answer them. Making art will change as the technology changes. Some art will feel strange to focus on in the future, but some artists will anyhow. Few artists will want to simply create exactly what’s been done before, but not zero. Many artists will attempt to run away from what technology can produce, but many artists won’t mind. Technology for art will have other effects. Photography will change how art is useful, then later become artistic itself, then later still be commoditized and people will think that makes it unartistic but art photography will still exist, and portrait painting will still too. Most people won’t practice doing art. Many people will consume art. Some will produce art. Artists will ultimately do what they always did: ask questions that he...
My favorite evals tool is a database
RT Austin Huang I’m happy to share what we’ve been working on recently at Collaborative Computing Inc. Our first product is live - http://collaborate.dev is a shared visual desktop for teams of multiple agents and people - http://collaborate.dev ^^^ try it now
If SaaS and software is dead, why is the Codex desktop app so valuable and different? Narrator: its not dead
😂🎉
@xlr8harder closed models are cutting their prices 80% to compete with small open models. this proves that open models have no impact and everyone just wants frontier intelligence for everything
View quoted postHappening in 10 minutes!
Final lesson in the series: How to turn eval results into a better model ✨ with @willccbb and @xeophon If you've put effort into evals, you should also consider if customizing your own model is right for you. I can't think of any better people to walk us through this, link to
RT dex if your business is struggling, building a software factory ain’t gonna save you. If your business is doing well, it should be obvious what to automate and in what order
RT Isaac Flath More stuff in the AI enabled notebook space! I am particularly excited about this one, because I really really loved R studio and this is made by the same group. I've been following it for a while and very excited to see it stable and released 🤩 https://opensource.posit.co/blog/2026-07-29_positron-jupyter-notebook-editor-ga/
The Positron Notebook Editor is out of beta and ready for all your notebook workflows! I am super proud of the work the team did to build this notebook experience from scratch. Something we decided was necessary to have a fantastic notebook experience for data science.
View quoted postThis new AI Notebook IDE looks really promising from the folks at @posit_pbc (the folks who made the beloved RStudio). It's super polished and has features like allowing you to control which cells the AI sees. Elastic License 2.0 https://opensource.posit.co/blog/2026-07-29_positron-jupyter-notebook-editor-ga/
Today @AAAzzam will walk through how to setup agent sandboxes in @modal , my favorite compute platform Modal has wicked fast cold start times. The devex makes remote compute feel local Link: https://maven.com/p/0684ab/don-t-build-agents-build-environments-instead notes/recording sent to people who sign up
RT Erik Drouhard Looking for my next role. Here’s what I’m about: I work at the intersection of design and code. I spent seven years building conversational AI tools—first at Nuance, working on Mix.dialog and the Verse design system, and later on Microsoft Copilot Studio. Most recently, I’ve been on Microsoft’s CoreAI team, prototyping LLM-powered experiences across Microsoft Foundry and the Azure portals. What I care about: making intelligent systems understandable, predictable, and useful—and figuring out how humans should actually work with agents. I do my best work in service of a team, using prototypes to answer hard questions early so everyone can move with more confidence. What I’m looking for: AI product design, design engineering, UX engineering, or hybrid roles where interaction design, product thinking, and technical execution genuinely overlap. I’m based in Worcester, MA and open to remote or Boston-area roles. Work and case studies: http://erikdrouhard.com DMs are open.
No slides 🥰
This will be a (virtual) whiteboard lecture! My last lecture in the free summer of evals series 😔 but it has been really fun to participate and see many folks attend and learn!
View quoted postHappening in 10 minutes!
If you're using LLMs for classification, check out this talk. @sh_reya shows how to optimize costs while maintaining performance. Classification is increasingly important in AI workflows, especially w/model routing. Those who sign up get the recording + notes
RT Shreya Shankar This will be a (virtual) whiteboard lecture! My last lecture in the free summer of evals series 😔 but it has been really fun to participate and see many folks attend and learn!
If you're using LLMs for classification, check out this talk. @sh_reya shows how to optimize costs while maintaining performance. Classification is increasingly important in AI workflows, especially w/model routing. Those who sign up get the recording + notes
If you're using LLMs for classification, check out this talk. @sh_reya shows how to optimize costs while maintaining performance. Classification is increasingly important in AI workflows, especially w/model routing. Those who sign up get the recording + notes https://maven.com/p/35471f/stop-paying-full-price-for-llm-classification
I think the E in FDE should be rebranded to educator. The whole point of AI is you can fish for yourself. Nudge towards upskilling instead of outsourcing
If you don't subscribe to Lenny's newsletter you are losing money. It's hard to find any offer with a stronger value / price ratio It's worth it just for the newsletter alone, but Lenny stacks it with 100s of free high quality AI subscriptions. I've saved > $5k this year
I continue to invest in http://LennysProductPass.com because if you've seen my recent research on the state of tech workers, there's a growing divide between people who are having the most fun and success of their entire careers thanks to AI, and people who are feeling stressed and
View quoted postRT Lenny Rachitsky I continue to invest in http://LennysProductPass.com because if you've seen my recent research on the state of tech workers, there's a growing divide between people who are having the most fun and success of their entire careers thanks to AI, and people who are feeling stressed and confused about their futures. For over seven years now, I’ve been sharing advice, building community, and capturing insights from industry leaders, but for many people, *access* is often the trickiest part. These tools are expensive, and there are so many of them. It’s hard to know where to invest your time and money. That’s why the curation of the Product Pass—only including products I genuinely love and recommend—and making every partner offer a full free year, are so important. A typical free trial gives you only enough time to poke around and get at the surface level of a product. But with a full year of access, you and your team can build real things, form durable habits, and have enough time to figure out which tools deserve a lasting place in your stack. Also, this time around, because life is about more than work, I’m expanding the Product Pass to include a free year of my favorite personal tools, like @wakingup, @Readwise, http://brain.fm, and @Mercury Personal. Here's everything you need to know about what's new in the Product Pass and how to grab the deals: https://www.lennysnewsletter.com/p/productpass-summer2026launch
If you thought your Lenny's Newsletter subscription couldn’t get any better, you ain’t seen nothing yet. Today, I’m adding 11 more incredible products to Lenny’s Product Pass: 1. @RunwayML 2. @WakingUp 3. @Higgsfield 4. @BrainFMapp 5. @Mercury (Personal) 6. @Resend 7.
View quoted postRT Randy Olson Ege and I are opening 10 slots this week for people stuck on an AI agent problem. You bring something real: an agent that won't behave, evals you can't trust, a skill your team can't share. We work through it with you for 30 minutes, and it costs nothing. Signup link in the replies.
Coding harnesses are super sticky Trying to pass an open weights model into a closed harness is not smooth I tried using K3 with Codex Desktop. Big mistake LOL
The fastest way to burn money is Sol-Ultra-Fast I don't know who can afford that combo
RT Shreya Shankar hot take: sometimes I wonder if the uncanny valley is actually what we need, so we can more easily distinguish human-AI interaction from human-human interaction. for example, writing. when I read something that feels mostly AI produced, I evaluate it differently. instead of asking, “do I trust this person’s taste and understanding of the topic? so much that if I spend time reading this, will I learn something interesting?” I switch to asking, “am I interested in this topic? and does the AI probably know more than I do here, without being so sloppy that I can still understand it?” in short, the uncanny valley is a good cue to tell us how to evaluate or make sense of the information; the uncanny value cues the appropriate epistemic frame
Voice models love replying with "Exactly!" even when it isn't called for
View quoted post