@LangChain Always hiring: https://t.co/D5Ut3loFO7
“Pair this with the right observability” Langsmith 🔥
@hwchase17 Pair this with the right observability and the system is unstoppable
View quoted postRT Prakalp Choubey Re @hwchase17 Pair this with the right observability and the system is unstoppable
when software makes decisions at runtime instead of build time, a 200 OK can still be a complete failure. Here is why traditional SDLC will not work and we need a new ADLC. https://x.com/i/article/2093398626408304640
View quoted postModels labs will create great harnesses and ecosystems for their models but will block model access to harnesses of other labs Only choice for a harness that works across models is one not associated with a lab Long live LangChain
We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care
View quoted postFollow @colifran_ for all things wiki! Wikis are a great way to represent knowledge in a simple and easy to understand way
seems remarkably similar to langchain wiki? also interesting how everything is collapsing to some shape of a loop and instruction file
RT george salapa 🜁 seems remarkably similar to langchain wiki? also interesting how everything is collapsing to some shape of a loop and instruction file
Would love feedback on our support for the new mcp spec!
we're adding support for the new MCP spec in @LangChain open source! the new API is built on top of FastMCP, so you can take advantage of their excellent devx for building MCP servers and clients. an early version is available in langchain==1.4.0a2; let us know what you think!
RT LangChain In 13 minutes, @jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs a month. Topics covered: ✅ Their four quadrant eval framework ✅ Why closing the production-to-eval loop is the hardest part ✅ How a data lake and long context changed what agents can do with data
RT LangChain .@PodiumHQ's agents seemed broken. LangSmith showed the real story: the agent was behaving rationally based on the context it had. Principal Software Engineer Walker Ward on tracing agent reasoning end to end.
this is why benchling will create the best ai agent for science, not ant/oai strongly believe in a many-agent world, each one domain specific
openwiki now supports okf 2.0
the claims runtime in openwiki v0.4.0 gives your wiki a persistent source of truth to help it forget and self-correct. but how would a reader know: 1⃣ what sources a page was built from? 2⃣ if a page is fully verified? 3⃣ who generated the page in the first place? to answer
RT Colin Francis the claims runtime in openwiki v0.4.0 gives your wiki a persistent source of truth to help it forget and self-correct. but how would a reader know: 1⃣ what sources a page was built from? 2⃣ if a page is fully verified? 3⃣ who generated the page in the first place? to answer that we migrated openwiki to support OKF v0.2 the claims runtime powers page-level sources and verification and openwiki separately stamps who generated the page and when. claims=content-level truth okf=page-level trust special thank you to @seo_dev2142 for kicking off okf v0.2 support in openwiki!
RT Caspar 𝚍𝚎𝚎𝚙𝚊𝚐𝚎𝚗𝚝𝚜 is evolving
Deepagents is becoming a multiplayer harness Auth, memory, etc If you’re building an agent that you want to expose to multiple users in same thread and are thinking about these issues - reach out! Would love to chat
View quoted postRT Vapi We’re excited to announce that 𝗛𝗮𝗿𝗿𝗶𝘀𝗼𝗻 𝗖𝗵𝗮𝘀𝗲 @hwchase17, Co-Founder & CEO of @LangChain, will be speaking at VapiCon 2026! 🎙️ Harrison co-founded LangChain with the goal of making it easier to use LLMs to develop context-aware reasoning applications. LangChain also builds LangSmith, a platform for agent engineering. Prior to starting LangChain, Harrison led the ML team at Robust Intelligence and the entity linking team at Kensho. He also studied statistics and computer science at Harvard. Join us in San Francisco this November to hear from Harrison and other leaders shaping what’s next in AI and voice. 📍 Fort Mason, San Francisco 📅 November 11–12, 2026 🎟️ Early bird tickets are limited — get yours at https://www.vapicon.ai/
How we think about evaluating wikis Including whether wikis are actually even helpful!
RT Nick Hollon the evals grind continues! this time on OpenWiki. we built WikiBench to answer two things: 1. how good is a given wiki? 2. does the wiki help at all? for the second, we ran the same questions three ways: wiki only, source only, and both. read on to find out which won.
RT Colin Francis a question we were interested in answering for openwiki 0.4.0 was: how do we make openwiki available to more devs? our answer was simple: coding agent integrations! you can now drive openwiki init and update using claude, codex, and opencode with almost no set up needed: 1⃣ npm install -g openwiki@latest 2⃣ openwiki integrations install claude|codex|opencode 3⃣ "use openwiki to create a wiki for this repo" or "use openwiki to update my wiki for this repo" go try openwiki. it is easier than ever. watch me drive openwiki update using codex 👇
Deepagents is becoming a multiplayer harness Auth, memory, etc If you’re building an agent that you want to expose to multiple users in same thread and are thinking about these issues - reach out! Would love to chat
RT Viv word on the street is that Ben Franklin used LangSmith to mine his traces and turn them into environments to hill-climb agents for science
RT Benjamin Tannyhill Big step forward for Engine! Some things I'm excited about in this release that we've shipped for our most avid users: - Engine now points more obviously to the issue in a trace. Engine identifies and clusters errors, but now helps you more quickly confirm its findings by guiding you to places of interest in your agent's logs. - We find a lot of our customers living entirely out of Slack + Linear. Engine now sends alerts to Slack, and syncs with Linear tickets, allowing you to continue tracking issues in whatever system is best, use your coding agent of choice, etc. - Engine's proposed fixes are much more effective at solving the issues that Engine has identified. Our biggest users are using Engine not only to find problems but to resolve them quickly and avoid asking c-bro to TAL at yet another problem. - We've introduced Analysis Levels (Reduced, Standard, and Extended) so your team can get the "right" amount of Engine based on your budget. DM if you have feedback!
LangSmith Engine now offers >2x performance on key internal benchmarks. Engine has already helped identify tens of thousands of issues in our customers’ agents. Now it offers: ✅ More accurate issue detection, clustering, and remediation ✅ Support for SaaS and self-hosted
View quoted postRT LangChain LangSmith Engine now offers >2x performance on key internal benchmarks. Engine has already helped identify tens of thousands of issues in our customers’ agents. Now it offers: ✅ More accurate issue detection, clustering, and remediation ✅ Support for SaaS and self-hosted deployments ✅ Reduced Analysis mode for cost-sensitive customers ✅ Integrations with @slackhq and @Linear ✅ Automatic closing of stale issues Learn more → https://www.langchain.com/blog/new-in-langsmith-engine-2x-better-issue-detection
its relatively easy to generate a wiki the first time, but how do you update them reliably? openwiki v0.4.0 improves in this dimension - it "forgets" better
openwiki v0.4.0 just went live 🎉 the focus? forgetting, accessibility, trust 1⃣ stale docs detection with forgetting and self-correction 2⃣ drive init and update through claude, codex, and opencode 3⃣ okf v0.2 support for provenance and verification 🧵1/4
View quoted postRT Colin Francis openwiki v0.4.0 just went live 🎉 the focus? forgetting, accessibility, trust 1⃣ stale docs detection with forgetting and self-correction 2⃣ drive init and update through claude, codex, and opencode 3⃣ okf v0.2 support for provenance and verification 🧵1/4
RT Nick Hollon Re @Vtrivedy10 and I have spent a lot of time thinking about how to efficiently build evals and environments. our eval-engineering skill detailed here helps spin the environment generation loop very quickly!
trying to think about how to most easily create evals we launched a skill to help with this iterative process
RT Viv this is a practical guide (+ updated skill) on how we use real world data like traces + human feedback to make synthetic environments + evals so we can measure, harness engineer, and post-train our agents a few main components: - World Knowledge gathering of the systems our agent will interact with - Using Specs to coordinate what will be built —> World Spec & Task Spec - Human-Agent collaboration to iteratively edit a Spec and infuse human feedback - An execution pipeline that takes an agreed on Task Spec and generates an environment + task - Fixing design flaws in the task and environment by actually running agents in the environment & mining the verifier scoring + traces @harrison_chase, @nick_hollon and I spent a ton of time going back and forth on design decisions and the overall flow - how much a coding agent should prompt for human feedback - what specs do we need and what do they contain - how to create and update world knowledge this new flow is packaged in our updated eval-engineering skill. we want to help every team own their intelligence which means owning the pipeline to convert their valuable data into Tasks+Environments that improve their agent over time check out the skill and reach out if you’re thinking about this! https://www.skills.sh/langchain-ai/langchain-skills/eval-engineering
RT Eric Rea Agree 100%. We started building agents and quickly realized we needed to build the system of record too. Agents need all the context and permissions you’d give an employee to complete work end-to-end. It goes both ways: the system of record makes the agent powerful, and agents force you to rethink the system of record.
Prediction: systems of record will need to become AI harnesses or face replacement by agents
View quoted postEvals matter
OpenRouter is an LLM gateway, not a router. Some of the investment in routers may be missing the plot. 80% of the work in building a good router is building a good eval. The routing logic is the easy part.
View quoted postA hard part of agents is connections - seems easy but hard to get right/seamless Managed deepagents tackles it for you More to come!
Starting today, a @LangChain's Managed Deep Agent deploy provisions your Slack app for you 🤯 No manifest. No OAuth redirects. No bot tokens to copy around. One command, and your agent says 👋 in Slack.
View quoted postRT Caspar just released managed deep agents 0.6.0. we now automate the ugly parts of deploying your agent to slack. it's like terraform for agents
RT Christian Bromann Starting today, a @LangChain's Managed Deep Agent deploy provisions your Slack app for you 🤯 No manifest. No OAuth redirects. No bot tokens to copy around. One command, and your agent says 👋 in Slack.
RT will brown skills are code, traces are data, environments are data, benchmarks are code, configs are code, one-shotted apps are data
you always gotta be asking whether it's code or data. applies to everything. json. markdown. tweets. group chats. emails. blogs. calendar notifications. zoom meetings
View quoted postRT Christian Bromann An agent is 3️⃣ layers. 👉 Business logic you write. 👉 A harness that runs the loop. 👉 Infrastructure that survives production. @hwchase17 breaks the whole stack in one video. https://youtu.be/TUJmfeGTr1Q
RT Viv launching another update to this skill soon - open holy grail question if anyone wants to riff! what goes in “How-to-make-tasks-harder[.]md”? some options in a list below👇 part of making a good eval is making sure it captures the real world. another part is making the actual task hard so there’s something to hill climb one way to calibrate “hard” is running a weaker model and smarter model, if everything passes all the time that’s not great -> perfect pass @ k isn’t a good learning signal a list of things that can make tasks harder, would love to compile more and hear other strategies - add more data (not super bullish on this as models are great brute forcers) - add information that needs to be discovered by search (ex: look across multiple tables to find missing information) - completeness (add many cases that all need to be passed, this works well. it can incentivize bad behavior though like over checking) - cross-domain tasks that combine 2 abilities. this is the best imo, agents are not good at this but it can look artificial and toyish humans are still useful here but would love to chat on thoughts
RT Sydney Runkle last night i spent some time building a browser agent with stagehand from @browserbase and deepagents from @LangChain i wanted to build something that required pretty complex web navigation to prove the value of agentic browser use! my agent plays a game called "map tap" where you are assigned a city around the globe, and have to drag the map, zoom in, and click as close to the given city as possible. you get points based on how close you are. 1. one shotted the agent w/ browserbase + langchain docs, it scored ~300/1000 points 2. asked my coding agent to review the trace and make it faster (smaller model) and more accurate (higher score) 3. came back 10 minutes later to see a perfect score screenshot in my trace here's a guide on how to build your own browser agent! https://docs.langchain.com/oss/python/integrations/tools/stagehand
RT LangChain OSS What's new in LangChain? 🚀 🧯 Standard exceptions: chat models now raise standard exception types across providers, so you can distinguish retryable errors (timeouts, rate limits, etc.) in a consistent way. Fully backward compatible. 🧩 Agent middleware: custom token_counter in ContextEditingMiddleware, plus non-retryable exceptions now skip retries in ModelRetryMiddleware (http://github.com/syyy44, http://github.com/Yigtwxx) 🔥 Fireworks document reranking: a new reranker integration (http://github.com/11adyy) 📊 Token & usage accounting: bug fixes for xAI and DeepSeek (http://github.com/aryansk, http://github.com/nazsats) ⚡ Portability & perf: lazy transformers import, and tighter grep scope for Anthropic (http://github.com/jtoman, http://github.com/Haaaarry) This is your harness! Thanks to everyone who made it better. 🙏
Open bot!
🎉 Introducing 𝙾𝚙𝚎𝚗 𝙱𝚘𝚝 An open source Grok Bot that works with ANY agent harness, designed for real companies. It includes: - AI Coworkers - Generative UI - Computer use (remote/local) - Agent-human handoffs - Full data recording, owned by you Repo →
View quoted postRT LangChain We are excited to be hosting Alex Atallah, Co-Founder & CEO of OpenRouter, at Interrupt NYC next month on September 24th! See the agenda and get your tickets: https://interrupt.langchain.com/nyc
very cool launch we did a webinar with jeff on "wiki" style memory and it's clear he'd thought about this problem a lot (webinar here: https://www.youtube.com/watch?v=Lsut4TCfygw)
I’ve been looking forward to today for 3 years Today we’re announcing Foundation - Chroma’s solution to memory Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https://trychroma.com/foundation
View quoted postRT Christian Bromann Managed Deep Agents is just getting started 🚀 @hwchase17 just published a set for banger videos that get you up to speed how shipping production agents today looks like.
🎓New YouTube playlist: Managed Deep Agents Gives an overview of Managed Deep Agents, and then each video dives deep into core concepts. Launching with six videos! 1⃣ Intro: https://youtu.be/xdrB53bgpp0 2⃣ Conceptual Overview: https://youtu.be/TUJmfeGTr1Q 3⃣ Quickstart:
RT Jeff Huber I’ve been looking forward to today for 3 years Today we’re announcing Foundation - Chroma’s solution to memory Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https://trychroma.com/foundation
RT Sydney Runkle so i'm currently hill climbing on map tap with my browser agent... any thoughts on 1) best model for simple image processing 2) what if any geographic context i should provide ? perfect score by eow would be pretty cool, not sure how fast an agent can do this though?
lots of observability & evals platforms not a lot of platforms that help close the loop and have an agent that suggests fixes, adds evals itself, etc
@hwchase17 BTW your NYC ads worked 😂 just learned that's why we chose LangSmith
RT YOЯNOC Re @hwchase17 BTW your NYC ads worked 😂 just learned that's why we chose LangSmith
RT Caspar The people who thrive in the AI age will be the ones who stay curious, keep learning, and use these new tools to expand what they're capable of I am proud to work at a company that values this mindset and creates resources like this to help others do the same
🎓New YouTube playlist: Managed Deep Agents Gives an overview of Managed Deep Agents, and then each video dives deep into core concepts. Launching with six videos! 1⃣ Intro: https://youtu.be/xdrB53bgpp0 2⃣ Conceptual Overview: https://youtu.be/TUJmfeGTr1Q 3⃣ Quickstart:
top tier webinar tmrw! @Vtrivedy10 (@LangChain) @willccbb (@PrimeIntellect) @AEllisBloor (@baseten) and I will will be jamming on how to automate more parts of the agent improvement loop (eval and environment engineering in particular) Come join us: https://events.langchain.com/webinar/Towards-Automating-Eval-and-Environment-Engineering/
+1 to this part of this is due to coding agent standards - agents.md and skills are just markdown files/directories!
agents are starting to look less like apps and more like directories instructions. skills. tools. memory. identity. channels. schedules. evals. the harness + runtime become infra underneath it all. really like where this abstraction is going
View quoted post🎓New YouTube playlist: Managed Deep Agents Gives an overview of Managed Deep Agents, and then each video dives deep into core concepts. Launching with six videos! 1⃣ Intro: https://youtu.be/xdrB53bgpp0 2⃣ Conceptual Overview: https://youtu.be/TUJmfeGTr1Q 3⃣ Quickstart: https://youtu.be/L54dR9qMKzc 4⃣ Instructions and Context Hub: https://youtu.be/8HIoZV7qlwA 5⃣ Skills: https://youtu.be/Lhru0yMI2as 6⃣ Tools: https://youtu.be/piBwvWgHqkY Playlist link: https://www.youtube.com/playlist?list=PLKYx7JmJ_IcA
RT Sergio R. agents are starting to look less like apps and more like directories instructions. skills. tools. memory. identity. channels. schedules. evals. the harness + runtime become infra underneath it all. really like where this abstraction is going
channels are how you interact with your managed deepagents eg slack pretty diagram 👇
View quoted postRT Mason Daugherty Allow me to introduce you to Managed Deep Agents * Works with all models * Connects to your company knowledge via a version-controlled, environment-aware ContextHub backend * Channels with Slack (& others) https://docs.langchain.com/langsmith/python/managed-deep-agents-overview
It's actually crazy watching everyone pivot in the complete wrong direction from what companies want. Having an army of agents/bots is counterproductive. It's a vanity metric. What teams want is a shared workspace where they can: - work with any model - build a "company
View quoted postnew onboarding for managed deep agents
I love working with Sydney and team and you will too! Come join us
we're hiring open source devs @LangChain looking for people who are building at the frontier of agents and are excited to quickly develop ownership apply: https://www.langchain.com/careers?ashby_jid=74e5f9f4-e44a-4594-ba26-abdf71bf287d#explore-jobs
View quoted postRT Sean Kerner 😮 wow that’s huge 💰
🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirable agent behavior and attach feedback that you can use in your agent improvement processes. In our benchmark, our tuned model beat frontier
Slack is the best ux
@hwchase17 if you dont have ur agents working overtime in slack/teams , ngmi
View quoted postRT Satyabrat Singh LangSmith’s new Tuned Evaluators look pretty interesting. Perceived Error can flag agent mistakes in production by picking up on user corrections and unresolved conversations. What’s interesting is it scores the whole conversation rather than individual steps and apparently does this cheaper and more accurately than a frontier model. Can make life easier for agent developers…
🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirable agent behavior and attach feedback that you can use in your agent improvement processes. In our benchmark, our tuned model beat frontier
channels are how you interact with your managed deepagents eg slack pretty diagram 👇
RT sonil the real unlock is not cheaper evaluation it is making specialized evaluators cheap enough to run continuously on production traces so evaluation stops being a launch checklist and becomes a permanent feedback loop that actually improves agents over time
🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirable agent behavior and attach feedback that you can use in your agent improvement processes. In our benchmark, our tuned model beat frontier
🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirable agent behavior and attach feedback that you can use in your agent improvement processes. In our benchmark, our tuned model beat frontier models at 82% lower cost. https://www.langchain.com/blog/introducing-langsmith-tuned-evaluators-starting-with-perceived-error
RT LangChain As model capabilities converge, more of the differentiation may come from the system around the model: the harness, tools, environments, evals, and feedback loops used to improve it. Join @hwchase17, co-founder and CEO of LangChain, @Vtrivedy10 Applied Researcher at LangChain, @willccbb , Applied Researcher at Prime Intellect, and Aaron Ellis-Bloor Applied Researcher at @baseten for a live panel session on Automating Eval & Environment Engineering. We’ll explore the improvement loop and what open models and model-harness co-design mean for companies that want to own their intelligence. Register now: https://events.langchain.com/webinar/Towards-Automating-Eval-and-Environment-Engineering/
RT DrCAO | AIWeb3 | ComputeFlux "Agents = model + harness + context" is a much better mental model than treating the LLM as the entire product. The model provides capability, but the harness determines tools, permissions and behavior; context determines what the agent knows; evals determine whether the system improves rather than merely changes. This matters for ownership too. Owning intelligence is not just downloading model weights. It means having meaningful control over the complete system that selects models, handles data, executes actions and learns from results.
gave a talk "owning your intelligence" - ty @sequoia @sonyatweetybird for having me talked about harnesses and evals and the role they play in owning your intelligence TLDR: > agents = model + harness + context > model - own the weights using something like @FireworksAI_HQ >
RT Daryl W Top 2 videos I’ve watched on harnesses.
gave a talk "owning your intelligence" - ty @sequoia @sonyatweetybird for having me talked about harnesses and evals and the role they play in owning your intelligence TLDR: > agents = model + harness + context > model - own the weights using something like @FireworksAI_HQ >
RT Sonya Huang 🐥 The race for the AI application layer is not only about UI, workflows, or GTM... it is a fight for the intelligence layer itself. https://x.com/i/article/2088383363795341312
RT Ebrahim Re @hwchase17 not every agent needs code execution, mine mostly needs read write on inventory data and that fake backend covers all of it
totally agree! here's how we architected deepagents to enable this deepagents runs connected to a "backend". this backend needs to expose filesystem like operations, but it does not have to be a filesystem. it could be a database, object storage, or a real filesystem - it just has to expose read/write/edit etc operations this backend could also be what we call a "sandbox". if a sandbox, it needs to expose an "execute" command which lets it execute code this backend is SEPARATE from where the agent loop runs. this allows us to "separate the brains from the hands" (https://www.anthropic.com/engineering/managed-agents) deepagents is built on top of langgraph, which means we can easily deploy it with MCP, a2a, and other standard endpoints we use this architecture to power many different types of experiences first, we can create a classic TUI like coding experience. we do this by giving deepagents a "sandbox" that is running locally in the same directory; deloying deepagents locally behind a light weight server; and then connecting to it with the TUI acting like a frontend. see dcode for an example of this https://docs.langchain.com/oss/deepagents/code/overview second, we can create a cloud coding experience. we can do this by running deepagents on LangSmith deployments for a production scale deployment, and connecting to a sandbox running on modal, daytona, e2b that is running elsewhere. we can then build a frontend to connect to langsmith deployments and let users interract with it there, and also expose it in slack to let users interract with it there. note: both slack and web ui connect to the same backend, so you can switch between them seamlessly. code: https://github.com/langchain-ai/open-swe of course - deepagents can be used to create agents that are NOT coding agents. a lot of agents still need to write and execute code, so this architecture is still very useful. but for some the code execution is overkill, and thats where you can swap to a "fake" backe...
I love agentic coding harnesses, but they shouldn't be primarily terminal-based. The terminal is great for quick and precise commands, but information density is extremely low and UI affordances are minimal. Maybe provision of TUIs is worthwhile for occasional use (when
View quoted postRT QUEEN 👑 Re @hwchase17 @LangChain I built an open-source CLI that scaffolds, runs, and deploys production LangChain agents — frontend and backend in one command. Every agent project starts the same way: wire up a LangGraph backend, scaffold a Next.js frontend, figure out CORS, hide API keys, set up a proxy. Two days gone before you write a single line of agent logic. langctl kills that. langctl new my-agent cd my-agent langctl dev That's it. Agent + chat UI running at localhost:3000. One URL. No CORS. No exposed keys. How it actually works: langctl new scaffolds a LangGraph Agent Server backend and a Next.js frontend, wired together through a built-in proxy. The browser only ever talks to /api/agent/* on the frontend's origin — the proxy attaches x-api-key server-side. Zero cross-origin requests, ever. langctl dev starts both processes, waits for the agent's health endpoint before booting the UI, and tears everything down cleanly on Ctrl-C. Like next dev, but the backend is an agent. Three chat UI options out of the box: assistant-ui (default) — threads, branching, composer runtime minimal — one hand-written Chat.tsx, zero UI deps ai-elements — shadcn-registry components (experimental) Other commands: langctl sync — regenerates langgraph.json from agent.yaml, preserves hand-written keys langctl doctor — verifies toolchain, ports, keys, and config before something breaks agent.yaml is your single source of truth. Dev and production differ by one env variable. The frontend source doesn't change. Phase 1 (scaffold + unified dev runtime) is live on PyPI. Deploy providers (LangSmith Cloud, Vercel) are next. Apache-2.0 licensed. Generated projects carry no license obligation to this tool. GitHub: https://lnkd.in/dcTqBPkE PyPI: http://pypi.org/project/langctl pip install langctl and build your next agent in minutes, not days. LangChain Langchain Community for Developers, India LangChain JS LangChain Developer LangChain Labs Formation LangChain FranceLan...
RT Caspar explaining some managed deep agents terminology channel: a connection between your agent and an external messaging service, e.g. slack- tag your agent in slack and receive its reply there sandbox: isolated execution environment with its own filesystem where your agent can run shell commands safely. MDAs use one by default middleware: hook-based way to extend and control the agent loop, e.g. summarize conversation before model. see our docs on prebuilt middleware for many great examples evals: repeatable tests that run your agent on example tasks, grade its performance, catch regressions as it evolves all covered in our docs! http://langch.in/mda
@caspar_br Hey, can anyone explain this? I don’t understand middleware, sandboxes, evals and channels in this context. Been using the rest.
View quoted postRT Ankush Gola Great talk from Harrison on “owning your intelligence”
gave a talk "owning your intelligence" - ty @sequoia @sonyatweetybird for having me talked about harnesses and evals and the role they play in owning your intelligence TLDR: > agents = model + harness + context > model - own the weights using something like @FireworksAI_HQ >
gave a talk "owning your intelligence" - ty @sequoia @sonyatweetybird for having me talked about harnesses and evals and the role they play in owning your intelligence TLDR: > agents = model + harness + context > model - own the weights using something like @FireworksAI_HQ > context - memory needs to be portable > harness - needs to be model agnostic. also needs to be good at bringing right context to llm. "right" context may depend on your use case, which is why an open/configurable harness helps > how to use middleware in langchain/deepagents to configure your harness > how to use langgraph to fully own your cognitive architecture > why evals/obs matters - some quotes from @satyanadella - “Create your private evals, because evals define what “good” looks like inside the organization” - “retain ownership of your organization’s memory, traces, feedbacks, decisions, and institutional context” - “you create your own continuous learning loop (i.e. hill climbing machine) that will allow your AI investments to compound the value of your firm” > how to use harbor for evals > tracing is important > evals + observability only matter so you can set up a data flywheel > data flywheel = run agent -> collect traces -> find interesting traces -> use those to improve > demo of langsmith engine which does exactly this! full video: https://www.youtube.com/watch?v=HI2q3ci3Iuc&list=PLaqC3GACblSs&index=2
RT 🎭 I wrote on this two years ago
basically every company will own their pipeline of using Trajectories to mine knowledge + data to improve (harness eng + finetune) their agents example of things you can do: - find good examples, use them to distill a smaller, cheaper model - make environments/tasks from
View quoted postIt’s beautiful
some great new LangSmith docs on traces vs threads vs trajectories (new concept) observability data is no longer just for observability - its also for memory & learning having a really clear mental model of this data is incredibly helpful! https://docs.langchain.com/langsmith/observability-concepts
new langchain oss release!
What's new in LangChain? 🚀 🔌 Support for OpenAI's 3.0 SDK (using httpx2) ✨ Support for gemini-3.7-flash Plus a wave of core reliability fixes from external contributors: 🛠️ Tool calling & structured output: clearer errors, respect for pydantic aliases when validating tool
View quoted postRT LangChain OSS What's new in LangChain? 🚀 🔌 Support for OpenAI's 3.0 SDK (using httpx2) ✨ Support for gemini-3.7-flash Plus a wave of core reliability fixes from external contributors: 🛠️ Tool calling & structured output: clearer errors, respect for pydantic aliases when validating tool inputs (http://github.com/RinZ27, http://github.com/kinch-tech, http://github.com/charan16) 📊 Usage metadata fixes: token usage callback bugs, and cost metadata when streaming OpenRouter (http://github.com/Solaris-star, http://github.com/iroiro147) 🧱 Content & prompt hardening: guard malformed Anthropic content blocks, and preserve non-str/non-dict items in prompt templates for content blocks (http://github.com/Oxygen56, http://github.com/zerafachris) Huge thanks to everyone who contributed! 🙏
RT Viv notes on "How to do data" from th Zai team that others should steal as they build their own environments - "Research agents collect task patterns from real work and turn them into runnable long-horizon environments" good environments come from faithfully simulating the real world. You already own a lot of that data in your Traces. Agents can use those + other data to bootstrap environments everyone can own this process for themselves by carefully using coding agents - "these pipelines still require a meaningful amount of human-in-the-loop" basically...look at the data one of the highest value things builders can do is selectively infuse their domain knowledge into the environment building process, this means inspecting rollouts and verifier structure to make sure it aligns with stuff they care about Agents help automate a lot but without humans the synthetic data pipeline doesn't worth end to end....yet
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: https://z.ai/blog/glm-5.3
RT Justin Lin Have been thinking about building a custom harness because nothing out there quite meets my needs for scaling a fully hybrid human + AI company, especially when it comes to compounding learning loops. This talk by @hwchase17 of @LangChain is a great intro to evals, traces and harness engineering.
An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. @LangChain founder @hwchase17 joined us at our @sequoia Own Your Intelligence to talk about the piece that often gets the least
View quoted postRT Mason Daugherty LangSmith Engine ;) https://www.langchain.com/langsmith/engine
agent traces are the most important new input to a self-improving company. but a pile of traces doesn't improve anything on its own. the full loop takes three components: 1. a company brain - shared memory of what the team knows: decisions, context, the why behind the work 2.
View quoted postRT Austin Hughes our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+ cost savings for customers great talk by @hwchase17 and @HeggieConnor
New Max Agency with @unifygtm CTO & co-founder @HeggieConnor. Some of my favorite tidbits: - how they cut 90-95% of costs two weeks before launch - why their subagents are just a function call. 🎧 Apple: https://podcasts.apple.com/nz/podcast/how-unify-cut-its-ai-agent-costs-95-in-two-weeks/id1891551672?i=1000783151404 🎧 Spotify: https://open.spotify.com/episode/6kWQouc2QmiHGk0vdiZEtd?si=ba6e241ca4a24faf ⏯️ YouTube:
View quoted postRT Tanuj Prakash Agents running in the background will not be called agents running in background. They'll just be productivity enhancers working for you while your main job is thinking. This is the right step for the future.
agents running in the background will be the future - lets work scale beyond people prompting them directly crons are one way to do this. first class support in managed deepagents
View quoted postRT Jim Very important explanation from @hwchase17 and helpful for enterprise leaders trying to figure out their AI architecture plans
Can use a lightweight classifier step to first decide if even worth running Then more expensive agent if that criteria is met
agents running in the background will be the future - lets work scale beyond people prompting them directly crons are one way to do this. first class support in managed deepagents
you can build agents that prompt themselves by giving them a schedule. with Managed Deep Agents, that only takes a few lines of code code snippet is from our shiny new MDA docs: https://langch.in/schedules
RT Caspar Broekhuizen you can build agents that prompt themselves by giving them a schedule. with Managed Deep Agents, that only takes a few lines of code code snippet is from our shiny new MDA docs: https://langch.in/schedules
talked about how harnesses & evals help you own your intelligence how they fit in to the big picture: owning your intelligence means three things: - open agent system (harness is a big part of this!) - compounding loop (evals are a big part of this!) - governed runtime (harness also important here - we see managed harnesses growing rapidly)
An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. @LangChain founder @hwchase17 joined us at our @sequoia Own Your Intelligence to talk about the piece that often gets the least
View quoted postone of the strengths of managed deepagents is how easy it is to define all the pieces of your production agent we updated our docs to reflect this. each component is its own file. you can easy click through to see where it lives, what it looks like, etc https://docs.langchain.com/langsmith/python/managed-deep-agents-overview#example-agent
RT Julia Schottenstein Had to marry him for @segall_max to write this @LangChain ftw!
If Albert Einstein were developing agents, he’d FOR SURE be building with @LangChain 🦜🦜🦜
RT Christian Bromann You don't stand up an agent. You upload a folder. Claude Code is a harness for your laptop. Managed Deep Agents is a harness for production. Want Slack? Add a file. Want a daily run? Add a file. Want memory? Add a file. LangSmith runs the harness. https://docs.langchain.com/langsmith/python/managed-deep-agents-overview
New Max Agency with @unifygtm CTO & co-founder @HeggieConnor. Some of my favorite tidbits: - how they cut 90-95% of costs two weeks before launch - why their subagents are just a function call. 🎧 Apple: https://podcasts.apple.com/nz/podcast/how-unify-cut-its-ai-agent-costs-95-in-two-weeks/id1891551672?i=1000783151404 🎧 Spotify: https://open.spotify.com/episode/6kWQouc2QmiHGk0vdiZEtd?si=ba6e241ca4a24faf ⏯️ YouTube: https://youtu.be/6898VdRtKDE
RT LangChain Next stop on the LangSmith Roadshow: Silicon Valley. ✅ A half-day, hands-on look at LangSmith Engine ✅ Hear from @hwchase17, @Expedia, + more ✅ Get hands-on experience building + deploying RSVP today: https://events.langchain.com/LangSmithRoadshow/SiliconValley/
RT Brace MDA makes building any type of agent super easy. Checkout @caspar_br new video on how to build a social media agent!
📹Build a social media agent This Managed Deep Agent scans Hacker News and optional X, drafts three posts, saves them to durable memory, and sends them to Slack. Learn how to use: > Slack channel integration > Custom tools > Memory Video: https://youtu.be/OpFXXSsEIBo
📹Build a social media agent This Managed Deep Agent scans Hacker News and optional X, drafts three posts, saves them to durable memory, and sends them to Slack. Learn how to use: > Slack channel integration > Custom tools > Memory Video: https://youtu.be/OpFXXSsEIBo
RT Simon Budziak LangChain is back in Poland. Two cities, one week apart. Warsaw, 15 Sept: https://luma.com/rc0d3pve Wrocław, 22 Sept: https://luma.com/yg8tn5dv Four talks each night on agents in production. Free. Thanks @amadaecheverria @coffeewlthkaran @hwchase17 @LangChain team
very helpful
im making a massive "Managed Deep Agents 101" video, diving deep into all the core concepts of managed deep agents: https://docs.langchain.com/langsmith/python/managed-deep-agents-overview would people prefer a single 2hr+ video, or a youtube playlist of like 12 ten minute videos?
View quoted postLangSmith Engine is hitting the road and the rails in NY + SF If you spot our billboards in-person over the next few months, send it our way!
really excited to see this release AND integrate it with deepagents!! https://www.langchain.com/blog/switchyard-agent-routing-benchmark
Lightning strikes for continuous and long-run agents! Nemotron 3.5 Lightning is smart, fast, efficient and open.
View quoted postim making a massive "Managed Deep Agents 101" video, diving deep into all the core concepts of managed deep agents: https://docs.langchain.com/langsmith/python/managed-deep-agents-overview would people prefer a single 2hr+ video, or a youtube playlist of like 12 ten minute videos?
RT LangChain Already building voice agents? Join our upcoming webinar to learn how to evaluate them across execution, outcomes, and experience—from tool use and task completion to latency, interruptions, and conversational friction. https://events.langchain.com/webinar/How-to-evaluate-voice-agents/
RT Sydney Runkle we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap!
We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Prime Agent, Deep Agents) on 30 challenging agentic tasks. Pi Agent was the cheapest harness and passed the most tasks 🧵🧵