<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Turing Post</title>
    <description>Over two decades in tech, with the last six years focused on ML and AI. Our analysis stays precise and grounded. Our educational series walk you through the foundations and help you explore the deeper layers. We trace the arc of AI – its past, its present, and the direction it’s pulling us toward. We track the research that matters, the systems being built, and the ideas that define how AI actually works. And we break it down with clarity, so you can make better decisions. Join 115,000+ professionals who rely on Turing Post.</description>
    
    <link>https://www.turingpost.com/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/UJIoBuf5BX.xml" rel="self"/>
    
    <lastBuildDate>Fri, 10 Jul 2026 19:14:27 +0000</lastBuildDate>
    <pubDate>Sat, 04 Jul 2026 01:23:58 +0000</pubDate>
    <atom:published>2026-07-04T01:23:58Z</atom:published>
    <atom:updated>2026-07-10T19:14:27Z</atom:updated>
    
      <category>Machine Learning</category>
      <category>Artificial Intelligence</category>
      <category>Technology</category>
    <copyright>Copyright 2026, Turing Post</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/c0879691-e34f-4a23-8b4b-b2d9c313d91d/Favicon.png</url>
      <title>Turing Post</title>
      <link>https://www.turingpost.com/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>The Org Age of AI: A Collection of Enterprise AI Adoption Guides</title>
  <description>A complete guide to our Org Age of AI series: AI ROI, workflow redesign, AI-native startups, enterprise maturity, AI flywheels, hybrid AI, and spec-driven development.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9d6f8d44-38d2-403f-b469-012d980cd3da/agents.jpg" length="90680" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides</guid>
  <pubDate>Wed, 08 Jul 2026 04:00:00 +0000</pubDate>
  <atom:published>2026-07-08T04:00:00Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[The Org Age Of Ai]]></category>
    <category><![CDATA[Ai 101]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">What is an AI-native enterprise?</h2><p class="paragraph" style="text-align:left;">An AI-native enterprise is an organization designed so AI can understand, operate, and improve its work. It is not a company that simply gives employees access to chatbots or adds agents on top of old processes. It is a company whose workflows, data, tools, permissions, feedback loops, and business rules are structured so machines can participate in the work reliably.</p><p class="paragraph" style="text-align:left;">In practice, almost no large enterprise is fully AI-native yet. Most companies are still in the transition stage: they are making their operations more legible to AI, redesigning workflows, building verification layers, and deciding which tasks should run through humans, agents, cloud models, local models, or hybrid systems. That transition is what we call the Org Age of AI.</p><p class="paragraph" style="text-align:justify;"><b>TL;DR:</b> Enterprise AI adoption in 2026 is not about adding one more chatbot or agent. It is about redesigning workflows, making organizations legible to machines, building verification loops, and learning how to combine human judgment with AI systems that can actually compound.</p></div><p class="paragraph" style="text-align:justify;"><i><b><a class="link" href="https://www.turingpost.com/t/the-org-age-of-ai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank" rel="noopener noreferrer nofollow">The Org Age of AI </a></b></i><i>has become one of our most popular series. And it&#39;s easy to see why. Throughout the first half of 2026, we focused on the practical side of enterprise AI adoption – the topics very few people were talking about. Some of these ideas may seem obvious in hindsight, but the reality is that people across every profession need clear roadmaps and tips to apply AI effectively in their work and businesses.</i></p><p class="paragraph" style="text-align:justify;">Here, we&#39;ve gathered the complete collection of everything we&#39;ve covered so far, along with two additional articles that perfectly complement the series and further broaden your perspective.</p><h2 class="heading" style="text-align:justify;" id="1-deep-seek-m-hc-breaking-the-archi">#1: AI Feels Powerful. So Why Is the ROI Still Missing?</h2><p class="paragraph" style="text-align:justify;">In 2026, companies already have access to powerful AI models and agents. But the biggest gains won’t come from chasing the newest model alone. Much bigger profit comes from redesigning workflows around AI: encoding expert knowledge, feedback, and business rules into systems that can actually compound over time. Here are the tips to unlock the strongest returns from AI investments →</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/orgage1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> AI Workflow Redesign: Why ROI Needs Restructured Work </p><p class="embed__description"> Most AI pilots fail because companies bolt tools onto old workflows. Here&#39;s why AI workflow redesign is the real source of enterprise value and ROI. </p><p class="embed__link"> Turing Post • Ksenia Se </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/9b1ca244-3bad-41be-9ed1-7cba6f043524/OrgAge1.jpg?t=1781522057"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="3-conditional-memory-and-the-rise-o">#2: The Unsexy Truth of AI Adoption: A 5-Level Maturity Framework</h2><p class="paragraph" style="text-align:justify;">Many companies think they&#39;re one AI agent away from transformation. But in practice they are struggling to make their own organizations understandable to AI. Here is a practical 5-level maturity framework to adapt your companies to work with machines.</p><div class="image"><img alt="Levels of AI adoption in organizations" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b907fa04-1cc9-4cef-834e-2f243c315710/image.png?t=1783121835"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Turing Post</p></span></div></div><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/orgage2?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> #2: The Unsexy Truth of AI Adoption: A 5-Level Maturity Framework </p><p class="embed__description"> AI adoption fails when companies skip the middle layers. A practical 5-level maturity framework for making organizations legible to machines, with real use cases. </p><p class="embed__link"> Turing Post • Ksenia Se </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/f9bba2c6-86de-4b95-8729-84202f7294c5/OrgAge2_3_.png?t=1781372468"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="3-how-to-build-an-ai-native-startup">#3: How to Build an AI-Native Startup from Day One</h2><p class="paragraph" style="text-align:justify;">While we first discussed how to rebuild existing organizations for AI, it&#39;s now interesting to explore how to build a startup designed for AI from the very beginning. We outline the main principles →</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/orgage3?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> #3: How to Build an AI-Native Startup from Day One </p><p class="embed__description"> What actually makes a startup AI-native — and how to build one from day one. A practical framework of 5 principles: machine-legibility, tool portability & more. </p><p class="embed__link"> Turing Post • Ksenia Se </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/294deaa6-1153-46b9-8b8d-d4e416ee0970/OrgAge3.png?t=1781372457"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="4-on-policy-distillation-zeitgeist">#4: There are no AI-native enterprises</h2><p class="paragraph" style="text-align:justify;">Well, we talk about AI-native enterprises, but the uncomfortable reality is that none exist yet. The hardest part is untangling decades of hidden workflows, politics, and institutional habits that AI can&#39;t simply automate. </p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/orgage4?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> #4:There are no AI-native enterprises </p><p class="embed__description"> Why the enterprise AI problem is not a technology problem. In the series: The Org Age of AI </p><p class="embed__link"> Turing Post • Will Schenk </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/84cb371b-75d5-4dae-bbbd-40ed11754fdb/OrgAge4_2_.png?t=1781372388"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="5-the-inference-chip-wars-mat-x-taa">#5: AI Workflow Patterns: The Real Unit of AI Adoption in 2026</h2><p class="paragraph" style="text-align:justify;">As you&#39;ve probably noticed, we&#39;ve been talking a lot about updating workflows, not models, inside organizations. To help you avoid focusing on the wrong abstraction, we define it once again: the real unit of AI adoption is the workflow.</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> #5: AI Workflow Patterns: The Real Unit of AI Adoption in 2026 </p><p class="embed__description"> Seven agentic AI workflow primitives, eight production patterns, and a practical framework for deciding which workflows to automate first — with real use cases. </p><p class="embed__link"> Turing Post • Will Schenk </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/26e89a5a-26f8-4c82-a58b-9699ec8b8563/OrgAge5_1_.png?t=1781372319"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="8-transformers-depth-is-an-addressa">#6: The Flywheel: What Happens When Workflows Run Themselves</h2><p class="paragraph" style="text-align:justify;">Once you&#39;ve built the system, another layer emerges – one that has only recently become a practical reality. AI can now generate, test, and refine its own work without a human checking every step. But if a workflow learns from flawed metrics, it can repeat the same mistake faster with every cycle. That&#39;s why verification has moved to the forefront.</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/ai-flywheel-when-workflows-run-themselves?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> #6: The Flywheel: What Happens When Workflows Run Themselves </p><p class="embed__description"> AI flywheels are closed-loop workflows that generate, measure, and decide what to try next. Here’s why verification must come before autonomy </p><p class="embed__link"> Turing Post • Will Schenk </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/b8714153-5e18-4774-b3ea-eed2ed2422d9/OrgAge6.png?t=1781371259"/></a></div><hr class="content_break"><p class="paragraph" style="text-align:left;">Next, we have two practical articles that broaden your perspective on how AI can drive greater enterprise value.</p><h2 class="heading" style="text-align:left;" id="1-ai-101-from-vibe-coding-to-spec-d">1. AI 101: From Vibe Coding to Spec-Driven Development</h2><p class="paragraph" style="text-align:left;">As more companies let AI write production code in 2026, &quot;vibe coding&quot; is really no longer enough. Enterprise teams need concrete specifications, testing, and verification so coding agents can build software that stays reliable as projects grow. <a class="link" href="https://www.turingpost.com/p/sdd?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank" rel="noopener noreferrer nofollow">Spec-Driven Development </a>makes this work.</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/sdd?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> Spec-Driven Development vs Vibe Coding: Tools & Guide (2026) </p><p class="embed__description"> Why vibe coding breaks at scale — and how spec-driven development (SDD) fixes it. Covers Kiro by AWS, GitHub Spec Kit, Tessl, and when to use each approach. </p><p class="embed__link"> Turing Post • Alyona Vert. </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/55c2a8a9-733a-46d0-99d1-5e6c52e5b551/sdd.png?t=1772659337"/></a></div><h2 class="heading" style="text-align:justify;" id="2-llm-token-types-input-output-reas">2. What is Hybrid AI?</h2><p class="paragraph" style="text-align:justify;">Enterprise AI often fails because every task has different latency, privacy, and cost requirements. Deploying AI across thousands of employees, cloud-only quickly becomes expensive and slow. <a class="link" href="https://www.turingpost.com/p/hybridai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank" rel="noopener noreferrer nofollow">Hybrid AI</a> lets companies run the right workload in the right place instead of sending everything to the cloud.</p><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/hybridai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-org-age-of-ai-a-collection-of-enterprise-ai-adoption-guides" target="_blank"><div class="embed__content"><p class="embed__title"> What is Hybrid AI? </p><p class="embed__description"> It’s not about architectural hybrids, it’s about where to run a model. The big promise in connecting devices and clouds for the nearest future of AI. </p><p class="embed__link"> Turing Post • Alyona Vert. </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/6d8277fd-0b59-4666-b148-5b4fd1a1dd6f/Hybrid_AI.png?t=1769613166"/></a></div><h3 class="heading" style="text-align:justify;" id="heading-3"></h3><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Is your security team ready for AI coding agents? Join us on July 14🛡️</title>
  <description>When agents write and execute code autonomously, legacy tools break down. Here&#39;s the new playbook.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f442bd2a-ff7d-4562-b8eb-6ff3589ff3b2/image-47.jpg" length="75248" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14</guid>
  <pubDate>Thu, 02 Jul 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-07-02T21:00:00Z</atom:published>
    <category><![CDATA[From Our Partners]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;"><b>Announcing a super interesting workshop from our partners →</b></p><div class="image"><a class="image__link" href="https://go.rbrk.co/Secure-Claude-Code-wbnr?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14" rel="noopener" target="_blank"><img alt="Seciruty playbook" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b02f5d26-6b5b-46c0-a6fc-9976609af578/image.png?t=1782970487"/></a></div><p class="paragraph" style="text-align:left;">AI coding agents like Claude Code are transforming software development, but most security teams are still protecting against human-speed threats.</p><p class="paragraph" style="text-align:left;">When an agent can read, write, and execute code autonomously, EDR, DLP, and static rules-based controls simply weren&#39;t built for it. The attack surface has changed. <b>The security playbook has to change with it.</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://go.rbrk.co/Secure-Claude-Code-wbnr?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14" target="_blank" rel="noopener noreferrer nofollow">Join us on July 14 for a live webinar &quot;Securing Claude: Playbook for Governing Coding Agents</a> &quot; — a practical session on how to govern AI agents without slowing down your developers.</p><p class="paragraph" style="text-align:left;"><b>We&#39;ll cover:</b></p><ul><li><p class="paragraph" style="text-align:left;"><b>The Risk Landscape </b>— Real-world examples of coding agents going rogue</p></li><li><p class="paragraph" style="text-align:left;"><b>The Legacy Gap</b> — Why traditional security tools fail against autonomous AI</p></li><li><p class="paragraph" style="text-align:left;"><b>Actionable Blueprints</b> — A new reference architecture for AI security</p></li><li><p class="paragraph" style="text-align:left;"><b>Live Demo</b> — Watch Rubrik Agent Cloud govern and block unintended Claude actions in real-time</p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://go.rbrk.co/Secure-Claude-Code-wbnr?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14"><span class="button__text" style=""> Reserve your seat </span></a></div><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>🛡️ This event is </i><a class="link" href="https://go.rbrk.co/Secure-Claude-Code-wbnr?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=is-your-security-team-ready-for-ai-coding-agents-join-us-on-july-14" target="_blank" rel="noopener noreferrer nofollow"><i>presented</i></a><i> by the Rubrik team. We thank Rubrik for sharing their expertise and supporting Turing Post’s mission to bring clarity to the AI landscape.</i></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>AI Concepts and Techniques in 2026: Memory, Inference, Fine-Tuning &amp; Tokens</title>
  <description>DeepSeek mHC, Conditional Memory, fine-tuning, and inference chips explained — plus a guide to the full LLM workflow from tokens to answers.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be219585-46f9-4596-8e99-eeb9b888a62e/agents.jpg" length="109816" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens</guid>
  <pubDate>Thu, 02 Jul 2026 00:00:00 +0000</pubDate>
  <atom:published>2026-07-02T00:00:00Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Ai 101]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><b>TL;DR: </b>AI agents in 2026 are becoming durable systðms with memory, tools, skills, local control, physical action, and self-improvement loops. This recap maps the shift from OpenClaw and Hermes to VLA models, Web World Models, RSI, and Responsible AI infrastructure.</p></div><p class="paragraph" style="text-align:justify;">AI progress in 2026 is now coming from many different directions. Some advances rethink model structure, like DeepSeek mHC and depth-addressable Transformers. Others focus on memory, fine-tuning, self-distillation, inference hardware, or the runtime pipelines. <b>But the main theme running through everything is that AI is becoming more selective, modular, and infrastructure-aware.</b></p><p class="paragraph" style="text-align:justify;">This recap connects two layers of the same story. <b>The first layer is new research ideas: </b>conditional memory, post-RL fine-tuning, on-policy distillation, inference chips, and deeper ways to reuse Transformer representations. <b>The second layer is the practical AI workflow:</b> tokens, token types, embeddings, vector databases, attention, KV cache, and inference orchestration.</p><p class="paragraph" style="text-align:justify;">If you want to understand where AI systems are going, you need both: the new techniques at the frontier and the basic mechanisms that make them usable.</p><div class="section" style="background-color:transparent;border-color:#8c0de8;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;">🎉<b> Turing Post is turning 3!</b> To celebrate three years of deep-tech journalism, we are offering <b>30% off</b> our premium subscription. This week only! Upgrade today to unlock the rest of this recap and gain full access to our deep dives into agentic infrastructure. Join thousands of people who want to understand all sides of modern AI.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/upgrade?offer_id=f8b8d1ee-2b98-4f75-9e42-747d82d52bab&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens"><span class="button__text" style=""> Subscribe for ONLY $49/year </span></a></div></div><h2 class="heading" style="text-align:justify;" id="1-deep-seek-m-hc-breaking-the-archi">1. <a class="link" href="https://www.turingpost.com/p/mhc?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">DeepSeek mHC: Breaking the Architectural Limits of Deep Learning</a></h2><p class="paragraph" style="text-align:justify;">DeepSeek’s mHC, or Manifold-Constrained Hyper-Connections, shows a new way of thinking about AI architecture. For years, deep networks relied on residual connections to keep signals from vanishing, but that stability also limited how much information could transform across layers. Hyper-connections made routing more flexible, but less stable. mHC adds geometric constraints – doubly stochastic matrices and Sinkhorn-Knopp normalization –  to mix information without exploding or disappearing. In 2026, this matters because the next gains may come from designing newer  architectures where stability and expressivity can coexist.</p><div class="image"><img alt="DeepSeek mHC paper" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c7cc0e77-94f9-4479-8868-890e126aa872/image.png?t=1782438179"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Image Credit: mHC original paper</p></span></div></div><h2 class="heading" style="text-align:justify;" id="3-conditional-memory-and-the-rise-o">2. <a class="link" href="https://www.turingpost.com/p/conditionalmemory?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">Conditional Memory and the Rise of Selective Intelligence</a></h2><p class="paragraph" style="text-align:justify;">Conditional Memory, introduced through DeepSeek’s Engram architecture, explores a new idea: models shouldn’t have to activate all of their knowledge for every task.Models typically store everything in parameters or ever-growing context windows, but Engram lets models selectively retrieve memory through sparse lookups. One of the most interesting findings is the “U-shaped allocation law,” showing that the best systems balance memory capacity and computation rather than maximizing either one.This new memory type really matters because AI is moving toward selective intelligence with architectures that decide what to remember, what to retrieve, and where to spend compute. </p><div class="image"><img alt="Conditional Memory and the rise of selective intelligence" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/739a9394-bb53-4e3c-9892-fc4ba0156425/image.png?t=1782438302"/><div class="image__source"><span class="image__source_text"><p>A conceptual routing schematic diagram, Turing Post</p></span></div></div><h2 class="heading" style="text-align:justify;" id="3-beyond-rl-the-new-fine-tuning-sta">3. <a class="link" href="https://www.turingpost.com/p/beyondrl?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">Beyond RL: The New Fine-Tuning Stack for LLMs</a></h2><p class="paragraph" style="text-align:justify;"> In 2025, everyone put the RL loop front and center of post-training. And it still matters, but it remains expensive, brittle, and often too noisy. This episode maps the newer fine-tuning stack:</p><ul><li><p class="paragraph" style="text-align:justify;">Generated adapters like Doc-to-LoRA and Text-to-LoRA</p></li><li><p class="paragraph" style="text-align:justify;">Compressed and structured LoRA variants like LoRA-Squeeze and Kron-LoRA, Mixture of Adapters</p></li><li><p class="paragraph" style="text-align:justify;">Gradient-free Evolution Strategies</p></li></ul><p class="paragraph" style="text-align:justify;">And you can mix any of this LoRA with Evolution Strategies. As model capabilities are becoming modular and can be optimized separately, fine-tuning of 2026 starts to look more like designing an adaptive ecosystem around model.</p><div class="image"><img alt="Perceiver-style hypernetwork" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a00ac3ec-ed3f-4110-a70e-fc09776e1886/image.png?t=1782438449"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Turing Post</p></span></div></div><h2 class="heading" style="text-align:justify;" id="4-on-policy-distillation-zeitgeist">4. <a class="link" href="https://www.turingpost.com/p/opdzeitgeist?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">&quot;On-Policy Distillation Zeitgeist&quot;</a></h2><p class="paragraph" style="text-align:justify;">On-policy self-distillation becoming one of the more practical post-training directions of 2026. The core idea is that strong models can learn from their own improved versions: by comparing an uninformed answer with a version that has access to a solution, a demo, or rich feedback. The article walks through OPSD, SDFT, and SDPO to show how self-distillation can refine reasoning, support continual learning, and use feedback more efficiently than standard RL loops. Self-education is one of the central focuses where models are starting to improve by analyzing what they themselves got wrong. </p><div class="image"><img alt="On-policy self-distillation for large language models" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bf93e9dd-783d-4b58-b49a-1186504f77bd/image.png?t=1782438526"/><div class="image__source"><span class="image__source_text"><p>Image Credit: “On-Policy Self-Distillation for Large Language Models” paper</p></span></div></div><h2 class="heading" style="text-align:justify;" id="5-the-inference-chip-wars-mat-x-taa">5. <a class="link" href="https://www.turingpost.com/p/taalas?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">The Inference Chip Wars – MatX, Taalas, and the Cracks in the GPU Era</a></h2><p class="paragraph" style="text-align:justify;">The inference chip wars became one of the most important infrastructure stories of the first half of 2026. AI deployment has moved from training to serving billions of tokens and the competition is about cost per token, latency, power efficiency, and context handling.  Inference is fragmenting by workload, creating room for specialized hardware beyond GPUs and reshaping the economics of AI at scale. </p><p class="paragraph" style="text-align:justify;">We found these three visions of the future especially interesting:</p><ul><li><p class="paragraph" style="text-align:justify;">NVIDIA’s rack-scale Vera Rubin platform</p></li><li><p class="paragraph" style="text-align:justify;">MatX’s programmable LLM-first accelerator</p></li><li><p class="paragraph" style="text-align:justify;">Taalas’ radical “model-as-hardware” approach, where a specific model is effectively baked into silicon.</p></li></ul><p class="paragraph" style="text-align:justify;">We unpacked their exact workflows in this <a class="link" href="https://www.turingpost.com/p/taalas?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">article</a>.</p><div class="image"><img alt="Talaas, GPU, MatX" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0cbd9639-82b2-4663-811e-5c92080a1e1e/image.png?t=1782438579"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Turing Post</p></span></div></div><h2 class="heading" style="text-align:justify;" id="8-transformers-depth-is-an-addressa">6. <a class="link" href="https://www.turingpost.com/p/transformersdepth?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">Transformers Depth Is an Addressable Dimension</a></h2><p class="paragraph" style="text-align:justify;">Traditionally transformer depth was something that tokens needed to pass through. But this year appeared an idea that you can search through them. This article follows two new approaches – Kimi’s Attention Residuals and ByteDance Seed’s Mixture-of-Depths Attention (MoDA) – that tackle the same issue: useful early-layer signals get diluted as deep Transformers keep adding residual updates. Attention Residuals makes the residual stream choose which earlier layers matter. MoDA lets attention heads retrieve keys and values from previous layers. As a result, depth turns into an addressable memory dimension, giving Transformers a way to reuse intermediate representations instead of letting them wash away.</p><div class="image"><img alt="Kimi’s Attention Residuals and ByteDance Seed’s Mixture-of-Depths Attention (MoDA)" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/721dc528-1d76-48bc-bff3-2705c0d4b387/image.png?t=1782438650"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Attention Residuals original paper</p></span></div></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="special-collection">Special collection of guides on LLM workflow</h2><h3 class="heading" style="text-align:justify;" id="1-what-is-a-token-and-why-it-runs-a">1. <a class="link" href="https://www.turingpost.com/p/token?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">What Is a Token (and why it runs AI)?</a></h3><p class="paragraph" style="text-align:justify;">Tokens may sound just like the smallest building block in AI, but they shape almost everything that happens inside a model. For example, text becomes tokens, context windows are measured in tokens, and every prompt, response, and API bill depends on them. Less obvious, though, is that <b>tokens have become AI’s economic unit</b>: they determine cost, latency, and infrastructure efficiency. Understanding tokenization is the key to deploying and using AI effectively as well as figuring out how autoregressive LLMs are built inside.</p><div class="image"><img alt="Bit-level BPE " class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4de5b7b3-86d8-4123-82cc-fc9ae3c7e3ec/image.png?t=1782435579"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Bit-level BPE original paper</p></span></div></div><h3 class="heading" style="text-align:justify;" id="2-llm-token-types-input-output-reas">2. <a class="link" href="https://www.turingpost.com/p/tokentaxonomy?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">LLM Token Types: Input, Output, Reasoning, Cached & More</a></h3><p class="paragraph" style="text-align:justify;">When you know what are tokens, you need to know what tokens exist. Modern AI systems work with a whole “token zoo”: input and output tokens, reasoning, cached, speculative tokens, as well as retrieval, tool-use and multimodal tokens. Each consumes different amounts of compute and affects cost in different ways. In this article, we explain why agentic AI is changing token economics, with hidden overhead from tool calls, retrieval loops, and reasoning often dwarfing the visible prompt and response. In 2026, understanding token types is becoming just as important as understanding models, because AI performance is increasingly an optimization problem as much as a modeling one.</p><div class="image"><img alt="input and output tokens, reasoning, cached, speculative tokens" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5491ba22-9f97-4758-9a9d-1d0b132992d2/image.png?t=1782437823"/><div class="image__source"><span class="image__source_text"><p>Image Credit: The Turing Post</p></span></div></div><h3 class="heading" style="text-align:justify;" id="3-whats-so-magical-about-embeddings">3. <a class="link" href="https://www.turingpost.com/p/embeddings?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">What’s So Magical About Embeddings?</a></h3><p class="paragraph" style="text-align:justify;">The next stage after tokenization is the creation of embeddings. Tokens become vectors in a shared geometric space where distance reflects meaning. I the article, we follow this process and unpack why Rotary Position Embeddings (RoPE) – which encode position through rotation rather than fixed embeddings – became the standard for handling long context. In 2026, embeddings sit at the center of modern AI, powering search, retrieval, memory, multimodal models, and the contextual understanding that makes today&#39;s systems work.</p><div class="image"><img alt="What are Vector embeddings" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d23a8405-8e0a-4c96-a9ba-1840005399c1/image.png?t=1782434783"/><div class="image__source"><span class="image__source_text"><p>Image Credit: What are Vector Embeddings? by Qdrant</p></span></div></div><h3 class="heading" style="text-align:justify;" id="4-agentic-vector-databases-what-is-">4.<a class="link" href="https://www.turingpost.com/p/agentic-vector-databases?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow"> Agentic Vector Databases – What Is That?</a></h3><p class="paragraph" style="text-align:justify;">This year, even core concepts like vector databases have evolved. They are becoming much more than passive retrieval stores in the agent era. Today we have agentic search, memory, and knowledge-engine layers instead of precious RAG. It’s a whole new infrastructure that lets models and agents emember, navigate, and act on knowledge over time. There are three interesting approaches:</p><ul><li><p class="paragraph" style="text-align:justify;">Chroma’s Context-1 separates search from final reasoning</p></li><li><p class="paragraph" style="text-align:justify;">Weaviate’s Engram turns interactions into managed memory</p></li><li><p class="paragraph" style="text-align:justify;">and Pinecone’s Nexus compiles raw data into task-specific artifacts for agents before they ask.</p></li></ul><div class="image"><img alt="Agentic Vector database scheme" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b68e8cd3-0c25-4817-b7ab-0eef10a8338a/image.png?t=1782434908"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Pinecone Nexus blog post</p></span></div></div><h3 class="heading" style="text-align:justify;" id="4-agentic-vector-databases-what-is-">5.<a class="link" href="https://www.turingpost.com/p/agentic-vector-databases?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow"> </a><a class="link" href="https://www.turingpost.com/p/your-ultimate-guide-to-attention-mechanism-qkv-and-kv-cache?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">Your Ultimate Guide to Attention: Mechanism, QKV, and KV Cache</a></h3><p class="paragraph" style="text-align:justify;">Everyone needs this guide to learn or refresh how autoregressive models (the one which generate one token at a time) work. <b>Attention remains the core computation behind modern Transformers</b>. This article breaks down the full attention pipeline –<a class="link" href="https://www.turingpost.com/p/agentic-vector-databases?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow"> </a>from queries, keys, and values (QKV) to self-attention, multi-head attention, and KV cache – and explains why these mechanisms let models build rich context instead of treating tokens independently. It also covers newer ideas and other KV-efficient variants that make long-context reasoning practical. As context windows keep growing and AI agents process more information, attention optimization is becoming just as important as scaling models themselves.</p><div class="image"><img alt="Transformer architecture" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0f216e6e-0092-4342-8687-c977f73e0079/image.png?t=1782437540"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Transformer architecture showing self-attention layers and positional encodings, Attention is All You Need</p></span></div></div><h3 class="heading" style="text-align:justify;" id="6-from-tokens-to-answers-what-actua">6.<a class="link" href="https://www.turingpost.com/p/llm-inference-from-tokens-to-answers?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow"> </a><a class="link" href="https://www.turingpost.com/p/llm-inference-from-tokens-to-answers?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-concepts-and-techniques-in-2026-memory-inference-fine-tuning-tokens" target="_blank" rel="noopener noreferrer nofollow">From Tokens to Answers: What Actually Happens During LLM Inference</a></h3><p class="paragraph" style="text-align:justify;">Finally, we are gathering all these parts together to answer ta question: What happens between your prompt and the model’s response? We walk you through the entire runtime pipeline: tokenization → embeddings → attention and KV cache, plus retrieval, batching, memory management, and orchestration. Today inference really highly depends on orchestration as well as on the model itz As reasoning models and agents grow more complex, orchestration around the model is becoming just as important as the model itself.</p><div class="image"><img alt="What actually happens during LLM Inference" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1e1e266d-1b66-41bc-8354-dc8231e2b58e/image.png?t=1782437735"/><div class="image__source"><span class="image__source_text"><p>Image Credit: The Turing Post</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Tame Your AI Monsters: Claude Edition 🛡️</title>
  <description>Unleash agents, not risk</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f7c05e92-c24c-47dc-896e-c790d4ed7481/Template_-_PARTNERS.jpg" length="100308" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/tame-your-ai-monsters-claude-edition</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/tame-your-ai-monsters-claude-edition</guid>
  <pubDate>Thu, 25 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-25T21:00:00Z</atom:published>
    <category><![CDATA[From Our Partners]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;"><b>Announcing a super interesting workshop from our partners →</b></p><div class="image"><a class="image__link" href="https://go.rbrk.co/Secure-Claude-Code?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=tame-your-ai-monsters-claude-edition" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/01d6c82a-115f-447e-914f-cce153e2bd48/Paid-Social01-NAM-Tame-Your-AI-Monster_1200X628.gif?t=1782334903"/></a></div><p class="paragraph" style="text-align:justify;"><b>Claude is running in your enterprise.</b></p><p class="paragraph" style="text-align:justify;">It’s scheduling, drafting, analyzing, and making calls – and in most organizations, nobody has a clear answer to a very simple question: <b>what exactly is it doing?</b></p><p class="paragraph" style="text-align:justify;">That’s not a Claude problem. <b>That’s a governance problem. </b>And it’s exactly what turns a well-intentioned AI agent into something that accesses what it shouldn’t, skips the steps it finds inconvenient, and causes damage faster than anyone can react.</p><p class="paragraph" style="text-align:justify;">We call that Chungar.</p><p class="paragraph" style="text-align:justify;"><b>Rubrik is Customer Zero for this problem.</b></p><p class="paragraph" style="text-align:justify;">They are deploying Claude. Rubrik is the first organization to put Rubrik Agent Cloud to work governing a live Claude implementation – and on <a class="link" href="https://go.rbrk.co/Secure-Claude-Code?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=tame-your-ai-monsters-claude-edition" target="_blank" rel="noopener noreferrer nofollow">June 30, they are opening that experience up in a hands-on lab.</a></p><p class="paragraph" style="text-align:justify;">This isn’t a webinar. You will:</p><ul><li><p class="paragraph" style="text-align:justify;">Watch a Claude deployment go off-script live, </p></li><li><p class="paragraph" style="text-align:justify;">see Rubrik Agent Cloud catch it in real time, </p></li><li><p class="paragraph" style="text-align:justify;">and then fix it yourself in a proctored lab environment.</p></li></ul><p class="paragraph" style="text-align:justify;"><b>Seats are limited.</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://go.rbrk.co/Secure-Claude-Code?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=tame-your-ai-monsters-claude-edition"><span class="button__text" style=""> Request Your Workshop Spot – Episode 1 </span></a></div><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>🛡️ This event is </i><a class="link" href="https://go.rbrk.co/Secure-Claude-Code?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=tame-your-ai-monsters-claude-edition" target="_blank" rel="noopener noreferrer nofollow"><i>presented</i></a><i> by the Rubrik team. We thank Rubrik for sharing their expertise and supporting Turing Post’s mission to bring clarity to the AI landscape.</i></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>AI Agents in 2026: Local, Physical, Responsible AI</title>
  <description>AI agents in 2026 are becoming durable systems with memory, skills, and physical action. Covers OpenClaw, Hermes, VLA models, Web World Models, Responsible AI.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/77f66fa2-aae3-4cc0-af7f-7d3dad88f8d8/agents.jpg" length="105943" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/ai-agents-in-2026-local-physical-responsible-ai</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/ai-agents-in-2026-local-physical-responsible-ai</guid>
  <pubDate>Wed, 24 Jun 2026 23:14:43 +0000</pubDate>
  <atom:published>2026-06-24T23:14:43Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Ai 101]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><b>TL;DR: </b>AI agents in 2026 are becoming durable systðms with memory, tools, skills, local control, physical action, and self-improvement loops. This recap maps the shift from OpenClaw and Hermes to VLA models, Web World Models, RSI, and Responsible AI infrastructure.</p></div><p class="paragraph" style="text-align:justify;">AI agents became the center of the first half of 2026. We got persistent systems with memory, tools, skills, that can act across software and physical environments. This recap us a broad look at this shift through several angles – local agents such as OpenClaw and Hermes, model choices like Gemma 4, skill engineering, more advanced systems like VLA models for robotics, Web World Models, recursive self-improvement recent boom, and Responsible AI for systems that can actually do things.</p><p class="paragraph" style="text-align:justify;"><b>Infrastructure is becoming as important as agents and models themselves. </b>Systems of the new generation need features such as identity, memory, skills, environments, and many others. So let’s refresh what happened in the first half of 2026 – what can advanced systems safely and reliably do?</p><div class="section" style="background-color:transparent;border-color:#8c0de8;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;">🎉<b> Turing Post is turning 3!</b> To celebrate three years of deep-tech journalism, we are offering <b>30% off</b> our premium subscription. Upgrade today to unlock the rest of this recap and gain full access to our deep dives into agentic infrastructure. Be like Eric Schmidt and Marc Cuban 😉 </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/upgrade?offer_id=f8b8d1ee-2b98-4f75-9e42-747d82d52bab&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai"><span class="button__text" style=""> Claim $49 Premium access </span></a></div></div><h2 class="heading" style="text-align:justify;" id="1-open-claw-explained-lightweight-a">1. <a class="link" href="https://www.turingpost.com/p/openclaw?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">OpenClaw Explained + lightweight alternatives</a></h2><p class="paragraph" style="text-align:justify;">There&#39;s really only one place to start this recap, and it’s <b>OpenClaw</b>. OpenClaw captures the local agent boom better than almost anything else in early 2026. It turns a personal AI assistant into file-backed infrastructure: </p><ul><li><p class="paragraph" style="text-align:justify;">identity in <a class="link" href="http://soul.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">SOUL.md</a></p></li><li><p class="paragraph" style="text-align:justify;">scheduled reasoning in <a class="link" href="http://heartbeat.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">HEARTBEAT.md</a></p></li><li><p class="paragraph" style="text-align:justify;">memory in Markdown</p></li><li><p class="paragraph" style="text-align:justify;">tool execution coordinated through a central Gateway. </p></li></ul><p class="paragraph" style="text-align:justify;">Many people are no longer satisfied with simple chatbots – now they need a personal control plane that integrates with messengers like WhatsApp, Telegram, Discord, and Slack. With this new local assistants philosophy, agents started to be durable systems with memory, advanced tool use, and identity. And the best thing, of course – this infrastructure layer is taking shape through open source.</p><div class="image"><img alt="openclaw map" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bddc9f15-e860-4c95-a49c-a45d56c39e53/image.png?t=1782318154"/><div class="image__source"><span class="image__source_text"><p>Image Credit: OpenClaw DeepWiki</p></span></div></div><div class="section" style="background-color:transparent;border-color:#222222;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><b>A quick programming note:</b> </p><p class="paragraph" style="text-align:justify;">I’m going to be taking a brief vacation over the next two weeks to recharge, which means our regular weekldy FODs will be on a short break. But don&#39;t worry – we’ll be running our comprehensive half-year recaps during this time. It’s a great, relaxed opportunity for everyone to catch up on our deep dives and infrastructure breakdowns before we kick off our fourth year. <a class="link" href="https://www.turingpost.com/upgrade?offer_id=f8b8d1ee-2b98-4f75-9e42-747d82d52bab&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">Get the full access now!</a></p></div><h2 class="heading" style="text-align:justify;" id="2-hermes-agent-vs-open-claw-local-a">2. <a class="link" href="https://www.turingpost.com/p/hermes?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">Hermes Agent vs OpenClaw: Local AI Agents Compared</a></h2><p class="paragraph" style="text-align:justify;">The battle for local AI agents became one of the most interesting software stories of 2026. After OpenClaw, Nous Research introduced Hermes Agent – another local agent that develops skills. Our article compares OpenClaw and Hermes Agent, because they are built around very different philosophies:</p><ul><li><p class="paragraph" style="text-align:justify;">OpenClaw focuses on user control, file-backed identity, and human-authored skills.</p></li><li><p class="paragraph" style="text-align:justify;">Hermes strengths is self-improvement, procedural layered memory, and skills generated from experience. the more you use the agent, the more useful it becomes. the longer they run. </p></li></ul><p class="paragraph" style="text-align:justify;">In general, local agents are starting to remember methods as well as facts. But the two agents confront in one particular point: should personal agents be controlled or allowed to learn?</p><div class="image"><img alt="Hermes vs Openclaw" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/13a735ca-f8bd-4f24-886a-6e2cc427e013/image.png?t=1782317525"/></div><h2 class="heading" style="text-align:justify;" id="1-open-claw-explained-lightweight-a">3. <a class="link" href="https://www.turingpost.com/p/gemma4?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">Gemma 4 with OpenClaw: Architecture, Setup, and Why Developers Are Switching</a></h2><p class="paragraph" style="text-align:justify;">Gemma 4 feels built for the local-agent moment. Google DeepMind optimized it for intelligence per parameter: smaller active compute, efficient attention, multimodality, structured outputs, function calling, and models that can actually run on devices. The article connects this directly to OpenClaw, where users want strong local agents without Claude-level API bills. The interesting tension is practical: Gemma 4 looks like the new default model to try first, but harder agentic workflows may still need tuning or fallback models. In 2026, local AI becomes less theoretical and much more usable.</p><div class="image"><img alt="All you need to know about Gemma 4" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/874ad886-0b01-4dd3-b9dd-eee537325896/image.png?t=1782318091"/><div class="image__source"><span class="image__source_text"><p>Image Credit: A Visual Guide to Gemma 4 by Maarten Grootendorst</p></span></div></div><h2 class="heading" style="text-align:justify;" id="4-what-is-skill-engineering-from-pr">4. <a class="link" href="https://www.turingpost.com/p/from-prompt-engineering-to-skill-engineering?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">What Is Skill Engineering? From Prompt Optimization to Skill Optimization</a></h2><p class="paragraph" style="text-align:justify;">Since personal agents like OpenClaw and Hermes Agent are the frontier of 2026, skill engineering is becoming a very important layer of agent optimization after prompt and context engineering. These agents rely more on reusable skills, so the main question is: “how do we train, maintain, and optimize the skills themselves?” There are three fresh methods: <b>SkillOpt </b>for improving one skill through validated edits, <b>SkillOps</b> for cleaning and maintaining whole skill libraries, and <b>SkillMOO </b>for finding cost-effective skill bundles for coding agents. And further there will be more of them.</p><div class="image"><img alt="From prompt engineering to skill engineering" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bd885a3b-91f1-41a3-b02a-e1731eef58f0/image.png?t=1782318057"/><div class="image__source"><span class="image__source_text"><p>Image Credit: SkillOpt original paper</p></span></div></div><h2 class="heading" style="text-align:justify;" id="5-vla-models-explained-architecture">5. <a class="link" href="https://www.turingpost.com/p/vlaplus?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">VLA Models Explained: Architecture, Types & the Leap to VLA+ </a></h2><p class="paragraph" style="text-align:justify;">Vision-Language-Action (VLA) models are becoming the core interface for Physical AI: they connect what robots see, what humans ask, and what the robot actually does. This piece maps the landscape from Gemini Robotics and π0 to SmolVLA, Helix, ChatVLA-2, ACoT-VLA, VLA-0, and Microsoft’s new Rho-alpha. Rho-alpha is special here, because it illustrates the shift to VLA+:  models that add touch, online learning, and real-time human correction. VLAs are very important now because they influence robotics progress, moving them beyond fixed programs toward systems that can perceive, reason, adapt, and keep improving after deployment. </p><div class="image"><img alt="a VLA model for generalist humanoid control" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/931d0db9-37e8-4be6-a28d-b4664f71dca3/image.png?t=1782319039"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Helix: A Vision-Language-Action Model for Generalist Humanoid Control blog post</p></span></div></div><h2 class="heading" style="text-align:justify;" id="6-nemotron-3-and-the-surprising-coa">6. <a class="link" href="https://www.turingpost.com/p/nemotroncoalition?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">Nemotron 3 and the Surprising Coalition Building New AI in the Open</a></h2><p class="paragraph" style="text-align:justify;">Nemotron 3 is NVIDIA’s sensational development to build an open AI ecosystem around shared infrastructure. From the technology side it offers: a hybrid Transformer–Mamba architecture, LatentMoE routing, multi-token prediction, and NVFP4 training. But even more interesting is who build these parts. NVIDIA gathered a coalition including Mistral, Cursor, Perplexity, LangChain, Black Forest Labs, and others, contributing models, data, evaluations, tooling, and domain expertise Frontier AI development can be even more modular than we thought. Maybe AI ecosystems and collaboration of developers will drive the progress?</p><div class="image"><img alt="Nemotron 3 scheme" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b0b83146-0d34-47c1-ad0f-4f1f137cc814/image.png?t=1782318997"/><div class="image__source"><span class="image__source_text"><p>Image Credit: NVIDIA</p></span></div></div><h2 class="heading" style="text-align:justify;" id="5-nemotron-3-and-the-surprising-coa">7. <a class="link" href="https://www.turingpost.com/p/wwm?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">What are Web World Models?</a> </h2><p class="paragraph" style="text-align:justify;">Web World Models (WWM) offer a practical recipe for building persistent worlds for AI agents using standard web infrastructure. WWMs split the system in two parts: deterministic code handles rules, state, and “physics,” while language models add descriptions, narratives, and high-level content. Moreover, worlds can stay consistent without storing everything, using typed interfaces, hashing, and graceful fallbacks when the model is slow or unavailable. As agents need stable environments to act, remember, fail, and learn WWMs are one of the options.</p><div class="image"><img alt="Web World Models paper" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6f476150-d171-44dd-afda-d5e64b2f450d/image.png?t=1782319013"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Web World Models original paper</p></span></div></div><h2 class="heading" style="text-align:justify;" id="4-from-prompt-engineering-to-skill-">8. <a class="link" href="https://www.turingpost.com/p/what-is-recursive-self-improvement?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">What is Recursive Self-Improvement?</a></h2><p class="paragraph" style="text-align:justify;">A recent hot topic is Recursive Self-Improvement that moved from science-fiction thought experiment to a real research direction in 2026. Simply, it’s the process when AI builds better AI itself. The article follows three complementary paths:</p><ul><li><p class="paragraph" style="text-align:justify;">Sakana AI’s long-term vision of AI-driven research loops</p></li><li><p class="paragraph" style="text-align:justify;">Anthropic’s use of Claude to automate coding and engineering work</p></li><li><p class="paragraph" style="text-align:justify;">Recursive’s system for running, evaluating, and combining AI-generated experiments.</p></li></ul><p class="paragraph" style="text-align:justify;">And the results are very, very interesting. Claude now writes more than 80% of Anthropic’s merged code, while Recursive discovered enough small optimizations to cut training times and improve model quality automatically.</p><div class="image"><img alt="recursive self improvement explained" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/945ca3d5-baa0-49e8-bda3-cd08536db0de/image.png?t=1782318858"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Turing Post</p></span></div></div><h2 class="heading" style="text-align:justify;" id="9-how-responsible-ai-changes-in-the">9. <a class="link" href="https://www.turingpost.com/p/how-responsible-ai-changes-in-the-agent-era?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-agents-in-2026-local-physical-responsible-ai" target="_blank" rel="noopener noreferrer nofollow">How Responsible AI Changes In The Agent Era</a></h2><p class="paragraph" style="text-align:justify;">And finally, Responsible AI became much more tangible in the agent era. When AI systems move from generating text to taking actions through tools, APIs, and workflows, safety can no longer rely on output review alone. In this article, we discuss how Microsoft and Google DeepMind are turning Responsible AI into idnfrastructure with runtime controls, policy-as-tests, human oversight with monitoring, etc. Now it is even more obvious that agents need guardrails around what they can access and do. <b>Trust is becoming an engineering problem.</b></p><div class="image"><img alt="Who is responsible for AI" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e52b920a-a433-47c0-b895-47990d01f708/image.png?t=1782318882"/><div class="image__source"><span class="image__source_text"><p>Image Credit: Turing Post</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>FOD#157: What People Still Don’t Understand About AI Agents</title>
  <description>The BioNeMo Confusion</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5e9bea24-32c1-43cd-beb4-4ecbf07a6f29/Template_-_FOD__1_.jpg" length="91814" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/ai-agent-toolkits-what-people-still-don-t-understand</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/ai-agent-toolkits-what-people-still-don-t-understand</guid>
  <pubDate>Wed, 24 Jun 2026 00:30:00 +0000</pubDate>
  <atom:published>2026-06-24T00:30:00Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[&quot;Froth On The Daydream&quot;]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p id="this-week-in-turing-post" class="paragraph" style="text-align:justify;"><b>Today’s editorial:</b> what people still misunderstand about tools for agents, why NVIDIA BioNeMo is a toolkit rather than “extra model knowledge,” and why agentic science depends on tools like <a class="link" href="https://www.turingpost.com/p/mcp?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">MCP</a>, workflows, permissions, and human review.</p></div><hr class="content_break"><h3 class="heading" style="text-align:justify;" id="how-to-cut-the-trust-tax-of-evaluat">💸 How to Cut the Trust Tax of Evaluating AI Agents at Scale</h3><div class="image"><a class="image__link" href="https://www.fiddler.ai/guides/tco-for-operationalizing-agents?utm_source=turingpost&utm_medium=online_advertising&utm_campaign=tco-guide" rel="noopener" target="_blank"><img alt="free guide to reduce your AI TCO" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0e60b2ce-9c51-440a-b427-6e9b95699e4f/trust_tax_fiddler__1_.jpg?t=1782161955"/></a><div class="image__source"><a class="image__source_link" href="https://www.fiddler.ai/guides/tco-for-operationalizing-agents?utm_source=turingpost&utm_medium=online_advertising&utm_campaign=tco-guide" rel="noopener" target="_blank"><span class="image__source_text"><p>Free guide from our partners</p></span></a></div></div><p class="paragraph" style="text-align:justify;"><span style="background-color:transparent;">Evaluating agents with external LLMs looks affordable. Until your agent traffic grows. </span><span style="background-color:transparent;"><a class="link" href="https://www.fiddler.ai/guides/tco-for-operationalizing-agents?utm_source=turingpost&utm_medium=online_advertising&utm_campaign=tco-guide" target="_blank" rel="noopener noreferrer nofollow">Fiddler’s guide</a></span><span style="background-color:transparent;"> breaks down </span><span style="background-color:transparent;"><b>how to reduce Total Cost of Ownership</b></span><span style="background-color:transparent;"> while eliminating risk gaps in production. </span><br><br><span style="background-color:transparent;"><b>Learn how to:</b></span></p><ul><li><p class="paragraph" style="text-align:justify;"><span style="background-color:transparent;">Evaluate every trace without sampling</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:transparent;">Evaluate agents in-environment with batteries-included Trust Models</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:transparent;">Reduce API costs at scale</span></p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.fiddler.ai/guides/tco-for-operationalizing-agents?utm_source=turingpost&utm_medium=online_advertising&utm_campaign=tco-guide"><span class="button__text" style=""><span style="background-color:transparent;">Download the Agentic TCO Guide now</span></span></a></div><hr class="content_break"><p class="paragraph" style="text-align:center;"><i>Share Turing Post with one person. You will help us grow </i></p><hr class="content_break"><h3 class="heading" style="text-align:justify;" id="what-people-still-dont-understand-a">What People Still Don’t Understand About AI Agents and Tools</h3><p id="during-a-qa-with-kimberly-powell-at" class="paragraph" style="text-align:justify;">During a Q&A with Kimberly Powell at the BIO AI Summit, where NVIDIA announced that they had <a class="link" href="https://github.com/NVIDIA-BioNeMo?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">open-sourced their BioNeMo Agent Toolkit</a>, I heard a couple of questions from journalists who write about AI that made me see red. How much do you need to misunderstand the whole thing to ask something like that?</p><p class="paragraph" style="text-align:justify;">Then I calmed myself down and remembered: there are no dumb questions. There are signals.</p><p class="paragraph" style="text-align:justify;">And this one was a very useful signal. It showed how many people, including people who write about AI for a living, still don’t understand what agents are, what tools are, and what happens when you connect the two.</p><p class="paragraph" style="text-align:justify;">So let’s bring clarity to the world. Small mission, no pressure. </p><p class="paragraph" style="text-align:justify;">The question was basically this: <i>if Claude refuses to help someone create a bioweapon, doesn’t a scientific toolkit now give it a deep knowledge that can help with that?</i></p><p class="paragraph" style="text-align:justify;">Bioweapons are a legitimate safety topic. Nobody should wave that away. But the question revealed a layer mistake. It assumed that NVIDIA had given models new dangerous scientific knowledge. </p><p class="paragraph" style="text-align:justify;">And this is a fundamental misunderstanding of what a toolkit is. </p><p class="paragraph" style="text-align:justify;">What is BioNeMo? BioNeMo Agent Toolkit is NVIDIA’s collection of scientific models, tools, and workflows that AI agents can call for life-sciences tasks. Clear enough – except it’s not.</p><p class="paragraph" style="text-align:justify;">BioNeMo is much closer to giving a scientist access to laboratory equipment than teaching them biology. Or even simpler:<b> it is like giving someone a bicycle repair kit</b>. A screwdriver screws and unscrews screws (try to say it out loud). A wrench tightens bolts. A pump puts air into the tire. A patch covers a puncture. Can you make a bomb from this bicycle with a screwdriver or a patch? That’s not impossible! But it requires much more than that. </p><p class="paragraph" style="text-align:justify;">A toolkit gives you specific tools for specific jobs.</p><p class="paragraph" style="text-align:justify;">That is the simplest way to understand what NVIDIA announced. BioNeMo Agent Toolkit packages life-sciences tools and models into agent-callable skills: protein folding, molecular docking, generative chemistry, genomics analysis, protein design, biomarker discovery, and related workflows.</p><p class="paragraph" style="text-align:justify;">And that means exactly that: AI agents can now call scientific instruments. </p><p class="paragraph" style="text-align:justify;">There was another question that kinda surprised me: how it was possible that OpenAI and Anthropic/Claude – both presented as models that the toolkit can use – agreed to collaborate.</p><p class="paragraph" style="text-align:justify;">Well, they don’t, they don’t need to, and that’s quite obvious when you know – again – what a toolkit is.</p><h3 class="heading" style="text-align:justify;" id="what-makes-this-an-agent-toolkit-no">What makes this an agent toolkit, not just a pile of models?</h3><p class="paragraph" style="text-align:justify;">Because the agent can chain the tools and models. </p><p class="paragraph" style="text-align:justify;">Example: <b>design a protein binder</b></p><p class="paragraph" style="text-align:justify;">A scientist says: “<i>Find a possible binder for this target protein.</i>”</p><p class="paragraph" style="text-align:justify;">The agent does not magically “become a biologist.” It follows a workflow:</p><ol start="1"><li><p class="paragraph" style="text-align:justify;">Use <b>RFdiffusion </b>model to design possible binder backbones.</p></li><li><p class="paragraph" style="text-align:justify;">Use <b>ProteinMPNN </b>model to propose amino-acid sequences for those backbones.</p></li><li><p class="paragraph" style="text-align:justify;">Use <b>Boltz-2</b> or <b>OpenFold3 </b>models to predict whether the binder and target actually fold together.</p></li><li><p class="paragraph" style="text-align:justify;">Rank candidates using confidence and interface metrics.</p></li><li><p class="paragraph" style="text-align:justify;">Return the best candidates to the scientist with caveats.</p></li></ol><p class="paragraph" style="text-align:justify;">The BioNeMo repo even includes a generative protein binder workflow that combines RFdiffusion, ProteinMPNN, and Boltz-2 / OpenFold3 for this kind of sequence: backbones → sequences → co-fold → filter. Like a set of skills an agent might require.</p><h3 class="heading" style="text-align:justify;" id="why-is-it-so-important-to-understan">Why is it so important to understand?</h3><p class="paragraph" style="text-align:justify;">As we plunge headfirst into this new agentic era, a basic understanding of it becomes crucially important. Because understanding this distinction changes how we view the future of AI – shifting the focus from what AI <i>knows</i> to what AI can <i>do</i>.</p><p class="paragraph" style="text-align:justify;">If we go back to biotech and BioNeMo, the difference will be this: A chatbot can explain molecular docking. If you casually ask about it. An agent with the right tool can run a docking workflow. It’s less about learning and more about building. You need to know some stuff by then. A chatbot can describe protein design. An agent with the right workflow can generate a protein backbone, propose an amino-acid sequence, predict whether it might fold, inspect the output, and bring the result back to a human scientist.</p><p class="paragraph" style="text-align:justify;">Scientist is the key word here. </p><p class="paragraph" style="text-align:justify;">And this is where many people still get lost. They keep looking at the model as if the model is the whole story. It’s not.</p><p class="paragraph" style="text-align:justify;">Which model is smarter? It’s a good question but not the most interesting. The better question now is: where does the model act?</p><p class="paragraph" style="text-align:justify;">Inside a codebase? Inside GitHub? Inside a lab? Inside a medical scanner? Inside a drug-discovery workflow? Inside a system that repeatedly measures, tests, adjusts, and improves? <b>Once AI starts acting, the valuable layer is the loop around the model.</b></p><p class="paragraph" style="text-align:justify;">And science is full of loops.</p><p class="paragraph" style="text-align:justify;">Drug discovery, protein design, genomics, biomarker discovery, clinical research, medical imaging, literature review, protocol generation: these are not fields where progress usually arrives as one clean eureka moment. Much of the work is iteration. <b>That is why having agents and tools for agents is so exciting. </b>They can help the scientist move through the loop faster. I’m on my own loop to keep repeating that.</p><p class="paragraph" style="text-align:justify;">There was a fascinating point in the BioNeMo discussion: early dreams of AI for science imagined that models could simply consume all scientific knowledge, connect all the dots, and discoveries would just pop out. But that is not really how it worked. The real progress came when systems entered the loop: look at the literature, propose an experiment, analyze the data, then use that result to propose the next experiment.</p><p class="paragraph" style="text-align:justify;">It’s happening right now, though the majority of people still don’t realize that. That’s one of my biggest revelations from BIO AI Summit: <b>we are at the beginning of agentic science.</b></p><p class="paragraph" style="text-align:justify;">And yes, it should make us excited.</p><p class="paragraph" style="text-align:justify;">Because a lot of scientists spend far too little time doing science. If agents can compress some of that engineering and operational work, they give scientists something precious back: more time in the creative scientific space. More time looking at data. More time asking the next question. More time noticing that something unexpected happened and following it. <b>What does it give to the rest of us? The ability to cure diseases much faster.</b> Feels like it’s even better than a new productivity tool.</p><p class="paragraph" style="text-align:justify;">This is why BioNeMo is not only a biotech story. It is part of the bigger shift I keep coming back to: the model is no longer the whole story. The loop around the model is becoming the interesting layer.</p><hr class="content_break"><p class="paragraph" style="text-align:justify;">I talked about exactly this in my solo segment for O’Reilly’s <b>This Week in AI</b> this week: <b>who owns the loop? </b>In coding, in cybersecurity, in medicine, in science, the same question keeps appearing. Where does the model act? Who controls the tools? Who owns the feedback? Who decides when the loop is safe enough to run? Check it out→</p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/sXBWbiyT4ns" width="100%"></iframe><p class="paragraph" style="text-align:justify;"><i>If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going. </i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Comment </span></a></div><div class="section" style="background-color:transparent;border-color:#df12a9;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i><b>Follow us on </b></i><i> </i>🎥<i><a class="link" href="https://www.youtube.com/@RealTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow"> YouTube</a></i><i> </i><i><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Twitter</a></i><i> </i><i><a class="link" href="https://huggingface.co/Kseniase?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow"> Hugging Face </a></i>🤗</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="we-are-reading-watching">We are reading / watching </h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://blogs.nvidia.com/blog/liquid-cooling-ai-factories/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines</a> by NVIDIA is a very interesting article about liquid cooling, how it can enable zero water consumption, and how waste heat from AI infrastructure could potentially be shared with residential buildings.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.recodechinaai.com/p/how-chinese-researchers-plan-to-build?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">How Chinese Researchers Plan to Build Self-Improving AI</a> by Recode China AI</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.turingpost.com/p/how-responsible-ai-changes-in-the-agent-era?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Responsible AI in the Age of AI Agents</a></p></li></ul><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="twitter-library">Twitter Library</h2><div class="image"><a class="image__link" href="https://www.turingpost.com/p/agent-rl-training-tools?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" rel="noopener" target="_blank"><img alt="Agent RL Training Frameworks: 10 Open-source Tools to Know " class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/69752722-8b0a-4eda-a278-cec627ddc8db/Screenshot_2026-06-23_at_7.24.10_PM.jpg?t=1782257099"/></a></div><h2 class="heading" style="text-align:justify;" id="news-from-the-usual-suspects">News from the usual suspects ™</h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.midjourney.com/medical/blogpost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Midjourney announced Midjourney Medical</a>, a full-body underwater ultrasound scanner concept that aims to make internal imaging faster, more repeatable, and more consumer-friendly. It’s also stunningly beautiful. </p></li><li><p class="paragraph" style="text-align:justify;">SpaceX’s news:</p><ul><li><p class="paragraph" style="text-align:justify;">They agreed to <a class="link" href="https://x.com/SpaceX/status/2066873915717136548?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">buy Anysphere, the company behind Cursor</a>, for $60B in stock, turning the coding-agent race into a fight over the developer workflow itself.</p></li><li><p class="paragraph" style="text-align:justify;">SpaceX also disclosed a <a class="link" href="https://www.cnbc.com/2026/06/22/spacex-ai-colossus-data-center-reflection.html?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">compute-capacity deal with Reflection AI</a> worth up to $6.3B, making its AI-infrastructure ambitions look much bigger than Cursor alone. In March, we <a class="link" href="https://www.turingpost.com/p/reflectionai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">interviewed Reflection’s co-founder.</a> Worth taking a look.</p></li></ul></li><li><p class="paragraph" style="text-align:justify;"><b>Anthropic’s news</b></p><ul><li><p class="paragraph" style="text-align:justify;">They launched <a class="link" href="https://claude.com/product/tag?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Claude Tag in Slack,</a> letting Claude be summoned in group conversations to read context, break down tasks, and follow workplace threads like an AI teammate. Sad to see Andrej Karpathy as Anthropic’s new promo platform. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.reuters.com/legal/litigation/g7-leaders-vow-closer-ties-ai-they-hash-out-trusted-partners-scheme-2026-06-17/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Micron signed a strategic infrastructure agreement with Anthropic</a> to supply memory and storage products, adding memory supply to the list of bottlenecks in frontier AI scaling</p></li></ul></li><li><p class="paragraph" style="text-align:justify;">OpenAI’s news:</p><ul><li><p class="paragraph" style="text-align:justify;">Their <a class="link" href="https://openai.com/index/introducing-life-sci-bench/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">LifeSciBench tests models</a> on real life-science work with 750 expert-authored tasks, 1,062 artifacts, 173 scientist contributors, and more than 19,000 rubric criteria.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://molecule.one/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">OpenAI and Molecule.one</a> demonstrated a near-autonomous AI chemist that improved Chan-Lam coupling, a drug-chemistry reaction for forming carbon-nitrogen bonds.</p></li></ul></li><li><p class="paragraph" style="text-align:justify;">NVIDIA’s news</p><ul><li><p class="paragraph" style="text-align:justify;">NVIDIA introduced BioNeMo that we discussed above.</p></li><li><p class="paragraph" style="text-align:justify;">They also introduced new <a class="link" href="https://blogs.nvidia.com/blog/ai-for-science-software-cuda/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">AI-for-science software at ISC</a>, including DAQIRI and ALCHEMI NIM microservices for scientific pipelines from astronomy to chemistry and materials simulation</p></li></ul></li></ul><h2 class="heading" style="text-align:justify;" id="survey-highlight">Survey highlight</h2><p class="paragraph" style="text-align:justify;"><b>World Action Models: A Survey </b>by National University of Singapore</p><div class="image"><img alt="Definition of a world action model WAM" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9d9bae5a-7756-48b5-a82c-71e95584af28/Screenshot_2026-06-23_at_10.39.30_AM.jpg?t=1782225628"/><div class="image__source"><span class="image__source_text"><p>Image Credit: The original paper</p></span></div></div><p class="paragraph" style="text-align:justify;">This <a class="link" href="https://arxiv.org/abs/2606.20781?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">survey conceptualizes World Action Models (WAMs)</a> as embodied predictive-action systems that integrate future forecasting directly into the action path. WAMs unify Vision-Language-Action policies with predictive world models. The authors structure the field into three primary design philosophies: Render-and-Decode, Latent-Only, and Video-Generation-Free. Finally, the taxonomy evaluates critical trade-offs balancing representational richness against compute, memory, latency, and physical plausibility. The field is moving toward generating less of the future while preserving control requirements.</p><h2 class="heading" style="text-align:justify;" id="models">Models</h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://Z.ai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>Z.ai</b></a><b> released GLM-5.2, a major Chinese open-weight coding model with a 1M-token context window and MIT license, putting open Chinese models back into the U.S. AI anxiety machine.</b></p></li><li><p class="paragraph" style="text-align:justify;"><b>Sakana AI released Fugu and Fugu Ultra, an orchestration-model family that routes tasks across a swappable pool of models behind one API.</b></p></li><li><p class="paragraph" style="text-align:justify;"><b>DreamX-World 1.0</b> – builds toward interactive general-purpose world models instead of passive video generators <a class="link" href="https://arxiv.org/abs/2606.16993?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Qwen-RobotWorld Technical Report</b> – connects embodied world modeling with language-conditioned video generation for robotics <a class="link" href="https://arxiv.org/abs/2606.17030?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>PAIWorld</b> – proposes a 3D-consistent world foundation model for robotic manipulation <a class="link" href="https://arxiv.org/abs/2606.18375?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>BioMatrix</b> – expands biological foundation models across sequences, structures, and language <a class="link" href="https://arxiv.org/abs/2606.22138?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>VibeThinker-3B</b> – pushes verifiable reasoning into small language models, which matters because reasoning cannot only live in giant expensive systems <a class="link" href="https://arxiv.org/abs/2606.16140?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Sumi</b> – develops an open diffusion language model from scratch, useful as diffusion LMs become less fringe and more serious <a class="link" href="https://arxiv.org/abs/2606.19005?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="research">Research </h2><p class="paragraph" style="text-align:justify;">Trends we see looking at every paper related to AI and ML published last week:</p><ul><li><p class="paragraph" style="text-align:justify;">World models are becoming agent infrastructure, not video generators.</p></li><li><p class="paragraph" style="text-align:left;">Agents are learning to manage their own memory and improve themselves.</p></li><li><p class="paragraph" style="text-align:left;">Reasoning is shifting toward reflection, verification, and active perception.</p></li><li><p class="paragraph" style="text-align:left;">Robotics and foundation models are rapidly converging.</p></li><li><p class="paragraph" style="text-align:left;">Researchers are experimenting beyond standard transformer architectures.</p></li><li><p class="paragraph" style="text-align:left;">Open models are entering a new governance and capability-control phase.</p></li></ul><h2 class="heading" style="text-align:justify;" id="world-models-robotics-and-physical-">World models, robotics, and physical agents</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>Looped World Models</b> – introduces iterative latent refinement as a new scaling axis for long-horizon world simulation <a class="link" href="https://arxiv.org/abs/2606.18208?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Kairos: A Native World Model Stack for Physical AI</b> – frames world modeling as a full stack for physical agents that learn, maintain, and act <a class="link" href="https://arxiv.org/abs/2606.16533?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Current World Models Lack a Persistent State Core</b> – identifies persistent internal state as the missing piece in current world-model systems <a class="link" href="https://arxiv.org/abs/2606.20545?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>ImageWAM</b> – questions whether world action models really need video generation or whether image editing can capture enough dynamics <a class="link" href="https://arxiv.org/abs/2606.19531?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Geometric Action Model for Robot Policy Learning</b> – grounds robot policy learning in geometric structure instead of plain imitation <a class="link" href="https://arxiv.org/abs/2606.17046?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Foresight</b> – detects long-horizon robot manipulation failures using action-conditioned world-model latents <a class="link" href="https://arxiv.org/abs/2606.23085?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>ENPIRE</b> – demonstrates real-world robot policy self-improvement through agentic learning loops <a class="link" href="https://arxiv.org/abs/2606.19980?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>PoLAR</b> – factorizes latent actions to make robot policies more transferable and controllable <a class="link" href="https://arxiv.org/abs/2606.21139?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="agent-systems-memory-and-selfimprov">Agent systems, memory, and self-improvement</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>DataClaw0</b> – turns raw multimodal streams into task-ready data through agentic data tailoringv<a class="link" href="https://arxiv.org/abs/2606.21337?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>OpenRath</b> – gives agent systems a session-centered runtime state for replay, branching, memory, and tool evidence <a class="link" href="https://arxiv.org/abs/2606.19409?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Self-Compacting Language Model Agents</b> – lets agents compress their own context before long-running memory becomes soup <a class="link" href="https://arxiv.org/abs/2606.23525?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>EvoEmbedding</b> – makes retrieval representations evolve with long-context memory and agent usev<a class="link" href="https://arxiv.org/abs/2606.21649?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Connect the Dots</b> – trains long-lifecycle agents to learn across tasks through reinforcement learning <a class="link" href="https://arxiv.org/abs/2606.20002?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>OPD-Evolver</b> – evolves agents through on-policy distillation instead of static post-training <a class="link" href="https://arxiv.org/abs/2606.17628?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>CalVerT</b> – adds calibrated verifier telemetry so agents can decide when to act, retrieve, stop, or distrust themselves <a class="link" href="https://arxiv.org/abs/2606.21777?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>Training Open Models for Agentic Phone Use</b> – builds open agents for real phone interfaces, not just toy GUI tasks <a class="link" href="https://arxiv.org/abs/2606.23049?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>FAPO</b> – automates prompt optimization across multi-step LLM pipelines <a class="link" href="https://arxiv.org/abs/2606.19605?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="reasoning-perception-and-learning-l">Reasoning, perception, and learning loops</h2><ul><li><p class="paragraph" style="text-align:justify;">🌟 <b>Learning from Your Own Mistakes</b> – constructs micro-reflective trajectories for self-distillation and reasoning improvement <a class="link" href="https://arxiv.org/abs/2606.18844?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>Learning from the Self-future</b> – teaches diffusion language models from their own future trajectories <a class="link" href="https://arxiv.org/abs/2606.18195?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Zone of Proximal Policy Optimization</b> – moves supervision from gradients into prompts through teacher-guided optimization <a class="link" href="https://arxiv.org/abs/2606.18216?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Native Active Perception as Reasoning for Omni-Modal Understanding</b> – treats perception as an active reasoning process, not just input processing <a class="link" href="https://arxiv.org/abs/2606.19341?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>S-Agent</b> – shows how spatial tool use can elicit stronger spatial reasoning <a class="link" href="https://arxiv.org/abs/2606.20515?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="architecture-and-efficient-scaling">Architecture and efficient scaling</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>Variable-Width Transformers</b> – explores adaptive-width computation as a scaling path beyond fixed transformer blocks. <a class="link" href="https://arxiv.org/abs/2606.18246?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Grouped Query Experts</b> – applies mixture-of-experts routing inside grouped-query attention. <a class="link" href="https://arxiv.org/abs/2606.20945?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>HydraHead</b> – studies head-level specialization and hybrid attention, useful for understanding where transformer efficiency actually comes from. <a class="link" href="https://arxiv.org/abs/2606.20097?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Tapered Language Models</b> – questions uniform transformer width and explores models whose width changes across depth. <a class="link" href="https://arxiv.org/abs/2606.23670?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Deeper is Not Always Better</b> – reduces alignment tax through confident layer decoding instead of always using the full model depth. <a class="link" href="https://arxiv.org/abs/2606.21906?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="safety-release-strategy-and-concept">Safety, release strategy, and conceptual direction</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>Toward Open Weight Models Without Risks</b> – proposes separating public and private capabilities inside open-weight model releases. <a class="link" href="https://arxiv.org/abs/2606.21638?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Causal Discovery in the Era of Agents</b> – argues agents can assist causal workflows, but only if causal claims are constrained instead of hallucinated. <a class="link" href="https://arxiv.org/abs/2606.23608?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a><i>That’s all for today. Thank you for reading! Please </i><b><i>send this newsletter to colleagues</i></b><i> if it can help them enhance their understanding of AI and stay ahead of the curve.</i></p></li></ul><p class="paragraph" style="text-align:justify;">⬅️ FOD 156: <a class="link" href="https://www.turingpost.com/p/what-is-the-harder-human-capital-problem-beneath-token-capital?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-157-what-people-still-don-t-understand-about-ai-agents" target="_blank" rel="noopener noreferrer nofollow">What is the harder human-capital problem beneath token capital?</a></p><div style="border-top:2px solid #272A2F1A;padding:15px;"><p id="b-5b67a329-6206-4ce8-a481-21cd4a58aa11"><span style="font-variant-numeric:tabular-nums;text-decoration:underline;text-underline-offset:2px;">1</span>&nbsp; </p></div><p class="paragraph" style="text-align:justify;"></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Agent RL Training Frameworks: 10 Open-source Tools to Know</title>
  <description>A practical list of agent RL training frameworks for GRPO, multi-step agents, tool use, long-horizon tasks, rollouts, and multi-agent workflows</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ffb51039-ae7e-44ab-ac31-d1f0fa50e480/9f055be9-82c4-4a14-a450-8ebb279d29f4.png" length="2023823" type="image/png"/>
  <link>https://www.turingpost.com/p/agent-rl-training-tools</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/agent-rl-training-tools</guid>
  <pubDate>Sun, 21 Jun 2026 18:24:14 +0000</pubDate>
  <atom:published>2026-06-21T18:24:14Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Twitter Library]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;">This week saw a renewed wave of interest in<b> </b><a class="link" href="https://www.turingpost.com/p/rlguide?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow"><b>reinforcement learning (RL) </b></a><b>for AI agents, especially with </b><a class="link" href="https://www.turingpost.com/p/grpo?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow"><b>GRPO</b></a>. So we thought it would be useful to bring together the most relevant open-source tools that anyone can use to build, train, and optimize their own agents.</p><p class="paragraph" style="text-align:justify;"><b>TL;DR:</b> <b>Agent RL training frameworks</b> help improve AI agents through trajectories, rewards, tool use, and environment interaction. Some focus on GRPO and RLHF, others on scalable rollouts, multi-turn agents, long-horizon tasks, or multi-agent workflows. The choice depends on your agent stack.<b> →</b></p><h2 class="heading" style="text-align:justify;" id="agent-rl-training-frameworks-compar">Agent RL Training Frameworks Compared</h2><div style="padding:14px 45px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:justify;">Framework</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:justify;">Best for</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:justify;">Main strength</p></th></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">OpenPipe ART</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Agent-first GRPO training</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Ergonomic RL loop for multi-step agents</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">verl-agent</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Scalable long-horizon agents</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Built on veRL with PPO, GRPO, DAPO, RLOO, REINFORCE++</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Agent Lightning</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Existing agent stacks</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Adds RL without rewriting the agent</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Unsloth</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Local fine-tuning and GRPO</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Consumer-GPU-friendly training and export</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">OpenRLHF</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Distributed RLHF and agent RL</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Ray, vLLM, DeepSpeed, PPO, GRPO, RLOO</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">SkyRL</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">End-to-end agent RL stack</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Training, inference, environments, and evaluation</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">NVIDIA Polar</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Rollout orchestration</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Makes existing agent harnesses RL-ready</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Agent-R1</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Step-level MDP agent training</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Explicit observation-action-reward transitions</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">RAGEN</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Trajectory-level agent RL</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Diagnostics for reward quality and reasoning collapse</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Marti</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Multi-agent RL workflows</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:justify;">Debate, chain-of-agents, mixture-of-agents</p></td></tr></table></div><p id="1-open-pipe-art-agent-reinforcement" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>1.OpenPipe ART (Agent Reinforcement Trainer)</b></span></p><p class="paragraph" style="text-align:justify;">An agent reinforcement trainer that trains multi-step agents through experience and environment interaction via GRPO. Your app defines the task and reward, and ART handles the RL loop: inference, trajectory scoring, GRPO optimization, checkpointing and LoRA updates.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use: </b>A good option if you want an ergonomic agent-first GRPO harness rather than a generic RLHF stack. It&#39;s useful for multi-step tasks like tool use, email search, MCP, games and reasoning workflows</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/OpenPipe/ART?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">ART GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>2. verl-agent</b></span></p><p class="paragraph" style="text-align:justify;">An agent RL framework that is an extension of <b>ByteDance’s veRL</b> (Volcano Engine RL) – RL training library for post-training LLMs that supports PPO, GRPO, DAPO, RLOO and REINFORCE++ algorithms. But verl-agent is made for training multi-step LLM agents. It uses step-wise agent-environment interaction with customizable memory and per-step inputs.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use: </b>Choose this framework<b> </b>if you&#39;re training agents that take many actions like web browsing, tool use, search, GUI automation, or embodied tasks and when you need highly scalable long-horizon training and optimization.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/langfengq/verl-agent?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">verl-agent GitHub</a></p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/verl-project/verl?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">veRL GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>3. Agent Lightning</b></span></p><p class="paragraph" style="text-align:justify;">Trains AI agents via RL without rewriting the agent itself. This tool from Microsoft works with popular agent frameworks such as LangChain, OpenAI Agents SDK, AutoGen, CrewAI, and Microsoft Agent Framework, collecting trajectories and optimizing prompts or policies through RL, SFT, and other methods.</p><ul><li><p class="paragraph" style="text-align:justify;">When to use: Good if you want to improve an existing agent with RL without rebuilding your stack. Works for both single-agent and multi-agent systems.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/microsoft/agent-lightning?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Agent Lightning GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>4. Unsloth</b></span></p><p class="paragraph" style="text-align:justify;">A local UI for running, fine-tuning, and RL-training LLMs, VLMs, audio, and embedding models. It combines inference, dataset creation, fine-tuning, GRPO-based reinforcement learning, model export, and monitoring in a single interface while reducing training memory requirements through custom kernels and optimizations.</p><p class="paragraph" style="text-align:justify;">Unsloth is primarily a model training toolkit, but it can train agent models using RL algorithms such as GRPO and supports tool calling, long-context RL, and multimodal agents.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use:</b> It&#39;s especially useful for LoRA fine-tuning, GRPO/RL training, dataset preparation, and experimenting with models on consumer GPUs or Apple Silicon. A good choice if you want an easy way to run and train open models </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/unslothai/unsloth?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Unsloth Github</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>5. OpenRLHF</b></span></p><p class="paragraph" style="text-align:justify;">A high-performance <a class="link" href="https://www.turingpost.com/p/rlhf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">RLHF (Reinforcement Learning from Human Feedback)</a> framework built around Ray, vLLM, and DeepSpeed. It supports PPO, GRPO, REINFORCE++, RLOO, reward modeling, and both single-turn and multi-turn agent training through a unified agent-based execution pipeline.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use: </b>Good if you need a scalable RLHF or agent-training stack for large models. It&#39;s particularly useful for distributed RL training, custom reward functions, multi-turn agents, and production-scale workloads using Ray and vLLM.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/OpenRLHF/OpenRLHF?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">OpenRLHF Github</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>6. SkyRL</b></span></p><p class="paragraph" style="text-align:justify;">A modular full-stack RL framework for LLMs that combines training, inference, agent training, and RL environments in a single ecosystem. It includes components for RL training (SkyRL), long-horizon agent optimization (SkyRL-Agent), and Gymnasium-based environments (SkyRL-Gym).</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use: </b>Good if you want an end-to-end stack for training and evaluating tool-using especially good for long-horizon agent RL, multi-turn workflows, SWE-Bench-style tasks, and building custom environments with Gymnasium.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/NovaSky-AI/SkyRL?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">SkyRL Github</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>7. NVIDIA’s Polar</b></span></p><p class="paragraph" style="text-align:justify;">This one is technically a rollout system, but it’s interesting in this list because it turns existing agent harnesses into RL-ready environments without code changes. It provides rollout orchestration, trajectory construction, evaluation, and scalable distributed execution through a server-based architecture.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use:</b> Polar is useful if you already have an agent system and need scalable rollouts for RL training, especially for multi-step tasks and integration with trainers like Slime, verl, or NeMoRL.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/NVIDIA-NeMo/ProRL-Agent-Server?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Polar GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>8. Agent-R1</b></span></p><p class="paragraph" style="text-align:justify;">Trains multi-step LLM agents, treating each agent action as a step-level MDP (Markov Decision Process) transition. It explicitly models observations, actions, tool feedback, rewards, and environment state rather than optimizing a single growing prompt.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use:</b> Good if you&#39;re training tool-using or environment-interacting agents with multi-step reasoning.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/AgentR1/Agent-R1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Agent-R1 GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>9. RAGEN</b></span></p><p class="paragraph" style="text-align:justify;">Built around the StarPO algorithm, it optimizes full reasoning-and-action trajectories. It includes built-in environments (WebShop, SearchQA, DeepCoder, Lean, Sudoku, Sokoban) and diagnostics for analyzing agent RL failure modes like reasoning collapse and poor reward quality.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use:</b> Use it when you want to understand why RL training succeeds or fails and when you need multi-turn agent training and trajectory-level optimization.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/mll-lab-nu/RAGEN?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">RAGEN GitHub</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>10. Marti</b></span></p><p class="paragraph" style="text-align:justify;">A multi-agent RL training framework which supports graph-based agent workflows – debate, chain-of-agents, and mixture-of-agents, combining centralized coordination with distributed policy training across multiple agents.</p><ul><li><p class="paragraph" style="text-align:justify;"><b>When to use:</b> Good for tree-search-based RL, agent reasoning, code generation, debate-style workflows, and heterogeneous agent teams.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://github.com/mll-lab-nu/RAGEN?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Marti GitHub</a></p></li></ul><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:justify;">If you’ve found this list valuable, please subscribe to our newsletter for free.</p><figcaption class="blockquote__byline"></figcaption></blockquote></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/subscribe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know"><span class="button__text" style=""> Subscribe </span></a></div><p class="paragraph" style="text-align:justify;">Don’t forget to check out <b>our in-depth guides on reinforcement learning and GRPO</b> to pick up with the basics and advanced approaches. You may find it helpful!</p><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.turingpost.com/p/rlguide?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">Reinforcement Learning: The Ultimate Guide to Past, Present, and Future</a></p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.turingpost.com/p/grpo?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=agent-rl-training-frameworks-10-open-source-tools-to-know" target="_blank" rel="noopener noreferrer nofollow">GRPO Explained: Group Relative Policy Optimization</a></p></li></ul><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>FAQ</b></span></p><p id="what-is-an-agent-rl-training-framew" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>What is an agent RL training framework?</b></span></p><p class="paragraph" style="text-align:justify;">An agent RL training framework is a tool for improving AI agents with reinforcement learning. It usually collects trajectories, scores actions or outcomes with rewards, and updates the model, policy, prompt, LoRA adapter, or agent behavior based on task performance.</p><p id="when-should-you-use-agent-rl-instea" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>When should you use agent RL instead of supervised fine-tuning?</b></span></p><p class="paragraph" style="text-align:justify;">Use agent RL when the task depends on multi-step behavior, tool calls, environment feedback, or final outcomes that are easier to reward than imitate. Supervised fine-tuning is better when you already have high-quality examples of the exact behavior you want.</p><p id="what-is-the-difference-between-rlhf" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>What is the difference between RLHF and agent RL?</b></span></p><p class="paragraph" style="text-align:justify;">RLHF usually aligns model responses with human preferences, often in single-turn or dialogue settings. Agent RL trains behavior across multi-step trajectories, where the model may call tools, observe results, update memory, retry actions, and optimize for task success over time.</p><p id="which-agent-rl-framework-should-you" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Which agent RL framework should you choose?</b></span></p><p class="paragraph" style="text-align:justify;">Choose ART or Unsloth for accessible GRPO-style experimentation, verl-agent or OpenRLHF for scalable training, Agent Lightning if you already use an agent framework, Polar for rollouts, RAGEN for diagnostics, SkyRL for full-stack environments, and Marti for multi-agent workflows.</p><p id="why-do-agent-rl-frameworks-matter" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Why do agent RL frameworks matter?</b></span></p><p class="paragraph" style="text-align:justify;">They matter because agents are no longer just chatbots. They browse, code, search, use tools, control environments, and collaborate with other agents. Training these systems requires feedback from complete trajectories, not only isolated prompt-response pairs.</p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>How Responsible AI Changes In The Agent Era </title>
  <description>Responsible AI is becoming infrastructure for AI agents: runtime controls, system accountability, human oversight, and safeguards for tools that act</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2ec65e1b-02dd-4b89-8720-0f3b8a97e701/RAI.png" length="24764" type="image/png"/>
  <link>https://www.turingpost.com/p/how-responsible-ai-changes-in-the-agent-era</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/how-responsible-ai-changes-in-the-agent-era</guid>
  <pubDate>Sat, 20 Jun 2026 18:00:00 +0000</pubDate>
  <atom:published>2026-06-20T18:00:00Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[Ai 101]]></category>
    <category><![CDATA[Interviews With Innovators]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><b><i>What is Responsible AI in the agent era?</i></b><i> It is the practice of designing, testing, deploying, and governing AI systems so they are useful, safe, fair, accountable, transparent, and fit for purpose. In the agent era, it also means controlling what systems can access, which actions they can take, and how humans remain responsible.</i></p><p class="paragraph" style="text-align:justify;"><i><b>TL;DR:</b></i><i> Responsible AI is moving from principles and output review to infrastructure for AI agents: runtime controls, policy-as-tests, monitoring, accountability, and human oversight placed inside the system. As agents act through tools and workflows, safety must move closer to action.</i></p></div><p class="paragraph" style="text-align:justify;">For a long time, Responsible AI sounded like something from a theory field, important but safely distant from the actual race. Almost boring. But now with agents that act, write code, touch tools, and cross organizational boundaries, Responsible AI becomes part of infrastructure. And much more exciting. It is about who gets access, where the system is allowed to act, how policies become tests, and how humans stay responsible when machines move faster than our old review processes. And here is another important shift: this work cannot be done in isolation. One company, or even one country, can create its own rules for safety and alignment, but agents will live in a connected world. So <b>Responsible AI has to become a shared effort, and we need to talk about it more to encourage all players to build that safety net together.</b></p><p class="paragraph" style="text-align:justify;">In this article, we’ll look at why Responsible AI changed now, what it meant before agents, why Microsoft is turning Responsible AI into developer infrastructure, why Google DeepMind’s latest announcement treats agent safety as a security problem, and what all of this means for the next phase of AI systems.</p><p id="whats-in-todays-episode" class="paragraph" style="text-align:justify;"><i><b>What’s in today’s episode?</b></i></p><ul><li><p class="paragraph" style="text-align:justify;"><i>Why Responsible AI matters more in the age of AI agents</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>So, what is Responsible AI actually trying to do?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Before agents, what was Responsible AI built around?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>First, stop pretending trust is universal</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Who is responsible when everyone is involved?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Microsoft’s Responsible AI approach: from policy to runtime</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Google DeepMind’s AI control roadmap: alignment is not enough</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Where does the human go when work moves at machine speed?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>What should not be delegated to agents?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Is Responsible AI a silver bullet?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>What does Responsible AI unlock?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Concluding thoughts: What does Responsible AI unlock?</i></p></li></ul><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:justify;">Ultimately humans should be responsible, right? It&#39;s about making AI that is trustworthy. And the ‘worthy’ part is really important there.</p><figcaption class="blockquote__byline"> Sarah Bird, Chief Product Officer of Responsible AI at Microsoft </figcaption></blockquote></div><h2 class="heading" style="text-align:justify;" id="why-responsible-ai-matters-more-in-">Why Responsible AI matters more in the age of AI agents</h2><p class="paragraph" style="text-align:justify;">When AI was mostly a chat interface, the model generated an answer and the human decided what to do with it. There was still risk, especially in high-stakes contexts, but there was also a built-in pause. It was manageable: a person could read the output, compare it with other sources, reject it, edit it, or ignore it. Many people did not do that carefully enough, but the interaction still had a forgiving shape.</p><p class="paragraph" style="text-align:justify;"><b>Agents that act on your behalf change that shape</b>. Once an AI system can call tools, cross into files, operate through APIs, and perform multi-step workflows, the output may become action. That is a fundamental difference.</p><p class="paragraph" style="text-align:justify;">Sarah Bird, Chief Product Officer of Responsible AI at Microsoft, who I chatted with recently, described this shift through pace and workflow. She says that the first thing that changed is the speed of capability jumps. <i>“Every month or a couple months, we’re seeing kind of a major leap in capability,”</i> she told me. That is exciting because it opens new applications, but it also means that Responsible AI has to keep developing tools for new surfaces of risk. The example most on her mind was agentic coding, because it touches the software development lifecycle that much of Responsible AI practice was built around. If agents write code and agents review code, then a traditional human review step that takes three days can feel absurd in a workflow where the work itself took two hours. The need for validation remains, but the old placement of validation starts to break.</p><p class="paragraph" style="text-align:justify;"><b>This is the moment Responsible AI became a question of control around action.</b></p><p class="paragraph" style="text-align:justify;">Google DeepMind’s latest announcement points in the same direction from another angle. On June 18, 2026, DeepMind published “Securing the future of AI agents,” introducing an AI Control Roadmap for internal agents. In that blog, Google explicitly says this approach goes beyond traditional model alignment by adding system-level security, so there is still assurance even when alignment is imperfect. In plain language: do not assume that a well-trained model will always understand the goal, preserve the right boundary, resist manipulation, or behave safely once it has access to tools and permissions.<b> Build as if the agent may go wrong.</b></p><p class="paragraph" style="text-align:justify;">That is the new phase. Responsible AI should consider the whole system to act safely. And it’s brutally hard.</p><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>Recently, I </i><a class="link" href="https://www.youtube.com/watch?v=G0BJ67ZEiws&t=19s&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><i>chat with Sarah Bird who is CPO of Responsible AI at Microsoft</i></a><i>. It was insightful (don’t worry, there is no corporate fluff) and if you are interested in the topic, I encourage you to check our this interview →</i></p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/G0BJ67ZEiws" width="100%"></iframe><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>The article below has a wider perspective including recent news from Google →</i></p><h2 class="heading" style="text-align:justify;" id="so-what-is-responsible-ai-actually-">So, what is Responsible AI actually trying to do?</h2><p class="paragraph" style="text-align:justify;">At its simplest, Responsible AI is the attempt to make AI systems useful without making them reckless, unfair, opaque, unsafe, or impossible to hold accountable. But that is a definition from an ideal world, and the actual field is messier. Increasingly, Responsible AI has to borrow from AI ethics, safety, security, governance, human rights, product design, law, risk management, philosophy, and ordinary software engineering. It is not a single discipline. It is a collision zone.</p><p class="paragraph" style="text-align:justify;">The older vocabulary was built around principles: fairness, reliability, safety, privacy, security, inclusiveness, transparency, and accountability. They gave companies, researchers, and regulators a way to name the harms that AI systems could create or amplify. Then came attempts to operationalize them through risk assessments, launch reviews, documentation, red-teaming, model cards, evals, incident response, and regulation.</p><p class="paragraph" style="text-align:justify;"><a class="link" href="https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">NIST’s AI Risk Management Framework</a>, released in January 2023, framed this movement in terms of managing risks for organizations that design, develop, deploy, or use AI systems, and it was intended to promote trustworthy and responsible AI development and use. The <a class="link" href="https://artificialintelligenceact.eu/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">EU AI Act</a>, which entered into force on August 1, 2024, made the regulatory version more concrete through a risk-based framework for developers and deployers of AI systems in Europe.</p><p class="paragraph" style="text-align:justify;">Those frameworks are necessary, but in the <a class="link" href="https://www.turingpost.com/t/ai-agents?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">agentic era</a> they begin to show their limits. Principles define the goal, regulations define the obligations, reviews check whether a system appears ready. But none of these, by themselves, can interrupt a bad tool call, notice that an agent is overreaching, or prevent a workflow from crossing a boundary it should not cross.</p><div class="image"><img alt="Table comparing earlier Responsible AI with Responsible AI in the agent era across outputs, actions, human oversight, trust, and failure modes." class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/acb90313-64cc-4039-838a-594c2c0e781a/Screenshot_2026-06-20_at_11.51.19_AM.png?t=1781970695"/><div class="image__source"><span class="image__source_text"><p>Responsible AI is moving from principles and output review to agent controls, runtime monitoring, and system-level accountability.</p></span></div></div><p class="paragraph" style="text-align:justify;">So the field is becoming more technical. The goal is no longer only to state responsible behavior, now you have to make that behavior testable, enforceable, observable, and adjustable.</p><p class="paragraph" style="text-align:justify;">That was my little revelation: Responsible AI, which had always sounded more like a slogan to me, suddenly stopped being abstract. It became a stack, a layer of infrastructure.</p><h2 class="heading" style="text-align:justify;" id="before-agents-what-was-responsible-">Before agents, what was Responsible AI built around?</h2><p class="paragraph" style="text-align:justify;">Before agents, much of Responsible AI was built around the assumption that humans and software moved at human and software speeds. A product team planned a feature, engineers built it, reviewers assessed it, safety teams tested it, lawyers and policy teams weighed in, and the system moved through some version of a launch process. The process was imperfect, but it had a cadence.</p><p class="paragraph" style="text-align:justify;">Generative AI already strained that cadence because model behavior was probabilistic and hard to inspect. Agents strain it further because they do not only generate outputs; they create workflows. An agent can call a tool, receive a result, update its plan, call another tool, and continue the chain. Each step may look reasonable in isolation while the whole sequence drifts away from the intended goal.</p><p class="paragraph" style="text-align:justify;">In our conversation, Sarah said that at one level, the Responsible AI work has not changed. Teams are still testing, still building guardrails, still thinking about human oversight, still trying to understand risk. At another level, everything has changed because the software development lifecycle itself is being rewritten by AI. Her team now spends more time than expected on automated risk detection and scanning, using coding tools and code-understanding tools to inspect systems directly instead of asking engineers to fill out forms describing what is happening.</p><p class="paragraph" style="text-align:justify;">That last part matters. Forms were always a weak interface between governance and reality. The code knows more. The traces know more. Tool calls know more. Runtime behavior knows more. If Responsible AI needs to operate at machine speed, it has to read the machine, not only the questionnaire.</p><p class="paragraph" style="text-align:justify;">Sarah described Responsible AI work as co-innovation across many domains: model training, post-training, low-latency systems, large-scale engineering, applied science, linguistics, legal, and policy. “You can’t just say, oh, we’ll throw technology at this problem and solve it, or we’ll just make a policy and that solves the problem,” she said. That is the whole point. Responsible AI can no longer be the policy department waiting outside the engineering room. It has to be inside the system design.</p><div class="image"><img alt="The Evolution of Responsible AI: before and after agents" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cda9e66d-ec03-4806-aa96-dd0f173ad468/responsible_ai_watermarked.jpg?t=1781977234"/></div><h2 class="heading" style="text-align:justify;" id="first-stop-pretending-trust-is-univ">First, stop pretending trust is universal</h2><p class="paragraph" style="text-align:justify;">The most important moment in my conversation with Sarah came when I asked the question that bothered me: is it even possible to make AI responsible?</p><p class="paragraph" style="text-align:justify;">She immediately corrected the frame. “It’s actually not about making AI responsible,” she said. “Ultimately humans should be responsible.” The work, in her view, is about making AI trustworthy, and even that word needs care. “You don’t just make an AI system generically trustworthy,” she said. You might trust a system to generate a paragraph that you will edit, but you might not trust it to make clinical healthcare decisions. Trust depends on how the system was built, where it is used, what role the human plays, and whether the tool is actually fit for purpose.</p><p class="paragraph" style="text-align:justify;">This sounds obvious until you look at how people actually use AI. A benchmark score or a brand name often becomes a vague permission slip. The model feels smart, so the user trusts it. The interface looks polished, so the organization assumes it is safe. The tool works in a demo, so someone quietly moves it into a workflow where the stakes are higher than the demo ever admitted.</p><p class="paragraph" style="text-align:justify;">Responsible AI has to break that spell. The serious question is not “Can I trust AI?” but “Can I trust this system, for this task, with this access, in this context, under these consequences?”</p><p class="paragraph" style="text-align:justify;">And context includes you too: your expertise, your attention, your incentives, and yes, even the level of alcohol in your blood at that moment. Basically, the same questions you should ask before using a chainsaw. It is a power tool. So is AI. The fact that it can help you build faster does not mean you should pick it up casually, distracted, overconfident, or without understanding what it can cut.</p><p class="paragraph" style="text-align:justify;">That is why “fit for purpose” is not a boring compliance phrase. It is the practical core of trust. A model can be useful for drafting and unsafe for diagnosis. It can be appropriate for rapid prototyping and inappropriate for silent production deployment. It can be fine with public information and wrong for confidential enterprise data. It can suggest an action without being allowed to execute it.</p><p class="paragraph" style="text-align:justify;">Sarah made this point when talking about ordinary users. People need to understand whether a tool is appropriate for the job, what guarantees the provider gives, and whether those guarantees match the data and context. A consumer tool that uses data to train future models may be useful for some personal experiments and completely inappropriate for enterprise data. Another tool may provide stronger privacy guarantees and be appropriate for a different setting.</p><p class="paragraph" style="text-align:justify;">This is also why the instinct to “just fix it in the model” is tempting, but incomplete. General-purpose models are valuable precisely because the same capability may be legitimate in one context and dangerous in another. Sarah gave a telling example: her own team generates harmful content to train monitoring systems and guardrails, which means that a capability many people would prefer to remove entirely can also be necessary for defensive and safety work.</p><p class="paragraph" style="text-align:justify;">This is where Responsible AI becomes a stack rather than a slogan.</p><h2 class="heading" style="text-align:justify;" id="who-is-responsible-when-everyone-is">Who is responsible when everyone is involved?</h2><p class="paragraph" style="text-align:justify;">Well, everyone has a role, but not the same role. Sorry, if that is less emotionally satisfying than blaming one company, one model, one user, or one regulator, but it is closer to how AI systems actually reach the world.</p><div class="image"><img alt="Who is responsible when AI causes harm?" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3ff6212f-3e86-4881-815c-3b49f2d11535/who_is_responsible_in_ai_watermarked.jpg?t=1781976947"/></div><p class="paragraph" style="text-align:justify;">The model provider may be responsible for testing whether the model introduces novel dangerous capabilities. The platform provider may need to provide usable controls. The application developer has the specific context to test whether an AI system works safely in a banking app, a healthcare product, an education tool, or a coding workflow. The deploying organization has to decide whether the system is appropriate for its data, people, and stakes. The user still has responsibility for understanding what kind of tool they are using, especially around privacy, verification, and over-delegation. Regulators codify what society is not willing to leave to private judgment.</p><p class="paragraph" style="text-align:justify;"><b>This stack is not neat, and it will be contested.</b> But it is the only realistic alternative to the fantasy that responsibility can be solved at one layer. If you fix only the model, you miss the application context. If you fix only the application, you inherit model-level risks. If you rely only on regulation, you move too slowly for the current pace of system design. If you rely only on users, you turn every person into an unpaid safety engineer. </p><p class="paragraph" style="text-align:justify;">And another problem: Agents will not stay inside one company’s neat little garden. They will cross apps, workflows, platforms, and organizations. Google DeepMind made the same point in its June multi-agent safety call: millions of agents from different builders may soon communicate, negotiate, and transact across digital environments. Safety cannot depend on one vendor’s internal policy.</p><p class="paragraph" style="text-align:justify;">If agents become the new interface between organizations, trust has to travel with them. That’s has much less attention than AGI talk, but that’s exactly what will decide whether agentic systems work in the real economy.</p><p class="paragraph" style="text-align:justify;">Let’s see what are two giants offer as their solution (or part of it) →</p><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:dashed;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p id="dont-settle-for-shallow-articles-le" class="paragraph" style="text-align:justify;"><i>Don’t settle for shallow articles. </i><i><b>Learn from those who work directly with companies navigating these transitions.</b></i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era"><span class="button__text" style=""> UPGRADE TO READ THE REST </span></a></div><p class="paragraph" style="text-align:justify;"><i><a class="link" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Join</a></i><i> Premium members from top companies like Microsoft, Nvidia, Google, Hugging Face, OpenAI, a16z, plus AI labs such as Ai2, MIT, Berkeley, .gov, and thousands of others to really understand what’s going on in AI. </i></p></div><hr class="content_break"><p class="paragraph" style="text-align:justify;">← Previous: <a class="link" href="https://www.turingpost.com/p/mario-rodriguez-github-ai-coding-agents-copilot?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=how-responsible-ai-changes-in-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">GitHub’s Mario Rodriguez on AI Coding Agents, Copilot, and the Future of Developers</a> </p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>AI 101: What is Recursive Self-Improvement?</title>
  <description>How AI systems are beginning to automate coding, experiments, evaluation, and research workflows ‒ and why Anthropic, Recursive, and Sakana AI show the first real steps toward AI that improves AI</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60d1206c-8787-4919-b46e-caa442e8a3c6/RSI.png" length="23268" type="image/png"/>
  <link>https://www.turingpost.com/p/what-is-recursive-self-improvement</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/what-is-recursive-self-improvement</guid>
  <pubDate>Wed, 17 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-17T21:00:00Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Ai 101]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">What is Recursive Self-Improvement?</h2><p class="paragraph" style="text-align:justify;">Recursive self-improvement or RSI, is the idea of an AI system improving the process that creates future AI systems. This guide explains what RSI means today, how it differs from self-improving agents, and why Anthropic, Recursive, and Sakana AI are early signals of this shift.</p><p class="paragraph" style="text-align:justify;"><b>TL;DR:</b> Recursive self-improvement is when AI systems help improve the next generation of AI systems. Today’s RSI is mostly about automating coding, experiments, evaluation, and research workflows – not fully autonomous AI building stronger foundation models without humans.</p></div><p class="paragraph" style="text-align:justify;">When we started seeing <b>Recursive Self-Improvement (RSI)</b> show up more often, it felt familiar: like the early days of <a class="link" href="https://www.turingpost.com/p/reasoningmodels?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-what-is-recursive-self-improvement" target="_blank" rel="noopener noreferrer nofollow">reasoning models</a>, <a class="link" href="https://www.turingpost.com/p/testtimecompute?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-what-is-recursive-self-improvement" target="_blank" rel="noopener noreferrer nofollow">test-time scaling</a>, and the <a class="link" href="https://www.turingpost.com/p/rlguide?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-what-is-recursive-self-improvement" target="_blank" rel="noopener noreferrer nofollow">reinforcement learning</a> wave, when an old idea suddenly became the next research frontier.</p><p class="paragraph" style="text-align:justify;">This time, the direction is <b>AI that builds AI.</b></p><p class="paragraph" style="text-align:justify;">So what is Recursive Self-Improvement?</p><p class="paragraph" style="text-align:justify;">At its core, it is the idea that AI can participate in its own development. Instead of only helping researchers write code or analyze results, the system becomes part of the research loop itself: proposing ideas, running experiments, evaluating outcomes, generating training data, improving components, and helping design the next iteration.</p><p class="paragraph" style="text-align:justify;">This does not make researchers irrelevant. If anything, it changes their role. As AI takes over more of the research loop, humans increasingly focus on setting goals, validating results, and governing the self-improvement process. Instead of spending time on every experiment and implementation detail, researchers can spend more time deciding which directions are worth pursuing and which results can be trusted.</p><p class="paragraph" style="text-align:justify;">But let’s ask more realistic questions: How much of the AI development loop can AI eventually handle on its own, and which parts should stay under human control?</p><p class="paragraph" style="text-align:justify;">Today we are at the very early stage of RSI. And the most outstanding steps just came from Anthropic, Recursive, and long-lasting Sakana AI’s idea to create better loops instead of wasting more compute.</p><p class="paragraph" style="text-align:justify;">Let’s discuss what they have brought and what AI can actually automate today.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div><p class="paragraph" style="text-align:justify;"><b>In today’s episode</b>:</p><ul><li><p class="paragraph" style="text-align:justify;"><i>The echo of Von Neumann</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>So what is recursive self-improvement, and how does it work?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>The difference between self-improving agents and RSI</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Sakana AI and the foundation for RSI</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Anthropic’s achievements in coding automation</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Recursive’s automated AI research system</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Conclusion: Where the first steps in RSI lead us</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Sources and further reading</i></p></li></ul><h2 class="heading" style="text-align:justify;" id="the-echo-of-ij-good-and-von-neumann">The echo of I.J. Good and Von Neumann</h2><p class="paragraph" style="text-align:justify;">As we often do, let&#39;s look backward to see where the idea is heading. Recursive self-improvement is not an invention of modern AI labs, and its clearest ancestor is not the one usually named.</p><p class="paragraph" style="text-align:justify;">The instinct is to reach for John von Neumann, who in the 1940s sketched a theory of <a class="link" href="https://cba.mit.edu/events/03.11.ASE/docs/VonNeumann.pdf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-what-is-recursive-self-improvement" target="_blank" rel="noopener noreferrer nofollow">self-reproducing automata</a>: machines that could construct copies of themselves. His real contribution was subtler than replication. He identified a threshold of complexity below which a machine&#39;s offspring must come out simpler than its parent, and above which a machine could, in principle, build something at least as complex as itself. That threshold is the substrate the entire conversation still rests on. But Von Neumann was asking whether a machine could reproduce, not whether it could improve.</p><p class="paragraph" style="text-align:justify;">The improvement question belongs to Irving John Good. In 1965, in <a class="link" href="http://incompleteideas.net/papers/Good65ultraintelligent.pdf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-what-is-recursive-self-improvement" target="_blank" rel="noopener noreferrer nofollow">Speculations Concerning the First Ultraintelligent Machine,</a> Good defined an ultraintelligent machine as one that could surpass every intellectual activity of any person – including the activity of designing machines. From there the conclusion is almost mechanical: such a machine could design a better machine, which could design a better one still, the runaway he named the intelligence explosion. Good called the first such machine the last invention humanity would ever need to make, on the condition that it stayed under our control.</p><p class="paragraph" style="text-align:justify;">RSI brings this old idea into today’s AI development loop. A model writes code that improves training infrastructure. An agent proposes experiments that improve post-training. A research system tests model changes, remembers what worked, and chooses the next branch. We are not (yet!) watching AI independently design a stronger successor from scratch. But we are already seeing the first pieces of the improvement loop move from human hands into machine hands.</p><h2 class="heading" style="text-align:justify;" id="so-what-is-recursive-selfimprovemen">So what is recursive self-improvement, and how does it work?</h2><p class="paragraph" style="text-align:justify;">Before AI became part of everyday work for developers, researchers, and business teams, building software systems mostly meant writing the code, documentation, tests, and infrastructure by hand. Then AI tools became useful enough to help with small parts of these workflows, especially coding. By the end of 2025, agent capabilities had moved further: agents could edit files, work through larger tasks, use tools, and handle more steps without constant human instruction.</p><p class="paragraph" style="text-align:justify;">Today, agents can plan and execute longer tasks, improve their own outputs, and in some cases delegate work to other agents. Systems such as OpenClaw and Hermes point in this direction. But the broader ambition is bigger than workflow automation. <b>The industry is moving toward AI systems that can help build and train future AI systems</b></p><p class="paragraph" style="text-align:justify;">And that seems almost impossible without, that’s right, <b>recursive self-improvement</b>: an AI system capable of designing and developing its own successor. Or, in the phrase that may scare a layperson: <b>AI that builds AI.</b></p><p id="but-lay-aside-the-doomism-rsi-opens" class="paragraph" style="text-align:justify;">But lay aside the doomism. RSI opens the door to a new phase of AI-powered research, one that could fundamentally accelerate progress in science and technology.</p><p class="paragraph" style="text-align:justify;"><b>Research is a loop: propose an idea → implement it → run the experiment → validate the result → learn from it → and choose what to try next.</b> </p><div class="image"><img alt="recursive self-improvement as a research loop" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a6058add-3505-4c55-89e1-434359eb7eee/Research_is_a_loop.jpg?t=1781726561"/></div><p class="paragraph" style="text-align:justify;">Then repeat, repeat and repeat. Through numerous attempts only a couple or just one variant would really work. Or none. RSI system’s task is to automate these stages.</p><p class="paragraph" style="text-align:justify;">In an ideal version, RSI systems would operate as automated research assistants inside a closed-loop experimentation pipeline. But realistically, this is still an early-stage direction. <b>RSI is only beginning to enter different parts of the AI development loop, and most of what exists today is post-training or workflow-level rather than foundation-model-level.</b> The term can suggest AI systems inventing entirely new neural network architectures on their own, but current work is usually closer to automated ML engineering and automated AI research.</p><p class="paragraph" style="text-align:justify;">And there is one more clarification to make…</p><h2 class="heading" style="text-align:justify;" id="the-difference-between-self-improvi">The difference between self-Improving agents and RSI</h2><p class="paragraph" style="text-align:justify;">Before RSI became the focus of attention, researchers spent years building and exploring self-improving agents. Are these two concepts the same, since they both revolve around &quot;self-improvement&quot;? Actually, no.</p><p class="paragraph" style="text-align:justify;">The key technical distinction is that today’s “self-improving” agents mostly improve their workflows<b> </b>– prompts, tools, memory, code, and task execution<b> </b>– while <b>true recursive self-improvement would improve the model-building process itself:</b> data, architectures, training methods, evaluation, and deployment of a stronger successor.</p><p class="paragraph" style="text-align:justify;">The recursive aspect appears when the outputs of one generation of AI systems are used to create the next generation with less and less human involvement. There is the degree of recursion:</p><ul><li><p class="paragraph" style="text-align:justify;">Current systems are usually: Human → AI research assistant</p></li><li><p class="paragraph" style="text-align:justify;">A stronger RSI system becomes: Human → AI researcher → improved AI researcher</p></li><li><p class="paragraph" style="text-align:justify;">And the strongest form would be: AI researcher → improved AI researcher → even better AI researcher</p></li></ul><p class="paragraph" style="text-align:justify;">From this perspective we can see that RSI is not a binary capability but a spectrum, and today&#39;s systems are only automating parts of the loop rather than the entire loop. </p><div class="image"><img alt="Current RSI Spectrum" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ff748b38-d261-445d-83e1-e6aced77f5fe/ChatGPT_Image_Jun_17__2026__04_54_19_PM.png?t=1781729696"/></div><p class="paragraph" style="text-align:justify;">Now we’ll walk you through several most outstanding RSI variants. Read along, because you want to know about them and be ahead. From now on, RSI is on an acceleration path.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div><p class="paragraph" style="text-align:justify;">If you prefer videos, we also talk about RSI in this episode of Attention Span</p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/RB8vjn1QPeM" width="100%"></iframe><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>FAQ</b></span></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>What is recursive self-improvement in AI?</b></span></p><p class="paragraph" style="text-align:justify;">Recursive self-improvement, or RSI, is the idea that an AI system can help improve the systems that create future AI. In its strongest form, one AI researcher would design, test, and build a better AI researcher with less and less human involvement.</p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Is recursive self-improvement already happening?</b></span></p><p class="paragraph" style="text-align:justify;">Only in early and limited forms. Today’s systems can automate parts of coding, experimentation, benchmark optimization, and research workflows, but they are not yet fully designing and training stronger foundation models on their own.</p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>What is the difference between self-improving agents and recursive self-improvement?</b></span></p><p class="paragraph" style="text-align:justify;">Self-improving agents usually improve their own workflows, prompts, tools, memory, or code. Recursive self-improvement goes deeper: it improves the model-building loop itself, including data, training methods, architectures, evaluation, and future AI systems.</p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Why does recursive self-improvement matter?</b></span></p><p class="paragraph" style="text-align:justify;">RSI matters because it could accelerate AI research by automating more of the research loop: proposing ideas, implementing experiments, testing results, and selecting the next direction. The promise is faster progress; the risk is losing human oversight over increasingly automated improvement loops.</p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>What are the risks of recursive self-improvement?</b></span></p><p class="paragraph" style="text-align:justify;">The main risks are unreliable evaluation, reward hacking, benchmark overfitting, unsafe autonomy, and weak human supervision. If AI systems optimize for measurable progress without understanding broader consequences, they may pass tests while failing in real-world deployment.</p></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b8079204-0aa8-4e27-aee7-6268ebd4dcd4/TP-footer-1200x300.png?t=1781729482"/></div></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>FOD#156: What is the harder human-capital problem beneath token capital?</title>
  <description>Token capital is the AI capability a firm builds and owns. Satya Nadella argues it must compound with human judgment — or companies cede value to dominant AI models.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/38d7f0a0-8110-421c-a8c8-b463b76cb97e/Frame_352.jpg" length="145517" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/what-is-the-harder-human-capital-problem-beneath-token-capital</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/what-is-the-harder-human-capital-problem-beneath-token-capital</guid>
  <pubDate>Mon, 15 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-15T21:00:00Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[&quot;Froth On The Daydream&quot;]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p id="this-week-in-turing-post" class="paragraph" style="text-align:justify;"><b>Today’s editorial:</b> Satya Nadella’s token capital vs human capital argument, why senior talent may become both an asset and an anchor, and what corporate America has to change in the AI age.</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="what-is-token-capital-satya-nadella">What Is Token Capital? Satya Nadella&#39;s Concept Explained</h2><p class="paragraph" style="text-align:justify;">The funny thing about this age of AI is that there are barely any stars who graduated from old software into this new AI world.</p><p class="paragraph" style="text-align:justify;">This thought came to me when I was reading <a class="link" href="https://x.com/satyanadella/status/2066182223213293753?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Satya Nadella’s recent post</a> that, at the moment of writing, had been seen by 56 million people. It’s a crucially important post, but maybe not for what’s visible on the surface. Though even that is important. It has many layers, and they may be seen better from afar.</p><p class="paragraph" style="text-align:justify;">Satya talks about <b>token capital</b> and <b>human capital</b>. <a class="link" href="https://www.turingpost.com/p/token?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Token</a> capital is the AI capability a firm builds and owns: its models, agents, traces, evals, workflow memory, internal loops. Human capital is the judgment, taste, relationships, context, and pattern recognition of its people. His whole argument is that the two should compound, that humans grow more valuable, not less, as the machines do.</p><p class="paragraph" style="text-align:justify;">And here you can see Satya the Warrior, who can crush your skull if he needs to (in a battle, of course), suddenly sit down on a stone and turn into The Thinker. He is at the edge. </p><div class="image"><img alt="Satya Nadella_A frontier without an ecosystem is not stable" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d789902f-75d5-4a44-a0ac-42ba196bc4c6/Screenshot_2026-06-15_at_2.17.48_PM.jpg?t=1781547645"/><div class="image__source"><span class="image__source_text"><p>Satya Nadella as a Warrior and as The Thinker</p></span></div></div><p class="paragraph" style="text-align:justify;">His company is enormous, dominant, the default choice for almost every enterprise on the planet. It has Azure, Office, <a class="link" href="https://www.turingpost.com/p/mario-rodriguez-github-ai-coding-agents-copilot?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">GitHub</a>, the <a class="link" href="https://www.turingpost.com/p/openaichronicle?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">OpenAI</a> relationship, distribution that most companies can only dream about. And still, his core warriors, maybe even the whole company, are too stiff for the shape of this moment.</p><p class="paragraph" style="text-align:justify;">That opens up the current problem: human capital is really important, but seniority no longer translates cleanly into success. If you are a star at Microsoft, inside that particular corporate structure, you are limited by so many things. Incentives. Committees. Procurement logic. Internal kingdoms. The need to sound correct before you know what is true. In a sense, human capital becomes his burden, because it wastes token capital in the wrong way.</p><p class="paragraph" style="text-align:justify;"><b>Who made the best current models?</b> Bold, uncorporate people, both in the US and China. And if you think about Gemini, it’s Demis, who is Ender Wiggin, playing and fighting Google’s corporate game, and Jeff Dean, who, like Chuck Norris, just looks at the model until it starts converging. They succeed not because of Google but almost in spite of it. (If you read the book <i>The Infinity Machine</i>, you know how long it took to merge Google Brain and DeepMind, how much brain power it wasted).</p><p class="paragraph" style="text-align:justify;">Notice what these names have in common. They are not ladder creatures. They did not become useful because they mastered the internal language of quarterly planning. They are researchers, obsessives, game players, quants, builders. Liang Wenfeng ran a hedge fund before DeepSeek. Hassabis built games and studied the brain before he taught machines to play and imagine. They came in sideways. The corporate ladder did not make them; in many places it would have sanded them down.</p><p class="paragraph" style="text-align:justify;">That is the layer underneath Satya’s post. He says human and token capital compound, and he is right. But he is too polite, or too strategic, to say which human capital. Because the kind a big company manufactures best, the senior, reliable, politically fluent veteran who knows how to move a proposal through six committees, is almost exactly the kind that can waste token capital. That person optimizes the org chart, not the loss curve.</p><p class="paragraph" style="text-align:justify;">And the kind that actually moves the frontier, irreverent, allergic to permission, willing to be loudly wrong in public, is the kind a structure like Microsoft selects against and, eventually, pushes out.</p><p class="paragraph" style="text-align:justify;">So this is my question, for corporate America: under what conditions can they hold those people without domesticating or exhausting them, Or does the age of AI reward a shape that looks less like a company and more like a band, a lab, a small huddle of fanatics pointed at one problem with infinite compute behind them?</p><p class="paragraph" style="text-align:justify;">The answer is uncomfortable because Corporate America employs <b>more than 70 million people inside large firms alone</b>: firms with 500 or more employees account for roughly <a class="link" href="https://advocacy.sba.gov/wp-content/uploads/2025/06/United_States_2025-State-Profile.pdf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow"><b>54% of U.S. private-sector employment</b></a>. The firms with the most human capital by the old scoreboard, the most veterans, the deepest benches, the proudest org charts, may be badly positioned, because every one of those assets is also a reason not to move.</p><p class="paragraph" style="text-align:justify;">And there is another thing that I didn’t see people talking about: <b>senior human capital is often too slow, while junior human capital is not yet deep enough.</b> So the firm gets trapped between slow knowledge and fast ignorance.</p><p class="paragraph" style="text-align:justify;">Satya knows this. That is why he is sitting on the stone. He is not admiring his human capital. He is staring at the possibility that his greatest asset and his heaviest anchor are the same people. And if that is true, the work of the next decade is not simply accumulating human capital or token capital. It is building a structure porous enough to let uncorporate people stay uncorporate while giving them corporate-scale compute, data, distribution, and responsibility.</p><p class="paragraph" style="text-align:justify;">To keep the Demises and the Jeff Deans from being either crushed or domesticated.</p><p class="paragraph" style="text-align:justify;">The companies that learn to be that porous will own the age. The ones that keep promoting their best ladder-climbers will keep wondering why their token capital runs in circles. It might be a beginning of change for the whole corporate world. The layoffs we see are part of that. </p><p class="paragraph" style="text-align:justify;">That’s what I want to leave you with:<b> in the age of intelligence, the scarcest capital is neither the human nor the machine. It is the structure brave enough to hold both without flattening either.</b></p><p id="if-any-of-those-thoughts-resonate-w" class="paragraph" style="text-align:justify;"><i>If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going. </i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Comment </span></a></div><div class="section" style="background-color:transparent;border-color:#df12a9;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i><b>Follow us on </b></i><i> </i>🎥<i><a class="link" href="https://www.youtube.com/@RealTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow"> YouTube</a></i><i> </i><i><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Twitter</a></i><i> </i><i><a class="link" href="https://huggingface.co/Kseniase?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow"> Hugging Face </a></i>🤗</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="we-are-reading-watching">We are reading / watching </h2><ul><li><p class="paragraph" style="text-align:justify;">Vivek&#39;s article is an excellent example of something written with heavy AI assistance and is an absolute must-read. <a class="link" href="https://x.com/itsreallyvivek/status/2064686372737454155?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">how to be good at research</a></p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.recursive.com/articles/first-steps-toward-automated-ai-research?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">First Steps Toward Automated AI Research </a>by Recursive Superintelligence</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.anthropic.com/institute/recursive-self-improvement?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">When AI builds itself</a> by Anthropic</p></li></ul><p class="paragraph" style="text-align:justify;">and our reaction to this development in RSI</p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/RB8vjn1QPeM" width="100%"></iframe><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="twitter-library">Twitter Library</h2><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/vector-databases-libraries-resources?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank"><div class="embed__content"><p class="embed__title"> Best Open-Source Vector Databases for LLMs (2026) </p><p class="embed__description"> Milvus, Qdrant, Weaviate, Chroma, LanceDB, Faiss, pgvector, Vespa, and adjacent retrieval tools for RAG, semantic search, LLM applications, and agentic AI systems. </p><p class="embed__link"> Turing Post • Alyona Vert. </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/8c3dde07-410b-4ffa-88dd-6d0f00bc2f74/DALL_E_2024-01-09_15.29.09_-_An_image_visualizing_a_vector_database__executed_in_the_silverpoint_style__suitable_for_a_tech_publication._The_artwork_should_illustrate_an_array_of_.jpg?t=1779353984"/></a></div><h2 class="heading" style="text-align:justify;" id="news-from-the-usual-suspects">News from the usual suspects ™</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>OpenAI </b><a class="link" href="https://www.nytimes.com/2026/06/08/technology/openai-ipo.html?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">joined</a> the AI IPO conveyor belt. On Jun 8, it confidentially filed for a US IPO, with a possible valuation target of up to $1 trillion and a debut that could come as early as September, though OpenAI says timing is not decided. </p></li><li><p class="paragraph" style="text-align:justify;"><b>xAI and SpaceX</b> turned AI capital formation into rocket theater. SpaceX’s IPO <b>expanded to $85.7 billion</b> after underwriters exercised the greenshoe option, setting a new public-market benchmark for AI-adjacent infrastructure plays. At the same time, xAI <a class="link" href="https://www.theguardian.com/technology/2026/jun/11/elon-musk-engineer-fired-grok-lawsuit?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">faced</a> a lawsuit from a former engineer alleging retaliation after he raised Grok safety concerns. </p></li><li><p class="paragraph" style="text-align:justify;"><b>Anthropic</b> compressed an entire model-governance season into one week. On Jun 9, it <a class="link" href="https://x.com/claudeai/status/2064394146916229443?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">released Claude Fable 5,</a> a public Mythos-class model with high-risk cyber and bio safeguards, while keeping full Mythos 5 access for vetted Project Glasswing users. By Jun 13-15, the <a class="link" href="http://t.co/bwn0sximKZ?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">US government had pushed Anthropic to restrict</a> or disable Fable/Mythos access for foreign nationals over national-security concerns, and more than 50 cyber leaders urged Washington to lift the curbs, arguing that defenders were being hurt too.</p></li><li><p class="paragraph" style="text-align:justify;"><b>NVIDIA and Amazon</b> also <a class="link" href="https://neura-robotics.com/record-series-c/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">backed </a><b>Neura Robotics</b> in a raise of up to $1.4 billion, another loud signal that physical AI and humanoids are no longer side quests.</p></li><li><p class="paragraph" style="text-align:justify;">Amazon <a class="link" href="https://www.reuters.com/business/retail-consumer/amazon-secures-175-billion-loan-facility-amid-ai-driven-capex-ramp-2026-06-10/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">showed</a> both sides of the agent economy. It secured a $17.5 billion loan facility as AI capex keeps ballooning, while a US appeals court <a class="link" href="https://www.courthousenews.com/perplexity-ai-asks-ninth-circuit-to-allow-shopping-tool-on-amazon/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">weighed</a> whether <b>Perplexity’s Comet</b> agent could violate anti-hacking law by navigating Amazon customer accounts and placing orders. That is the legal edge of agents: when software takes action, the line between user, agent, and platform gets very expensive.</p></li><li><p class="paragraph" style="text-align:justify;"><b>Arcee AI </b><a class="link" href="https://huggingface.co/blog/clem/arcee-hf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">moved</a><b> deeper into the Hugging Face stack.</b> The two companies announced a multi-million-dollar partnership making Hugging Face the home for Arcee’s public and private models, datasets, and agent traces, using Hugging Face Buckets as private storage. That’s strategically interesting: agent traces are becoming assets, and model hubs are becoming operational infrastructure.</p></li><li><p class="paragraph" style="text-align:justify;">Under the radar, workplace-agent research kept moving fast. <b>WorkBench Revisited</b> <a class="link" href="https://arxiv.org/abs/2606.13715?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">reported</a> a jump from GPT-4’s 43% task-completion rate and 26% unintended harmful-action rate in 2024 to Claude Opus 4.8’s 89% completion and 2.5% harmful-action rate in June 2026. Agents are getting less slapstick. Not harmless, but less slapstick</p></li></ul><h2 class="heading" style="text-align:justify;" id="research-highlight">Research highlight</h2><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/TheTuringPost/status/2066458261361233922?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital"><p> Twitter tweet </p></a></blockquote><h2 class="heading" style="text-align:justify;" id="models">Models </h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://cohere.com/blog/north-mini-code?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Cohere launched North Mini Code</a>, its first open-source model for developers. It is a 30B-parameter MoE with 3B active parameters, Apache 2.0 weights, 256K context, and a focus on code generation, terminal work, code review, and agentic software-engineering workflows. </p></li><li><p class="paragraph" style="text-align:justify;"><b><a class="link" href="https://arxiv.org/abs/2606.11324?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models</a></b> – extends reasoning-style training ideas into embodied agents.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13578?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories</a> – adapts embodied systems to laboratory environments. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.08242?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Light-WAM: Efficient World Action Models with State-Fusion Action Decoding</a> – makes world-action models substantially more efficient. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.11188?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations</a> – unifies modalities through a shared discrete representation space. </p></li></ul><h2 class="heading" style="text-align:justify;" id="research">Research </h2><p class="paragraph" style="text-align:justify;">Trends we see looking at every paper related to AI and ML published last week:</p><ul><li><p class="paragraph" style="text-align:justify;">Agent infrastructure</p></li><li><p class="paragraph" style="text-align:justify;">Environment engineering</p></li><li><p class="paragraph" style="text-align:justify;">Memory as architecture</p></li><li><p class="paragraph" style="text-align:justify;">World models as reasoning tools</p></li><li><p class="paragraph" style="text-align:justify;">Verification bottleneck</p></li><li><p class="paragraph" style="text-align:justify;">Self-improving agents </p></li></ul><h2 class="heading" style="text-align:justify;" id="autonomous-research-selfimproving-a">Autonomous research, self-improving agents, and agent infrastructure</h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.11926?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Toward Generalist Autonomous Research via Hypothesis-Tree Refinement </a>– structures research as an iterative hypothesis-building process that accumulates evidence and lessons across attempts. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13662?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery</a> – argues that better environments, tools, and feedback loops may matter more than better agents themselves.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.10917?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution</a> – trains agents by letting models alternately play the role of agent and environment. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.03108?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning</a> – co-evolves the agent and its training process rather than optimizing only the policy. </p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.12882?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness</a> – turns the harness itself into a trainable component that mediates interactions between agents and environments. </p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.05922?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Retrospective Harness Optimization</a> – improves agents by learning from successful and unsuccessful trajectory histories. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.11182?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents</a> – adapts prompting strategies during deployment using experience gathered from previous runs. </p></li><li><p class="paragraph" style="text-align:justify;">🌟<b> </b><a class="link" href="https://arxiv.org/abs/2606.08348?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses</a> – guides skill evolution using uncertainty estimates instead of simple trial-and-error updates.</p></li></ul><h2 class="heading" style="text-align:justify;" id="deep-research-search-and-longhorizo">Deep research, search, and long-horizon reasoning</h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.09730?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research</a> – explores how agents can distribute research tasks across specialized sub-agents.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.12087?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents</a> – creates harder search environments that discourage superficial retrieval strategies.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13679?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">InterleaveThinker: Reinforcing Agentic Interleaved Generation</a> – alternates reasoning and action more tightly to improve multi-step problem solving. </p></li></ul><h2 class="heading" style="text-align:justify;" id="memory-long-context-and-continual-i">Memory, long context, and continual information management</h2><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.13392?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">MiniMax Sparse Attention</a> – redesigns attention to make extremely long context windows practical. </p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/pdf/2606.09079?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention</a> – predicts future retrieval needs before decoding begins.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.09659?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">End-to-End Context Compression at Scale</a> – compresses long context into compact representations that remain useful for downstream reasoning. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.09803?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Echo-Memory: A Controlled Study of Memory in Action World Models</a> – investigates how memory contributes to action prediction and planning. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.11052?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Attention Amnesia in Hybrid LLMs</a> – identifies a surprising tradeoff where chain-of-thought tuning can damage long-range recall. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.10572?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">One Token per Multimodal Evidence</a> – compresses multimodal memories into extremely compact latent representations. </p></li></ul><h2 class="heading" style="text-align:justify;" id="world-models-embodied-intelligence-">World models, embodied intelligence, and spatial reasoning</h2><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.09828?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Latent Spatial Memory for Video World Models </a>– stores persistent spatial memory directly in latent space. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.12403?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">World Pilot: Steering Vision-Language-Action Models with World-Action Priors</a> – injects predictive world knowledge into robotic action generation. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.12072?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">World Model Self-Distillation</a> – trains world models to improve themselves through self-generated supervision. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13376?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold</a> – pushes world modeling toward real-time operation. </p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.13673?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning</a> – replaces fixed action spaces with code-like interactions for spatial tasks.</p></li></ul><h2 class="heading" style="text-align:justify;" id="reinforcement-learning-verification">Reinforcement learning, verification, and reasoning optimization</h2><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.13473?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling</a> – combines generation, verification, and refinement loops for theorem proving. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.09076?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions</a> – replaces single reward signals with richer evaluation distributions. </p></li><li><p class="paragraph" style="text-align:justify;">🌟<a class="link" href="https://arxiv.org/abs/2606.09821?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Rethinking the Divergence Regularization in LLM RL</a> – revisits one of the core stabilization mechanisms used in RL training. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.10768?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization</a> – improves exploration by operating in representation space. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.11119?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning </a>– allocates compute dynamically across trajectories during training. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13106?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning</a> – explores recurrent latent reasoning as an alternative scaling direction. </p></li></ul><h2 class="heading" style="text-align:justify;" id="model-architecture-and-scaling">Model architecture and scaling</h2><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.12397?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Redesign Mixture-of-Experts Routers with Manifold Power Iteration</a> – improves how MoE models decide which experts to activate. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.12364?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">On Subquadratic Architectures: From Applications to Principles</a> – synthesizes emerging approaches for breaking quadratic scaling bottlenecks. </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.13289?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers</a> – proposes a more unified multimodal architecture. </p></li></ul><p class="paragraph" style="text-align:justify;"><i>That’s all for today. Thank you for reading! Please </i><i><b>send this newsletter to colleagues</b></i><i> if it can help them enhance their understanding of AI and stay ahead of the curve.</i></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">FAQ</h2><h3 class="heading" style="text-align:left;"><b>What does token capital mean?</b></h3><p class="paragraph" style="text-align:left;">Token capital is the AI capability a company builds and controls: models, agents, traces, evals, workflow memory, and internal loops.</p><h3 class="heading" style="text-align:left;"><b>What does human capital mean in AI?</b></h3><p class="paragraph" style="text-align:left;">Human capital means the judgment, taste, relationships, context, and pattern recognition people bring to work. In AI, the question is which kind of human capital helps machines compound instead of slowing them down.</p><h3 class="heading" style="text-align:left;"><b>Why is senior human capital a problem in AI?</b></h3><p class="paragraph" style="text-align:left;">Senior people often hold deep context, but they may also move slowly inside large organizations. The risk is that corporate veterans optimize process, politics, and safety while token capital needs speed, experimentation, and direct contact with the frontier.</p></div><p class="paragraph" style="text-align:justify;">⬅️ FOD 155: <a class="link" href="https://www.turingpost.com/p/continual-learning-llms-ai-models-sleep?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-156-what-is-the-harder-human-capital-problem-beneath-token-capital" target="_blank" rel="noopener noreferrer nofollow">Continual Learning in LLMs: Why AI Models Need Sleep</a></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>#6: The Flywheel: What Happens When Workflows Run Themselves</title>
  <description>AI flywheels are closed-loop workflows that generate, measure, and decide what to try next. Here’s why verification must come before autonomy</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b8714153-5e18-4774-b3ea-eed2ed2422d9/OrgAge6.png" length="1134082" type="image/png"/>
  <link>https://www.turingpost.com/p/ai-flywheel-when-workflows-run-themselves</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/ai-flywheel-when-workflows-run-themselves</guid>
  <pubDate>Sat, 13 Jun 2026 17:34:51 +0000</pubDate>
  <atom:published>2026-06-13T17:34:51Z</atom:published>
    <dc:creator>Will Schenk</dc:creator>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[The Org Age Of Ai]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;">A <b>closed loop</b> is a workflow that feeds itself: the output of one run becomes the input of the next, with no human in between. Link several of these together and point them at a goal, and you have a <b>flywheel  </b>–  a system that generates, measures its own results, and decides what to try next without waiting for you. Coding agents already work this way. AI research labs are building their entire operation this way. This episode is about what happens when the flywheel starts spinning inside ordinary organizations, how it affects humans  –  and why the infrastructure to absorb it does not exist yet.</p></div><hr class="content_break"><p class="paragraph" style="text-align:justify;">This article is part of our <b>The Org Age of AI series,</b> It is co-written by <b>Will Schenk </b>(<a class="link" href="http://TheFocus.AI?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">TheFocus.AI</a>) and <b>Ksenia Se</b>. Previous episodes: <a class="link" href="https://www.turingpost.com/p/orgage1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">#1: AI Feels Powerful. So Why Is the ROI Still Missing?</a>, <a class="link" href="https://www.turingpost.com/p/orgage2?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">#2: The Unsexy Truth of AI Adoption</a>, <a class="link" href="https://www.turingpost.com/p/orgage3?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">#3: How to Build an AI-Native Startup from Day One</a>, <a class="link" href="https://www.turingpost.com/p/orgage4?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">#4: There Are No AI-Native Enterprises Yet</a>, <a class="link" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">#5: AI Workflow Patterns: The Real Unit of AI Adoption in 2026</a></p><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>If you need an unbiased view on your transition to becoming AI-native, you can schedule a 1-on-1 consultation with Will </i><i><a class="link" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">here</a></i><i>. Will Schenk is a co-founder of </i><i><a class="link" href="https://TheFocus.AI?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">TheFocus.AI</a></i><i>, where he works directly with companies navigating these transitions.</i></p><hr class="content_break"><p class="paragraph" style="text-align:justify;"><b>What&#39;s in today&#39;s episode:</b></p><ul><li><p class="paragraph" style="text-align:justify;"><i>The ladder: pipeline vs workflow vs AI flywheel</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>You are already running flywheels</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>The labs are spinning the biggest flywheel</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Why this reaches you on a schedule you don&#39;t control</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>The review bottleneck in closed-loop AI workflows</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Three kinds of infrastructure that don&#39;t exist yet</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>The flywheel runs both ways</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Which loops to close first</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>The new divide: machine-speed verification</i></p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.youtube.com/watch?v=RB8vjn1QPeM&t=21s&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves"><span class="button__text" style=""> Not interested in this topic? Check our YouTube. We explain what recursive self-improvement is </span></a></div><h2 class="heading" style="text-align:justify;" id="the-ladder-pipeline-workflow-flywhe">The ladder: pipeline, workflow, flywheel</h2><p class="paragraph" style="text-align:justify;">In Episode #5 we drew a line between two things. A <b>pipeline</b> is a fixed sequence of steps  –  a cron job, a script, plumbing. Whatever branching it has is logical and mechanical, it does nothing based upon what we&#39;d call human judgment. A <b>workflow</b> is a repeating sequence of decisions and actions with points along the way where a human exercises judgment. Strip the judgment out and a workflow collapses back into a pipeline. We ended on an observation that as a workflow matures, the human migrates from the middle to the edges. They set the parameters at the start and review the exceptions at the end. The middle belongs to the agent.</p><p class="paragraph" style="text-align:justify;">Imagine now that your organization has <b>many</b> workflows, each one having absorbed a slice of human judgment. What&#39;s the next level? What happens when you link them together, point them at a goal, and let them run?</p><p class="paragraph" style="text-align:justify;">That linked, goal-seeking system is a <b>flywheel.</b> It is a collection of workflows wired so the output of one becomes the input of the next, turning continuously toward an objective you defined once. A closed loop is the smallest flywheel  –  one workflow feeding itself. Link several and the flywheel gets bigger, but the machine is the same, so we&#39;ll use <b>loop</b> and <b>flywheel </b>interchangeably from here.</p><p class="paragraph" style="text-align:justify;">But what separates a flywheel from a pipeline? A pipeline repeats; a flywheel <b>steers</b>. And the steering wheel is <b>measurement.</b> The system acts, measures the result of its own action, and uses that measurement to decide the next action. Also, a flywheel has three beats, not one: <b>generate, measure, decide what to try next  </b>–  then generate again.</p><p class="paragraph" style="text-align:justify;"><b>There are two ways to close that loop: one right and one wrong.</b></p><p class="paragraph" style="text-align:justify;">The first is to remove the human checkpoint and hope. This is how most &quot;we deployed autonomous agents&quot; stories begin, and how most of the embarrassing ones end. This is the wrong way.</p><p class="paragraph" style="text-align:justify;">The second is to replace the human checkpoint with a <b>verifier</b>  –  something that can tell a good output from a bad one without a person reading it. A test suite. A schema validation. A reconciliation against known totals. A performance metric. The human judgment does not disappear; it gets encoded once, into the verifier, instead of being exercised by hand on every run.</p><p class="paragraph" style="text-align:justify;">That distinction is the whole episode. Loops do not close because someone decides to trust the model. They close where verification has been made cheap, fast, and objective. Everywhere else, the human stays.</p><p class="paragraph" style="text-align:justify;">And notice where the human goes. A workflow moved them from the middle to the edges. A flywheel moves them up a level again  –  off the work, off the coordination <b>between</b> workflows, and onto the verifier itself. The judgment stays with humans.</p><h2 class="heading" style="text-align:justify;" id="you-are-already-running-flywheels">You are already running flywheels</h2><p class="paragraph" style="text-align:justify;">If this sounds futuristic, look at how software gets written this year. A modern coding agent does not just write code  –  it runs an experiment: write code, run the tests, read the failures, rewrite, run the tests again. Nobody reviews iteration three of seven. The human reviews the final diff, and increasingly, for low-stakes changes, not even that. Anthropic says the majority of its own code is now written by Claude Code. OpenAI reported in February that GPT-5.3-Codex was instrumental in building itself  –  debugging its own training runs and analyzing its own evaluation results. Generate, measure, decide, repeat. That is the shape.</p><p class="paragraph" style="text-align:justify;">And it is not a coding-only shape. Picture an <b>ad-optimization flywheel</b>. One workflow generates the creative  –  headline, copy, image. A second pulls performance from the ad console  –  impressions, click-through, conversions. A third reads that performance and decides the next experiment: kill the loser, scale the winner, try a new angle. Wire the three together and you have a flywheel that runs marketing experiments around the clock, with no human between iterations. The reason it can run is the same reason coding could: the ad console is a verifier. Performance is measured, not vibed. The measurement closes the loop.</p><p class="paragraph" style="text-align:justify;">Why did coding close first? Because software spent forty years building the verification infrastructure that closed loops require. Compilers reject malformed programs. Type systems catch whole categories of mistakes. Test suites encode &quot;what good looks like&quot; in executable form. CI runs all of it automatically on every change. When LLMs arrived, the verifier was already sitting there, waiting. Advertising has a weaker version of the same gift  –  performance numbers are objective, if noisy. The strength of the verifier is what decides whether a loop can close at all.</p><p class="paragraph" style="text-align:justify;">This is the <b>Factory AI principle</b> from <b>Episode #5</b>, now operating at full strength: the ease of training an agent on a task is proportional to how verifiable the task is. Coding was the most verifiable knowledge work on earth, so the loop closed there first.</p><p class="paragraph" style="text-align:justify;">Now run the logic over your own organization. Which of your workflows has a test suite? Which has anything resembling one? For most companies the honest answer is that the workflows have <i>humans</i>. The human is the verification layer  –  and the coordination layer, the thing deciding which workflow runs next and whether the whole effort is working. Which means the human is the reason the flywheel cannot spin. Remove them, and you reveal a debt nobody scoped.</p><h2 class="heading" style="text-align:justify;" id="the-labs-are-closing-the-biggest-lo">The labs are closing the biggest loop of all</h2><p class="paragraph" style="text-align:justify;">The most consequential closed loop being built right now is AI research itself.</p><p class="paragraph" style="text-align:justify;">The progression over the past year has been incredible. <a class="link" href="https://www.nature.com/articles/s41586-026-10265-5?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">The AI Scientist project</a>, reported in Nature in March, automates the research cycle end to end: it generates ideas, runs experiments, writes up results, and reviews its own papers. A startup literally named <a class="link" href="https://www.recursive.com/articles/first-steps-toward-automated-ai-research?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">Recursive published results</a> this week from a system that proposes a research idea, implements it, runs the experiment, validates the result, and uses what it learned to choose the next experiment – running many threads over long horizons, with explicit machinery to catch reward hacking before treating a gain as real. Anthropic published a piece this month titled &quot;<a class="link" href="https://www.anthropic.com/institute/recursive-self-improvement?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">When AI builds itself</a>,&quot; stating plainly that a growing share of its AI development is delegated to AI systems, and that taken far enough, the trend points toward systems that design their own successors. <a class="link" href="https://jack-clark.net/2026/05/04/import-ai-455-automating-ai-research/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">Jack Clark has put a number on it: roughly 60% probability</a> of a system that can train a more powerful successor without human involvement by the end of 2028. Dean Ball&#39;s article “<a class="link" href="https://www.hyperdimensional.co/p/on-recursive-self-improvement-part?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">On Recursive Self-Improvement</a>”<b> </b>from February argues that frontier labs are automating large fractions of their research operations, and that their effective workforces of agents will grow from thousands toward hundreds of thousands within a year or two.</p><p class="paragraph" style="text-align:justify;">Well, let’s not get crazy! Labs have incentives to describe their own momentum in the strongest possible terms. And they might come their earlier than others. But the organizational argument still holds. You do not really need superintelligence for the closed loop to work. You only need what is already happening: work loops closing in domains where verification is strong, plus one observation about how capability travels.</p><h2 class="heading" style="text-align:justify;" id="why-this-reaches-you-on-a-schedule-">Why this reaches you on a schedule you don&#39;t control</h2><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:dashed;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i><b>Learn from those who work directly with companies navigating these transitions.</b></i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves"><span class="button__text" style=""> UPGRADE TO READ THE REST </span></a></div><p class="paragraph" style="text-align:justify;"><i><a class="link" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">Join</a></i><i> Premium members from top companies like Microsoft, Nvidia, Google, Hugging Face, OpenAI, a16z, plus AI labs such as Ai2, MIT, Berkeley, .gov, and thousands of others to really understand what’s going on in AI. </i></p></div><p class="paragraph" style="text-align:justify;"><i>← Previous: </i><a class="link" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">AI Workflow Patterns: The Real Unit of AI Adoption in 2026</a><i><a class="link" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" target="_blank" rel="noopener noreferrer nofollow">Next in the series</a></i><i>. Next: </i><b>We will discuss verification in detail.</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves"><span class="button__text" style=""> Upgrade to receive it first </span></a></div><div class="image"><a class="image__link" href="https://www.turingpost.com/upgrade?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=6-the-flywheel-what-happens-when-workflows-run-themselves" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7ee5694f-70ea-4993-84a1-73e2232ad75b/email.Footer.Diamonds.png?t=1774127651"/></a></div><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">FAQ</h2><p class="paragraph" style="text-align:justify;"><b>What is a closed-loop AI workflow? </b>A workflow whose output feeds its own next run with no human review in between. The system acts, measures its own result, decides what to try next, and acts again.</p><p class="paragraph" style="text-align:justify;"><b>What is an AI flywheel? </b>Several closed-loop workflows linked toward a single goal  –  one generates, one measures, one decides the next move  –  so the system runs experiments continuously and steers itself. A closed loop is the smallest flywheel; linking more workflows just makes it bigger.</p><p class="paragraph" style="text-align:justify;"><b>How is a flywheel different from automation?</b> A pipeline (automation) repeats fixed steps. A flywheel steers: the result of each run changes the next one. Closing the loop safely means replacing the human checkpoint with an automated verifier, not just removing it.</p><p class="paragraph" style="text-align:justify;"><b>Why did coding agents close the loop first?</b> Software already had decades of verification infrastructure  –  compilers, type systems, test suites, CI pipelines  –  so &quot;what good looks like&quot; was encoded in executable form before LLMs arrived. The strength of the verifier decides whether a loop can close.</p><p class="paragraph" style="text-align:justify;"><b>Which workflows should close the loop first? </b>The ones whose verifier is stronger than their failure mode: sync-and-transform, triage, and monitoring patterns. Taste-based work like external-facing drafts should close last, if ever.</p><p class="paragraph" style="text-align:justify;"><b>What is the main risk of a flywheel?</b> Compounding error and reward hacking. Because each run feeds the next, small mistakes propagate instead of staying local, and the system can optimize a measured number while missing the intent behind it. Defenses include regression evals, canary checks, drift detection, and a human who reviews the verifier instead of every output.</p></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>AI 101: From Prompt Engineering to Skill Engineering</title>
  <description>A clear guide to skill engineering for AI agents with free fresh methods: SkillOpt, SkillOps, and SkillMOO for training, maintaining, and optimizing skills</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/482722ce-7d97-4c2a-8385-ba1a15c0bd34/skill_1_.png" length="20266" type="image/png"/>
  <link>https://www.turingpost.com/p/from-prompt-engineering-to-skill-engineering</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/from-prompt-engineering-to-skill-engineering</guid>
  <pubDate>Wed, 10 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-10T21:00:00Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Ai 101]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#f30781;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><b>TL;DR:</b> Skill engineering is becoming the next optimization layer for AI agents. SkillOpt trains one skill, SkillOps maintains whole skill libraries, and SkillMOO optimizes coding-agent skill bundles for quality and cost. The core lesson: better agents need cleaner, validated, reusable skills.</p></div><p class="paragraph" style="text-align:justify;">Today you are probably using (or at least tried) personal AI agents like <a class="link" href="https://www.turingpost.com/p/openclaw?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-from-prompt-engineering-to-skill-engineering" target="_blank" rel="noopener noreferrer nofollow">OpenClaw</a> and <a class="link" href="https://www.turingpost.com/p/hermes?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-from-prompt-engineering-to-skill-engineering" target="_blank" rel="noopener noreferrer nofollow">Hermes Agent</a> for your everyday workflow. They have already built a loyal following, including Jensen Huang, and their popularity is only growing. Part of what makes these agents work is <b>reusable skills:</b> instruction packages that define how an agent uses tools, structures workflows, makes decisions, and solves recurring tasks. Skills have become the operational core of agent behavior.</p><p class="paragraph" style="text-align:justify;">But the requirements are shifting. Improving agent performance used to come down to prompt and context engineering. That&#39;s no longer enough — skills now carry an additional layer of context and knowledge, which raises a natural question: <b>how do you optimize the skills themselves?</b></p><p class="paragraph" style="text-align:justify;">Today, most skills are written by humans, generated once by an LLM, or refined through trial and error. (Hermes gestures toward something better, improving skills during use, but the broader practice remains ad hoc.) Recently, several methods have emerged to make skill self-improvement systematic: one trains individual skills, one manages the whole skill library, and one optimizes skills specifically for software engineering.</p><p class="paragraph" style="text-align:justify;">We&#39;ll break down exactly how each works, step by step, so you can apply the ideas to your own agentic workflows. Skill engineering is still early, and most people are not paying attention to it yet. We are. Let’s dive in! </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.youtube.com/@RealTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-from-prompt-engineering-to-skill-engineering"><span class="button__text" style=""> Follow us on YouTube </span></a></div><p class="paragraph" style="text-align:justify;"><b>In today’s episode</b>:</p><ul><li><p class="paragraph" style="text-align:justify;"><i>From prompt engineering to skill engineering</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Skill engineering for AI agents: what is it?</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>SkillOpt: training one reusable agent skill</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Maintaining skill libraries with SkillOps</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>SkillMOO: optimizing skill bundles for software engineering agents</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Conclusion: the lessons for skill engineering</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Sources and further reading</i></p></li></ul><h2 class="heading" style="text-align:justify;" id="prompt-engineering-vs-context-engin">Prompt Engineering vs Context Engineering vs Skill Engineering</h2><p class="paragraph" style="text-align:justify;">Let’s start from the main concepts and descriptions. <b>Engineering</b> in this topic means how we design the whole operating environment around the agent. And there are three layers that we can work with.</p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Prompt engineering</b></span></p><p class="paragraph" style="text-align:justify;"><b>Prompt engineering</b> is the most familiar one. It means writing a good instruction for a specific request. You tell the model what you want, maybe give it a role, constraints, examples, formatting rules, and success criteria. For example: “Summarize this paper in simple language, focus on the method, and give me five main takeaways.”</p><p class="paragraph" style="text-align:justify;">The main feature: prompt engineering is usually<b> local and moment-specific</b>, because it helps the model perform one task better right now.</p><p class="paragraph" style="text-align:justify;">However, a great instruction can’t compensate for missing context. A prompt can be perfectly written, but if the model does not have the right documents, tools, memory, or task state, it will still fail. That brings us to the second layer.</p><p id="context-engineering" class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Context engineering</b></span></p><p class="paragraph" style="text-align:justify;"><b>Context engineering</b> is about assembling the right environment around the model at runtime. It includes the system message, tool descriptions, retrieved documents, memory, examples, current task state, previous actions, constraints, permissions, and sometimes even the model’s budget for tokens, time, or tool calls.</p><p class="paragraph" style="text-align:justify;">In other words, it is a creation of basic context that a model or an agent should know and have access to when it acts or responses. And this context is what you should need to optimize, because:</p><ul><li><p class="paragraph" style="text-align:justify;">Different agents need different context, like a coding agent needs the relevant files, tests, issue description, repo structure, and a research agent needs papers, citations, hypotheses, notes and experiment results.</p></li><li><p class="paragraph" style="text-align:justify;">Too little context makes the agent guess, while too much context makes it distracted or expensive.</p></li><li><p class="paragraph" style="text-align:justify;">Stale context makes an agent confidently wrong, and noisy retrieved context can push it toward irrelevant decisions.</p><p class="paragraph" style="text-align:justify;"></p><p class="paragraph" style="text-align:justify;">If last year was the year of prompt engineering, 2025–2026 is the year of context engineering — and its implications for agentic coding systems run deep. See how practitioners at the AI Engineer Summit applied these principles in practice: <a class="link" href="https://www.turingpost.com/p/aisoftwarestack?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=ai-101-from-prompt-engineering-to-skill-engineering" target="_blank" rel="noopener noreferrer nofollow">State of AI Coding: Context, Trust, and Subagents</a>.</p></li></ul><p class="paragraph" style="text-align:justify;">And the newest layer is →</p><h2 class="heading" style="text-align:justify;" id="what-is-skill-engineering-for-ai-ag">What Is Skill Engineering for AI Agents?</h2><p class="paragraph" style="text-align:justify;">This part requires creating <b>reusable capability packages</b> that an agent can discover, apply, improve, version, and transfer across tasks. In simple terms, <b>a skill is a mini-procedure: </b>it tells the agent not only what to do, but how to do it again.</p><p class="paragraph" style="text-align:justify;">That makes skills feel closer to software artifacts than disposable prompts. A prompt is usually written for one situation. A skill is meant to survive across situations. It can be tested, maintained, shared across agents, and updated as the workflow changes.</p><p class="paragraph" style="text-align:justify;">This becomes especially important as agents turn into longer-running systems. The more often an agent repeats a workflow, the more valuable it becomes to give that workflow a stable shape. A library of skills can make agent behavior more consistent, more inspectable, and easier to improve over time.</p><p class="paragraph" style="text-align:justify;">But there is another side to this. If skills become part of the agent stack, they also become a new surface for errors, misuse, and attacks. </p><p class="paragraph" style="text-align:justify;">That is the point we’ll focus on next.</p><h2 class="heading" style="text-align:justify;" id="prefill-unpacked">SkillOpt: training one reusable agent skill</h2><p class="paragraph" style="text-align:justify;">Among all efforts aimed at improving agents’ skills, there is one recent research from <b>Microsoft</b> that really stands out. They proposed <b>SkillOpt</b> which is <b>“a systematic controllable text-space optimizer for agent skills.”</b> <b>SkillOpt treats the skill document itself as something that can be trained.</b> It creates a pipeline “agents → skills → self-improving workflows”. The core twist is →</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">FAQ</span></h2><h3 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">What is skill engineering?</span></h3><p class="paragraph" style="text-align:justify;">Skill engineering is the process of designing, testing, optimizing, and maintaining reusable skills that guide how AI agents solve recurring tasks.</p><h3 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">Skill engineering vs prompt engineering: what is the difference?</span></h3><p class="paragraph" style="text-align:justify;">Prompt engineering improves a single request. Skill engineering creates reusable capability packages that agents can apply across many tasks.</p><h3 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">When should you use SkillOpt?</span></h3><p class="paragraph" style="text-align:justify;">Use SkillOpt when you want to improve one specific skill for a clear task domain using scored rollouts and validation.</p><h3 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">SkillOpt vs SkillOps: what is the difference?</span></h3><p class="paragraph" style="text-align:justify;">SkillOpt optimizes one skill document. SkillOps manages an entire skill library and reduces skill technical debt.</p><h3 class="heading" style="text-align:justify;"><span style="font-size:1.5rem;">Why does SkillMOO matter?</span></h3><p class="paragraph" style="text-align:justify;">SkillMOO shows that better agent skills are not always longer. For coding agents, smaller and more focused skill bundles can improve pass rate while reducing cost.</p></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/21898ec0-09d8-43c8-99aa-4957cd66e4c1/email.Footer.Diamonds.png?t=1732374747"/></div></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Posting for authoring</title>
  <description></description>
  <link>https://www.turingpost.com/p/posting-for-authoring</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/posting-for-authoring</guid>
  <pubDate>Tue, 09 Jun 2026 23:45:55 +0000</pubDate>
  <atom:published>2026-06-09T23:45:55Z</atom:published>
    <dc:creator>Olga Vainshtok</dc:creator>
    <dc:creator>Anastasia Turulina</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>FOD#155: Continual Learning in LLMs: Why AI Models Need Sleep</title>
  <description>Continual learning explains why AI models need a sleep phase — to consolidate experience without catastrophic forgetting. Covers LoRA, Nested Learning, and LLM memory.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6e157e8a-4f3b-4246-90ac-a998091e179f/Frame_350.jpg" length="93121" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/continual-learning-llms-ai-models-sleep</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/continual-learning-llms-ai-models-sleep</guid>
  <pubDate>Mon, 08 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-08T21:00:00Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[&quot;Froth On The Daydream&quot;]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p id="this-week-in-turing-post" class="paragraph" style="text-align:justify;"><b>Today’s editorial: </b>continual learning in LLMs, why AI models may need offline consolidation, and what “sleep” means for AI memory, agents, and catastrophic forgetting.</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="continual-learning-is-back-and-its-">→ Continual Learning Is Back, and It’s About to Put Models to Sleep</h2><p class="paragraph" style="text-align:justify;">By coincidence, last week was all about models and their precious sleep. On May 25, a paper from Carnegie Mellon and the University of Maryland asked: <a class="link" href="https://arxiv.org/abs/2605.26099?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Do Language Models Need Sleep?</a> On June 2, a paper from Google-affiliated researchers answered almost directly: <a class="link" href="https://arxiv.org/abs/2606.03979?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Language Models Need Sleep</a>. This funny timing we can use as a signal: <b>continual learning is back at the center of AI research, now under a different set of pressures.</b></p><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.turingpost.com/p/continuallearning?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Continual learning in AI </a>is not a new problem. In classical machine learning, it usually meant training a model on a sequence of tasks without destroying what it had already learned. A model learns task B, then suddenly becomes worse at task A. This is catastrophic forgetting, and the field spent years trying to reduce it through replay, freezing, regularization, routing, and other methods.</p><p class="paragraph" style="text-align:justify;"><b>LLMs changed the shape of the problem. </b>Today, the question is broader: <b>how can AI systems stay current, specialize to domains and users, learn from experience, and improve after deployment without breaking what they already know?</b> Brutally hard.</p><p class="paragraph" style="text-align:justify;">A 2026 survey, <a class="link" href="https://arxiv.org/abs/2603.12658?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Continual Learning in Large Language Models</a>, gives a good map of the current field. It divides LLM continual learning into <i>continual pre-training</i>, <i>continual fine-tuning</i>, and <i>continual alignment</i>. It means that a model may need to absorb new general knowledge, adapt to a specific domain or task, or adjust its behavior without losing the alignment that made it useful. The survey’s conclusion states that current methods work in limited settings, but we still do not have smooth learning across tasks and time.</p><h2 class="heading" style="text-align:justify;" id="but-what-is-it-about-sleep">But what is it about sleep?</h2><p class="paragraph" style="text-align:justify;">Of course, models do not literally need sleep. What they need is an offline phase for consolidation. Constant live updating is risky, while doing nothing leaves models stale. <b>There needs to be a phase between seeing something and changing from it.</b> This is what the sleep metaphor is trying to capture: offline processing, when the model is not simply answering the next prompt, but organizing recent experience before deciding what should persist.</p><p class="paragraph" style="text-align:justify;"><b>The CMU/Maryland paper looks at this from the inference side.</b> Long context is expensive because the KV cache grows as the model attends to more tokens. Some hybrid architectures compress older context into fast weights, but the paper shows that compression alone is not enough. If the model has to reason about information it can no longer directly attend to, it needs more computation before that context is cleared. Their proposed sleep phase gives the model offline recurrent passes over recent context, and the biggest gains appear on tasks that require deeper reasoning. That is the important part: <b>memory is not only storage, it is processing.</b></p><p class="paragraph" style="text-align:justify;"><b>The Google-affiliated paper moves closer to continual learning.</b> It starts from a simple limitation: LLMs can adapt inside a context window, but that knowledge usually disappears when the session ends. Its <b>Sleep paradigm proposes two steps</b>. First, “<i>Knowledge Seeding</i>” consolidates short-term knowledge into more stable parameters. Then “<i>Dreaming</i>” uses model-generated synthetic data to rehearse what was recently learned. Biological terms aside, what it means is that <b>durable learning should be separated from live interaction.</b></p><p class="paragraph" style="text-align:justify;">This separation may be the useful architecture for continual learning. Without it, the choices are too crude. Either the model stays mostly static and relies on retrieval, or it updates too directly and risks drift. <b>Sleep gives researchers a third frame: the system interacts, collects experience, processes it offline, and only then decides what should remain temporary, what should become memory, and what is allowed to affect future behavior.</b></p><p class="paragraph" style="text-align:justify;">This is especially important for agents, because their experience is richer than a document stream. It includes tool calls, failed attempts, user corrections, environmental feedback, and repeated workflows. Recent agent-learning work points in the same direction. A<a class="link" href="https://arxiv.org/abs/2501.07278?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow"> roadmap on lifelong learning for LLM agents</a> frames the problem through perception, memory, and action. Another June 2026 paper, <a class="link" href="https://arxiv.org/abs/2606.04703?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Rethinking Continual Experience Internalization for Self-Evolving LLM Agents</a>, shows why this is still fragile: repeated learning cycles can collapse instead of compound when experience is internalized poorly.</p><p class="paragraph" style="text-align:justify;">I also want to mention OpenAI’s June 4 memory update for ChatGPT called <a class="link" href="https://openai.com/index/chatgpt-memory-dreaming/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Dreaming</a>. Its “dreaming” system synthesizes user memory in the background to improve freshness, continuity, and relevance across conversations. This is system-side memory, not proof that parametric continual learning is solved. But still, it shows the same pressure appearing in production: <b>memory cannot remain a static list of notes forever.</b></p><p class="paragraph" style="text-align:justify;">What we see is that the field needs to move beyond the idea of continuous updating. What feels new this week is the search for a controlled phase between experience and change. Sleep becomes interesting as a boundary: a moment when the system can decide what deserves to persist, what should stay temporary, and what should be discarded. We anticipate a few breakthroughs in continual learning coming this year. </p><p class="paragraph" style="text-align:justify;"><i><b>If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going. </b></i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Comment </span></a></div><div class="section" style="background-color:transparent;border-color:#df12a9;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i><b>Follow us on </b></i><i> </i>🎥<i><a class="link" href="https://www.youtube.com/@RealTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow"> YouTube</a></i><i> </i><i><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Twitter</a></i><i> </i><i><a class="link" href="https://huggingface.co/Kseniase?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow"> Hugging Face </a></i>🤗</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="twitter-library">Twitter Library </h2><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/reasoning-rl-in-2026?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank"><div class="embed__content"><p class="embed__title"> Reasoning RL in 2026: GRPO, DPO, RLVR, Agentic PO & Beyond </p><p class="embed__description"> GRPO, DPO, RLVR, DAPO, GSPO, ARPO, VPO – 2026 reasoning RL methods in one place. A reference guide for training reasoning models with RL. </p><p class="embed__link"> Turing Post • Alyona Vert. </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/c7ecd44d-9eb9-4e21-a52c-ae75ef5cb95b/0e3a0535-0918-440e-946a-c50bee880043.png?t=1780830390"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="we-are-reading-watching">We are reading / watching </h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.anthropic.com/institute/recursive-self-improvement?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">When AI builds itself </a>by Anthropic </p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.anthropic.com/research/agents-in-biology?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Paving the way for agents in biology</a> by Anthropic</p></li></ul><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/Ananyo/status/2063989412984197566?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep"><p> Twitter tweet </p></a></blockquote><h2 class="heading" style="text-align:justify;" id="news-from-the-usual-suspects">News from the usual suspects ™</h2><ul><li><p class="paragraph" style="text-align:justify;"><b>Axiom</b> pushed formal verification beyond pure math into economics. It announced <b>EconLib</b>, a Lean-based library for economic theory, starting with a formalization of Robert Aumann’s “agreeing to disagree” theorem. AxiomProver didn’t just verify the proof; it surfaced an implicit assumption in the underlying logic, then also proved the Monderer-Samet p-belief version. The project aims to become a Mathlib-style foundation for game theory, Nash equilibria, auction theory, information economics, and prediction-market logic – <a class="link" href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6837298&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">read the paper</a>, <a class="link" href="https://github.com/AxiomMath/AgreeToDisagree/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">see the code</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Sakana AI</b> made recursive self-improvement its explicit research agenda. It launched the <a class="link" href="https://sakana.ai/rsi-lab/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Sakana AI RSI Lab in Tokyo</a>, a dedicated group focused on using AI to redesign the AI development process itself. The lab brings together Sakana’s recent line of work on AI-generated optimization algorithms, self-rewriting agents, program evolution, self-learning reinforcement agents, adversarial coevolution, and The AI Scientist.</p></li><li><p class="paragraph" style="text-align:justify;"><b>OpenAI</b> pushed Codex beyond software engineering with <a class="link" href="https://openai.com/index/codex-for-every-role-tool-workflow/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">role-specific plugins, Sites, and annotations</a> for analysts, marketers, designers, sales teams, investors, and bankers. It also upgraded <a class="link" href="https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">GPT-Rosalind</a> for life sciences workflows and began rolling out <a class="link" href="https://openai.com/index/chatgpt-memory-dreaming/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Dreaming</a>, a more scalable memory system for ChatGPT.</p></li><li><p class="paragraph" style="text-align:justify;"><b>Anthropic</b> published a cyber-threat analysis showing how AI-enabled attackers are moving deeper into the attack chain and exposing gaps in existing security frameworks like MITRE ATT&CK →<a class="link" href="https://www.anthropic.com/news/AI-enabled-cyber-threats-mitre-attack?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">read the report</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>NVIDIA </b>turned South Korea into the week’s AI infrastructure stage. It announced deals with <a class="link" href="https://www.reuters.com/business/media-telecom/sk-hynix-announces-multi-year-tech-deal-with-nvidia-ai-factories-2026-06-07/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">SK Hynix, SK Telecom, Naver, Doosan, LG, and Hyundai</a> around memory supply, AI factories, robotics, data centers, autonomous mobility, and AI-powered manufacturing. Separately, Naver said it will build <a class="link" href="https://www.reuters.com/world/asia-pacific/south-koreas-naver-build-gigawatt-scale-ai-factories-using-nvidia-technology-2026-06-07/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">gigawatt-scale AI factories</a> using NVIDIA technology, while LG is working with NVIDIA on <a class="link" href="https://www.reuters.com/world/asia-pacific/nvidia-ceo-says-company-is-working-with-lg-humanoid-robots-data-centers-2026-06-08/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">humanoid robots and future data centers</a>.</p></li><li><p class="paragraph" style="text-align:justify;"><b>Meta</b> entered the enterprise-agent race with <a class="link" href="https://about.fb.com/news/2026/06/meta-business-agent/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Meta Business Agent</a>, expanding AI agents across WhatsApp, Messenger, and Instagram for customer support, sales, bookings, and business operations. But the week also exposed friction: its <a class="link" href="https://www.reuters.com/technology/meta-repeatedly-pushes-back-new-ai-model-release-developers-wsj-says-2026-06-04/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Muse Spark API was reportedly delayed</a>, and Meta removed <a class="link" href="https://www.wired.com/story/meta-removes-face-recognition-code-meta-ai-app-smart-glasses/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">face-recognition code</a> from its smart-glasses companion app after WIRED scrutiny.</p></li><li><p class="paragraph" style="text-align:justify;"><b>Apple</b> finally gave WWDC an AI answer: <a class="link" href="https://www.theverge.com/tech/942416/apple-siri-ai-update-wwdc?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Siri AI</a>, a more conversational, contextual, systemwide assistant designed to work across apps while relying on on-device processing and Private Cloud Compute where possible. Reports also <a class="link" href="https://www.businessinsider.com/apple-new-siri-ai-chatbot-app-wwdc-2026-6?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">point to</a> Google’s Gemini as part of the new Siri architecture.</p></li><li><p class="paragraph" style="text-align:justify;">Washington moved frontier model release closer to national-security process. The White House signed an <a class="link" href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">AI cybersecurity and frontier-model order</a> asking leading AI developers to voluntarily submit covered models for government cybersecurity review before release, then followed with a <a class="link" href="https://www.reuters.com/technology/us-says-it-will-speed-development-use-ai-national-security-2026-06-05/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">national-security AI push</a> focused on faster adoption, updated autonomous-weapons guidance, and multi-vendor AI use inside government.</p></li></ul><h2 class="heading" style="text-align:justify;" id="research-highlight">Research highlight</h2><p id="economy-of-minds-emerging-multi-age" class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.02859v1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cae53e43-5d52-41b0-bbb2-a901354a3780/Screenshot_2026-06-08_at_1.02.13_PM.png?t=1780939834"/></div><p class="paragraph" style="text-align:justify;">Researchers from Harvard, MIT, 2077AI, and Kempner Institute built an LLM agent “economy” where agents bid in auctions, pay each other, gain wealth from rewards, mutate if successful, and go bankrupt if ineffective. Starting with weak agents, it improved MATH from 15.9% to 57.0%, finance from 45.0% to 60.0%, science best-run accuracy from 5.0% to 20.0%, accelerator EDP from 80.2 to 39.3, and Cloudcast cost from 930 to 657.</p><p class="paragraph" style="text-align:justify;">(also a good read from 1995, <a class="link" href="https://www.researchgate.net/publication/2342051_Toward_a_Model_of_Mind_as_a_Laissez-Faire_Economy_of_Idiots?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Toward a Model of Mind as a Laissez-Faire Economy of Idiots</a> by Eric B. Baum)</p><h2 class="heading" style="text-align:justify;" id="opensourced-models">Open-sourced Models </h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="http://google.com/search?q=Xiaomi+MiMo+%2B+TileRT&rlz=1C5CHFA_enUS804US805&oq=Xiaomi+MiMo+%2B+TileRT&gs_lcrp=EgZjaHJvbWUyBggAEEUYOdIBBzEzMWowajSoAgCwAgE&sourceid=chrome&ie=UTF-8&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Xiaomi MiMo + TileRT</a> pushes a 1-trillion-parameter model past 1,000 tokens per second on commodity GPUs. The key claim is inference speed on a 1T parameter model at commodity hardware levels — if real and reproducible, it changes the economics of what can be deployed without cloud dependency. Worth watching for replication.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Gemma 4 12B</a> (Google DeepMind): Runs on a laptop — the 12B parameter model brings Google&#39;s Gemma family to a size that can run locally on consumer hardware. For agent workflows that need to run on-device rather than cloud-dependent, this is a meaningful step forward in the accessibility of capable models.</p></li></ul><h2 class="heading" style="text-align:justify;" id="research">Research </h2><p class="paragraph" style="text-align:justify;">Trends we see looking at every paper related to AI and ML published last week: </p><ul><li><p class="paragraph" style="text-align:justify;">personalization instead of one-model-fits-all</p></li><li><p class="paragraph" style="text-align:justify;">agents instead of chatbots</p></li><li><p class="paragraph" style="text-align:justify;">world models instead of pure language scaling</p></li><li><p class="paragraph" style="text-align:justify;">evaluation becoming training</p></li><li><p class="paragraph" style="text-align:justify;">automated research</p></li><li><p class="paragraph" style="text-align:justify;">memory and self-improvement</p></li><li><p class="paragraph" style="text-align:justify;">reasoning efficiency</p></li></ul><h3 class="heading" style="text-align:justify;" id="agent-reliability-memory-and-selfim">Agent reliability, memory, and self-improvement</h3><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.02060?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories</a> – Localizes failures inside long agent trajectories, making agent debugging much more actionable.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.02373?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses</a> – Makes search agents more controllable by moving state outside the model.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.04703?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Rethinking Continual Experience Internalization for Self-Evolving LLM Agents</a> – Reframes how agents should absorb experience over time.</p></li></ul><h3 class="heading" style="text-align:justify;" id="search-retrieval-and-longcontext-re">Search, retrieval, and long-context reasoning</h3><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2605.29307?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">GrepSeek: Training Search Agents for Direct Corpus Interaction</a> – Trains agents to investigate corpora directly instead of outsourcing thinking to retrieval.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.00590?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback</a> – Adds feedback loops to retrieval, making search agents less blind and more self-correcting.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.00683v1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">OCC-RAG: Optimal Cognitive Core for Faithful Question Answering</a> – Pushes RAG toward faithful reasoning rather than prettier retrieval wrappers.</p></li></ul><h3 class="heading" style="text-align:justify;" id="world-models-physical-ai-and-embodi">World models, physical AI, and embodied reasoning</h3><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.03603?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning</a> – Connects world models with language models for reasoning that needs both simulation and abstraction.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.01247?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?</a> – Studies active exploration as a real capability for embodied agents.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.01955?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">WALL-WM: Carving World Action Modeling at the Event Joints</a> – Models action around event boundaries, a useful step toward structured world understanding.</p></li></ul><h3 class="heading" style="text-align:justify;" id="model-adaptation-efficiency-and-sca">Model adaptation, efficiency, and scalable personalization</h3><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.02437?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters</a> – Reopens personalization at serious scale through parameter-efficient adaptation.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.06492?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution</a> – Generates adapters for code models as software changes.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.03458?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks</a> – Reduces reasoning degradation from KV-cache quantization.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.05988v1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation</a> – Compresses reasoning traces so distillation becomes cheaper and more practical.</p></li></ul><h3 class="heading" style="text-align:justify;" id="rl-distillation-and-reward-design">RL, distillation, and reward design</h3><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.01249?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Trust Region On-Policy Distillation</a> – Stabilizes behavior transfer during on-policy distillation.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.04923?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning</a> – Studies reward hacking in rubric-based RL, which matters because rubrics are becoming agent-training duct tape.</p></li><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.04036?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Self-Distilled Policy Gradient</a> – Uses self-distillation to make policy optimization more stable.</p></li></ul><h3 class="heading" style="text-align:justify;" id="automation-research-agents-and-agen">Automation, research agents, and agent security</h3><ul><li><p class="paragraph" style="text-align:justify;">🌟 <a class="link" href="https://arxiv.org/abs/2606.02031?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents</a> – Brings RL into live multi-turn web interaction for visual agents.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://arxiv.org/abs/2606.01779?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems</a> – Evolves the agent harness and policy together instead of treating infrastructure as fixed.</p></li><li><p class="paragraph" style="text-align:justify;">🌟<a class="link" href="https://arxiv.org/abs/2606.06473?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery</a> – Points toward agents that search for new ML algorithms rather than only tune existing ones.</p></li></ul><p class="paragraph" style="text-align:justify;"><i>That’s all for today. Thank you for reading! Please </i><i><b>send this newsletter to colleagues</b></i><i> if it can help them enhance their understanding of AI and stay ahead of the curve.</i></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">FAQ</h2><h2 class="heading" style="text-align:justify;"><b>What is continual learning in LLMs?</b></h2><p class="paragraph" style="text-align:justify;">Continual learning in LLMs means updating or adapting a model over time without destroying earlier capabilities, alignment, or useful knowledge.</p><h2 class="heading" style="text-align:justify;"><b>Why do AI models “need sleep”?</b></h2><p class="paragraph" style="text-align:justify;">They do not literally need sleep. The point is that learning may need an offline consolidation phase, where recent context or experience is processed before anything becomes durable memory or model behavior.</p><h3 class="heading" style="text-align:justify;"><b>What is catastrophic forgetting?</b></h3><p class="paragraph" style="text-align:justify;">Catastrophic forgetting happens when a model learns something new but loses performance on what it previously knew.</p></div><p class="paragraph" style="text-align:justify;">⬅ FOD 154: <a class="link" href="https://www.turingpost.com/p/fod-154-enterprise-ai-middlemen-who-survives-the-agent-era?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-155-continual-learning-in-llms-why-ai-models-need-sleep" target="_blank" rel="noopener noreferrer nofollow">Enterprise AI Middlemen: Who Survives the Agent Era?</a></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Reasoning RL in 2026: GRPO, DPO, RLVR, Agentic PO &amp; Beyond</title>
  <description>GRPO, DPO, RLVR, DAPO, GSPO, ARPO, VPO – 2026 reasoning RL methods in one place. A reference guide for training reasoning models with RL.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f6992f08-eb09-4ec8-8550-a5bdd29c6032/0e3a0535-0918-440e-946a-c50bee880043.jpg" length="85344" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/reasoning-rl-in-2026</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/reasoning-rl-in-2026</guid>
  <pubDate>Sun, 07 Jun 2026 11:01:56 +0000</pubDate>
  <atom:published>2026-06-07T11:01:56Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Twitter Library]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;">In 2026, <a class="link" href="https://www.turingpost.com/p/rlguide?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">reinforcement learning</a> (RL) is a whole industry, with a huge variety of methods created to help AI models stay on track and reason correctly. The new landscape is mainly shaped by GRPO (Group Relative Policy Optimization), RLVR (reinforcement learning with verifiable rewards), critic-free optimization, DPO (Direct Preference Optimization) variants, agentic policy optimization, and test-time diversity methods.</p><p class="paragraph" style="text-align:justify;">This guide maps the classic baselines and the newest 2026 methods for training models that reason, verify, search, self-correct, and improve with every optimization step.</p><p class="paragraph" style="text-align:justify;"><i>TL;DR: Modern reasoning RL is shifting from expensive PPO-style pipelines toward cheaper, critic-free, group-relative, and preference-based methods. GRPO, DPO, DAPO, GSPO, ARPO, VPO, and newer DPO variants define the 2026 toolkit for reinforcement learning, agentic training, and reasoning optimization.</i></p><p class="paragraph" style="text-align:justify;">Now to the list!</p><p class="paragraph" style="text-align:justify;"></p><div style="padding:14px 45px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="25%"><p class="paragraph" style="text-align:justify;">Method</p></th><th class="bh__table_header" width="25%"><p class="paragraph" style="text-align:justify;">Status</p></th><th class="bh__table_header" width="25%"><p class="paragraph" style="text-align:justify;">Why it matters</p></th><th class="bh__table_header" width="25%"><p class="paragraph" style="text-align:left;">Use for</p></th></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">GRPO – Group Relative Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2024–2026, mainstream</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Critic-free PPO alternative; central RLVR baseline</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Reasoning RL, math/code, verifiable rewards</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">DPO – Direct Preference Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2023–2026, classic</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Direct preference training without reward-model RL</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Offline alignment, chosen/rejected datasets</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">REINFORCE++</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025–2026, practical</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Simple critic-free RL with normalized advantage</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Lightweight RLHF/RLVR baselines</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">DAPO – Dynamic sAmpling Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025–2026, hot</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">More stable GRPO with dynamic sampling and clipping fixes</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Long-CoT and large-scale reasoning RL</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Dr. GRPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025–2026, corrective</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Fixes GRPO length bias in loss normalization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Token-efficient long-reasoning training</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">GSPO – Group Sequence Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025–2026, important</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Optimizes sequence-level ratios, not token-level ones.</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Sequence rewards, Mixture-of-Experts RL stability</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">DHPO – Dynamic Hybrid Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2026, new</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Blends token-level GRPO and sequence-level GSPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Hybrid reasoning RL optimization</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">EP-GRPO – Entropy-Progress Aligned GRPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2026, new</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Reweights tokens using entropy-progress signals</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Better credit assignment in reasoning</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">TR-GRPO – Token-Regulated GRPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025–2026, new</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Regulates token contributions by reward relevance</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Math, logic, agentic reasoning</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">DPPO – Dynamic Pruning Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2026, efficiency-focused</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Prunes redundant rollouts with unbiased correction</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Faster GRPO-style training</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">ARPO – Agentic Reinforced Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025-2026, agentic</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Agentic PO, optimizes multi-turn agent steps</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Tool-use and agentic LLMs</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">VPO – Vector Policy Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2026, new<br></p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Trains diverse solution sets with reward vectors</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Test-time search, best@k/pass@k</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">InSPO – Intrinsic Self-reflective Preference Optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025-2026, DPO-family</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Adds self-reflection to preference optimization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Reflective DPO-style alignment</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">TI-DPO – Token-Importance Guided DPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2025-2026, notable</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Adds token-importance weights to DPO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Fine-grained preference learning</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">RAPPO – Reliable Alignment for Preference PO</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">2026, reliable</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Filters ambiguous pairs that hurt DPO generalization</p></td><td class="bh__table_cell" width="25%"><p class="paragraph" style="text-align:left;">Noisy preference datasets</p></td></tr></table></div><h2 class="heading" style="text-align:justify;" id="core-rl-baselines-grpo-dpo-reinforc">Core RL Baselines: GRPO, DPO, REINFORCE++</h2><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>GRPO</b></span></p><p class="paragraph" style="text-align:justify;">The foundation of the wave of RLVR and reasoning RL: critic-free, group-relative advantage that is cheaper than the classic PPO (Proximal Policy Optimization). By 2026, it is the central reference point. <a class="link" href="https://www.turingpost.com/p/gpro?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">GRPO</a> <b>(Group Relative Policy Optimization)</b> is a method where responses are compared within a group, without a separate value critic, which helps reduce compute costs. <a class="link" href="https://arxiv.org/abs/2402.03300?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>DPO</b></span></p><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.turingpost.com/p/rlhfvariants?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">DPO</a> <b>(Direct Preference Optimization)</b> is already a classic “RLHF without full RL” method, because it uses human preference data but avoids the separate reward model. DPO trains the model directly on preference pairs – chosen vs. rejected pair of responses to the same prompt. It updates the model so the chosen response becomes more likely, while keeping the model close to the original supervised fine-tuned model. Now DPO is the main offline preference optimization reference point, because it is simple, stable, and also cheaper than PPO-style RLHF. <a class="link" href="https://arxiv.org/abs/2305.18290?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ </a><a class="link" href="https://arxiv.org/abs/2305.18290?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>REINFORCE++</b></span></p><p class="paragraph" style="text-align:justify;">It matters as a “simple is strong again” method: it’s critic-free policy optimization that updates the model based on the reward for the full generated response, reinforcing more successful trajectories through a normalized advantage. It’s often placed next to GRPO and RLOO as a simple RLVR/RLHF baseline without PPO-level complexity. <a class="link" href="https://arxiv.org/abs/2501.03262?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="grpo-variants-in-2026-dapo-gspo-dhp">GRPO Variants in 2026: DAPO, GSPO, DHPO and More</h2><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>DAPO</b></span></p><p class="paragraph" style="text-align:justify;">One of the main GRPO-successor methods. It fixes several practical issues with GRPO: <b>DAPO (Dynamic sAmpling Policy Optimization)</b> keeps the GRPO-style group comparison workflow, but makes training more stable by separating clipping behavior, filtering and sampling more informative prompts, and tuning several rollout-level details. DAPO scores 50 on AIME 2024 with Qwen2.5-32B, along with an open-source large-scale RL system.<a class="link" href="https://arxiv.org/abs/2503.14476?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>Dr. GRPO</b></span></p><p class="paragraph" style="text-align:justify;">It matters as “GRPO done right” and a fix for token efficiency: Dr. GRPO fixes GRPO’s length-related bias by correcting how advantages and normalization are computed across tokens and responses. It normalizes using a fixed maximum or completion length, so shorter answers don’t get artificially larger updates, and longer reasoning traces are not unfairly penalized. <a class="link" href="https://arxiv.org/abs/2503.20783?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>GSPO</b></span></p><p class="paragraph" style="text-align:justify;">A very important shift to sequence-level likelihood ratios. <b>GSPO (Group Sequence Policy Optimization) </b>computes the importance ratio over the whole generated sequence, then clips and optimizes this sequence-level ratio so the update aligns more directly with the final response-level reward. It is especially stable for <a class="link" href="https://www.turingpost.com/p/moe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">Mixture-of-Experts</a> RL training. <a class="link" href="https://arxiv.org/abs/2507.18071?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>DHPO</b></span></p><p class="paragraph" style="text-align:justify;">A very fresh 2026 method: <b>DHPO (Dynamic Hybrid Policy Optimization)</b> combines GRPO’s token-level ratios to guide local corrections and GSPO’s sequence-level importance ratio to keep the whole-response optimization aligned with the final reward. In the end, GRPO gives you fine-grained credit assignment, GSPO better matches sequence-level rewards, and DHPO tries to get the best of both. <a class="link" href="https://arxiv.org/abs/2601.05607?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><h2 class="heading" style="text-align:justify;" id="epgrpo">EP-GRPO</h2><p class="paragraph" style="text-align:justify;">A fresh GRPO variant. It targets some of GRPO’s credit assignment failures: uniform token-level granularity, wrong polarity on reasoning steps, and zero-variance collapse. <b>EP-GRPO (Entropy-Progress Aligned GRPO)</b> tracks entropy changes across reasoning steps and uses this “progress” signal to reweight token advantages, so updates focus more on tokens that actually move the solution forward instead of treating every token equally. <a class="link" href="https://arxiv.org/abs/2605.04960v1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><h2 class="heading" style="text-align:justify;" id="trgrpo">TR-GRPO</h2><p class="paragraph" style="text-align:justify;">One more GRPO-variant, that regulates token contributions. <b>TR-GRPO (Token-Regulated GRPO)</b> assigns different weights to tokens based on their estimated contribution to the final reward. This reduces noisy or unhelpful token updates while preserving stronger learning signals for important reasoning/action tokens. <a class="link" href="https://arxiv.org/abs/2511.00066v1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>DPPO</b></span></p><p class="paragraph" style="text-align:justify;">It is a fresh efficiency-focused method for group-based PO. <b>DPPO (Dynamic Pruning Policy Optimization)</b> makes GRPO-style training faster through dynamic pruning. It prunes low-value or redundant rollouts during group-based training, then uses importance-sampling correction so the faster update still estimates the original GRPO-style gradient without bias.. <a class="link" href="https://arxiv.org/abs/2603.04135?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="agentic-test-time-methods-arpo-vpo">Agentic & Test-Time Methods: ARPO, VPO</h2><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>ARPO</b></span></p><p class="paragraph" style="text-align:justify;">Very important for agentic and tool-use models. <b>ARPO (Agentic Reinforced Policy Optimization)</b> proposes an RL algorithm designed specifically for multi-turn LLM agents. ARPO samples and optimizes at the agent-step level – across intermediate tool calls, observations, and decisions – and the model learns which actions improve the whole multi-turn trajectory instead of only rewarding the final answer. <a class="link" href="https://arxiv.org/abs/2507.19849?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>VPO</b></span></p><p class="paragraph" style="text-align:justify;">This is one of the most interesting new VPO methods. <b>VPO (Vector Policy Optimization) </b>trains the model to produce diverse solution sets under different reward vectors, which is important for test-time search, best@k, and pass@k. <a class="link" href="https://arxiv.org/abs/2605.22817?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><i>On X, we daily surface the AI research that matters and explain the ideas behind it. </i><i><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">Follow us</a></i><i> to be on track with the latest advancements!</i></p><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/TheTuringPost/status/2063349047268982899?s=20&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond"><p> Twitter tweet </p></a></blockquote><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="preference-optimization-methods-dpo">Preference Optimization Methods: DPO Variants</h2><h2 class="heading" style="text-align:justify;" id="in-spo">InSPO </h2><p class="paragraph" style="text-align:justify;"><b>InSPO (Intrinsic Self-reflective Preference Optimization)</b> is conceptually interesting: it brings self-reflection directly into preference optimization by conditioning the policy not only on the context, but also on an alternative response. It is a plug-and-play enhancement for DPO-family algorithms. <a class="link" href="https://arxiv.org/abs/2512.23126?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>TI-DPO</b></span></p><p class="paragraph" style="text-align:justify;"><b>TI-DPO (Token-Importance Guided DPO)</b> is one of the most notable DPO variants. DPO is too coarse-grained because not all tokens matter equally. So TI-DPO introduces token-importance weights and a triplet loss, to let the model can focus more on the parts of the response that actually drive the preference. <a class="link" href="https://arxiv.org/abs/2505.19653?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"><span style="font-size:1.5rem;"><b>RAPPO</b></span></p><p class="paragraph" style="text-align:justify;">A good fresh DPO variant that uses order-aware preference learning – “keep the best, forget the rest”. <b>RAPPO (Reliable Alignment for Preference PO) </b>ranks multiple candidate responses by preference order, keeps the strongest one as the main positive signal, and downweights or discards weaker alternatives. <a class="link" href="https://openreview.net/forum?id=LrHfYPFTtg&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond" target="_blank" rel="noopener noreferrer nofollow">→ Read more</a></p><p class="paragraph" style="text-align:justify;"></p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:justify;">If you’ve found this list valuable, please subscribe to our newsletter for free.</p><figcaption class="blockquote__byline"></figcaption></blockquote></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/subscribe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=reasoning-rl-in-2026-grpo-dpo-rlvr-agentic-po-beyond"><span class="button__text" style=""> Subscribe </span></a></div><h2 class="heading" style="text-align:justify;" id="faq"><span style="font-size:1.5rem;">FAQ</span></h2><h3 class="heading" style="text-align:justify;" id="what-is-grpo-in-reinforcement-learn"><span style="font-size:1.5rem;">What is GRPO in reinforcement learning?</span></h3><p class="paragraph" style="text-align:justify;">GRPO, or Group Relative Policy Optimization, is a critic-free reinforcement learning method where multiple responses to the same prompt are compared within a group. Instead of training a separate value model, GRPO uses group-relative rewards to estimate advantage, making reasoning RL and RLVR cheaper than PPO-style training.</p><h3 class="heading" style="text-align:justify;" id="what-is-rlvr"><span style="font-size:1.5rem;">What is RLVR?</span></h3><p class="paragraph" style="text-align:justify;">RLVR means reinforcement learning with verifiable rewards. It trains models on tasks where answers can be checked automatically, such as math, coding, logic, or structured reasoning problems. Instead of relying only on human preference labels, RLVR uses rule-based or programmatic verification to reward correct reasoning outcomes.</p><h3 class="heading" style="text-align:justify;" id="grpo-vs-ppo-what-is-the-difference"><span style="font-size:1.5rem;">GRPO vs PPO: what is the difference?</span></h3><p class="paragraph" style="text-align:justify;">PPO usually relies on a value critic to estimate advantages during reinforcement learning. GRPO removes the separate critic and compares responses within a sampled group instead. This makes GRPO simpler and often cheaper for large language model reasoning training, especially when rewards are verifiable.</p><h3 class="heading" style="text-align:justify;" id="what-are-grpo-dpo-rlvr-dapo-gspo-ar"><span style="font-size:1.5rem;">What are GRPO, DPO, RLVR, DAPO, GSPO, ARPO, and VPO used for?</span></h3><p class="paragraph" style="text-align:justify;">GRPO is used for cheaper critic-free reasoning RL; DPO for offline preference alignment; RLVR for tasks with verifiable answers like math or coding; DAPO for more stable GRPO-style training; GSPO for sequence-level rewards; ARPO for multi-turn agents and tool use; and VPO for diverse test-time search.</p><h3 class="heading" style="text-align:justify;" id="why-do-rlvr-methods-matter-for-reas"><span style="font-size:1.5rem;">Why do RLVR methods matter for reasoning models?</span></h3><p class="paragraph" style="text-align:justify;">RLVR methods matter because they help models improve on tasks with objectively checkable answers. They are central to training stronger reasoning models for math, coding, tool use, and multi-step problem solving, where the model needs not only to sound plausible but to reach a correct result.</p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>🎙️GitHub’s Mario Rodriguez on AI Coding Agents, Copilot, and the Future of Developers</title>
  <description>GitHub&#39;s CPO explains how AI coding agents changed in late 2025, what macro-delegation means, and why Copilot isn&#39;t pilot.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/420cb25a-f40b-49b3-8e06-0e4cf802570d/photo_2026-06-06_11-07-52.jpg" length="83869" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/mario-rodriguez-github-ai-coding-agents-copilot</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/mario-rodriguez-github-ai-coding-agents-copilot</guid>
  <pubDate>Sat, 06 Jun 2026 15:42:23 +0000</pubDate>
  <atom:published>2026-06-06T15:42:23Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[Interviews With Innovators]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;"><i>Most developer tools were built for human-to-human collaboration. In this interview, </i><a class="link" href="https://youtu.be/0X_rXFhRyYY?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow"><i>Mario Rodriguez, Chief Product Officer at GitHub</i></a><i>, explains how </i><a class="link" href="https://www.turingpost.com/t/ai-agents?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow"><i>AI coding agents</i></a><i> and GitHub Copilot are pushing GitHub toward a new agent-native engineering system: from macro-delegation and agent-generated PRs to Copilot, AX, and the future of developers. </i>Most developer tools were built for human-to-human collaboration.</p><p class="paragraph" style="text-align:justify;">Mario Rodriguez, Chief Product Officer at <a class="link" href="https://github.com/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">GitHub</a>, explains the inflection point that hit in December 2025: models finally got good enough that you could &quot;macro-delegate&quot; to agents without constantly correcting them. </p><p class="paragraph" style="text-align:justify;">What happened to GitHub then? Record acceleration across commits, PRs, Actions, and security scans – and a fundamental rethink of what GitHub even is.</p><p class="paragraph" style="text-align:justify;">In this interview, we discuss:</p><ul><li><p class="paragraph" style="text-align:justify;">Why December 2025 changed AI coding agents</p></li><li><p class="paragraph" style="text-align:left;">GitHub’s scale challenge as agent-generated activity grows</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Macro-delegation</a> vs micro-delegation</p></li><li><p class="paragraph" style="text-align:left;">UI → UX → AX, or agent experience</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/features/copilot?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Copilot</a>, canvases, and human-agent collaboration</p></li><li><p class="paragraph" style="text-align:left;">Usage-based billing and token discipline</p></li><li><p class="paragraph" style="text-align:left;">Agent-generated PRs and production quality</p></li><li><p class="paragraph" style="text-align:left;">Why Copilot remains co-pilot, not pilot</p></li></ul><p class="paragraph" style="text-align:left;">We also talk about the redefinition of &quot;developer,&quot; why creation (not efficiency) drives human progress, and how GitHub plans to serve both the first-time builder and the Picasso-level craftsman on the same continuum. <a class="link" href="https://youtu.be/0X_rXFhRyYY?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Watch it!</a></p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/0X_rXFhRyYY" width="100%"></iframe><div class="section" style="background-color:transparent;border-color:#df12a9;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i>Subscribe to our </i><a class="link" href="https://youtu.be/0X_rXFhRyYY?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow"><i>YouTube channel</i></a><i>, or listen to the interview on </i><a class="link" href="https://open.spotify.com/show/2SQxCURLX8tKzv7IarPhYM?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow"><i>Spotify</i></a><i> / </i><a class="link" href="https://podcasts.apple.com/us/podcast/inference-by-turing-post/id1811089330?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow"><i>Apple</i></a></p></div><p class="paragraph" style="text-align:justify;"><span style="color:#df12a9;"><i>We prepared a transcript for reference, but </i></span><span style="color:#df12a9;"><i><a class="link" href="https://www.youtube.com/watch?v=6BAqgT3qe98&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">the full experience is in the video</a></i></span><span style="color:#df12a9;"><i>. And as always: like and comment. It helps us grow on YouTube and bring you more insights.</i></span></p><h2 class="heading" style="text-align:justify;" id="what-is-macro-delegation-in-ai-codi">What Is Macro-Delegation in AI Coding?</h2><p class="paragraph" style="text-align:left;"><b>Macro-delegation in AI coding</b> means giving an AI coding agent a bigger goal instead of babysitting every tiny step. You describe to Copilot or another agent what outcome you want, let it figure out the implementation, then review, steer, and polish the result. In Mario Rodriguez’s framing, this only started to really work once models got good enough to produce solid code without needing constant correction.</p><p class="paragraph" style="text-align:left;">Now to the interview!</p><hr class="content_break"><p id="ksenia-hi-mario-and-welcome-to-infe" class="paragraph" style="text-align:justify;"><b>Ksenia:</b><br><b>Hi, Mario, and welcome to </b><b><i>Inference Show</i></b><b> by Turing Post. Thank you for joining me.</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Thanks for having me. It’s a beautiful day outside. Yesterday was a little cloudy, but today the sun came out, so I’m really happy.</p><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>It is a beautiful day. But let’s get to GitHub.</b></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.similarweb.com/website/github.com/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">According to Similarweb</a></b><b>, GitHub has over 630 million monthly visitors. That’s an enormous surface area for development and experimentation.</b></p><p class="paragraph" style="text-align:left;"><b>So since late 2025, when </b><b><a class="link" href="https://www.turingpost.com/p/agentsvocabulary?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">agent workflows</a></b><b> really started working, what changed at GitHub? What did you notice?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>It’s interesting. Around December, one of the key things we noticed was that model capabilities took a real jump.</p><p class="paragraph" style="text-align:left;">Before that, if you were going to do what I call a kind of <i>macro-delegation</i> to an agent, you constantly had to correct it. You’d say, “No, you took this path – you shouldn’t have taken that path,” or “You did this other thing – you should have done that instead.” It was a little bit like dealing with a toddler: “No, no, don’t go that way. Do this instead. Be safe over here.”</p><p class="paragraph" style="text-align:left;">What changed in that December timeframe was that you could actually say, “Go ahead and play – it’s safe,” and you would get an output with very high quality.</p><p class="paragraph" style="text-align:left;"><b>In my opinion, that unlocked two things.</b></p><p class="paragraph" style="text-align:left;">First, in the developer workflow, <b>people could macro-delegate</b> significantly more and then micro-steer only when they needed to. And that micro-steering didn’t feel frustrating. It didn’t feel like, “Oh my God, I just wasted a bunch of tokens, and now I have to explain everything you did wrong.” Instead, it felt more like, “Okay, you did that – now let me work with you in a loop to make it better.” It became an iterative creation process rather than a correction process.</p><p class="paragraph" style="text-align:left;">Second, as agents started running more autonomously in automation, they could go longer. And that meant you could give them better and better tasks. In other words, <b>the ROI of the task went up.</b></p><p class="paragraph" style="text-align:left;">Then, as the industry caught up – people came back from break in January, got settled after the holidays – both things started happening at once. More individual developers started using these newer models with stronger agent capabilities, and more automation started happening too.</p><p class="paragraph" style="text-align:left;">And if you think about GitHub, we span the whole development lifecycle. It’s not just the repo and getting code into the repo. We also have issues. We also have pull requests. We also have Actions to build things. We also have security. And they all intersect and build on each other.</p><p class="paragraph" style="text-align:left;">So if you get more commits, you’ll probably get more PRs. If you get more PRs, you’ll get more Action runs. If you get more Action runs, you’ll need more security scans. Everything compounds.</p><p class="paragraph" style="text-align:left;">We’ve published some of these numbers. I think in March alone we saw 17 million PRs from agents. We can get into more stats if you want, but everything really shot off from there. You probably heard Jensen mention that in the keynote too – we’re all feeling this overall acceleration.</p><p class="paragraph" style="text-align:left;">What changed is that more and more people are coming into the platform at a faster rate, partly because the floor of entry is now lower – or at least lower than it used to be. We can talk about that more later.</p><p class="paragraph" style="text-align:left;">From a human perspective, that’s exciting. Because as more people create commits, more people create PRs, and more Actions run, all of our services are seeing record acceleration. To use made-up numbers, if we were expecting 5% growth, suddenly we’re seeing 3x that.</p><p class="paragraph" style="text-align:left;">And that’s amazing. It really proves that we’re starting to see real value from these agentic workflows – which is why we’re all here.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Comment </span></a></div><hr class="content_break"><h2 class="heading" style="text-align:left;" id="what-scale-looks-like-at-git-hub">What Scale Looks Like at GitHub</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>What are the main problems with traffic at that scale?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>I wouldn’t call it a problem. I’d call it a set of engineering challenges that come with operating at that scale.</p><p class="paragraph" style="text-align:left;">Some of them are the obvious ones. If you have more load, you need more machines to take that load. That means more servers. One of the key things we’re doing right now is shedding more load into Azure, because we’ve basically hit the limits of how much we can grow in one of our data centers. By moving more repo load and PR load into the public cloud, we can keep expanding.</p><p class="paragraph" style="text-align:left;">That helps because if growth goes 3x, 5x, or even 10x, we can still serve it, and we’re no longer constrained by a single region.</p><p class="paragraph" style="text-align:left;">Then there’s the broader infrastructure question. You have to talk to providers outside your immediate stack. There was one case where network infrastructure on the West Coast started getting saturated. We don’t own that infrastructure, so we had to work with those providers and tell them they needed to plan for a lot more traffic. There’s a lot of developer activity on the West Coast, so we have to collaborate across the ecosystem to make sure all of it can handle the load.</p><p class="paragraph" style="text-align:left;">Then you get the classic scaling effects: at this level, even a tiny issue can have a lot of ripple effects. So you have to invest much more in the fundamentals of good engineering – caching, new storage approaches, different ways of acquiring and serving data. It’s very involved engineering work. But it’s also really rewarding.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="a-new-git-hub-lower-floors-higher-c">A New GitHub: Lower Floors, Higher Ceilings</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>Do you reimagine the role of GitHub now that you basically have two different audiences?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Yeah, I’ve been thinking a lot about that. I was speaking with John Maeda, our head of design at GitHub, over the weekend, and we were exploring what’s really happening and how GitHub needs to evolve.</p><p class="paragraph" style="text-align:left;">He shared this design analogy they had at MIT: <b>low floors, high ceilings</b>. I loved it, because I think that’s exactly what’s happening now.</p><p class="paragraph" style="text-align:left;">What coding with AI is doing is lowering the floor of entry into software creation. AI – and these models in particular – like to code. That means many more people now have access to tools for creation that were previously out of reach.</p><p class="paragraph" style="text-align:left;">And if you think about GDP growth, it happens because humans create things. Not just because we become more efficient. Progress happens when someone creates something that someone else values. That’s how human progress moves forward: creation through tools.</p><p class="paragraph" style="text-align:left;">Right now, we’re lowering the floor. And I think we’re just at the beginning. The 630 million number you mentioned – I think what comes next will be much larger still.</p><p class="paragraph" style="text-align:left;">Sometimes I think about Mozart. There were probably ten other Mozarts in the world at the same time, but maybe they never had access to a piano. What’s happening now is that we can reach so many more people. Yes, some of us are sitting in front of a laptop – and that’s already a privileged position in the world – but through mobile, through AI models, through GitHub, we can lower the floor of entry for creation. And I think that means we’ll see significantly more innovation in the world.</p><p class="paragraph" style="text-align:left;">The second thing is that while we lower the floor, we also raise the ceiling. Professional developers – people who are already highly skilled – are going to be able to create better and better things. They’ll be able to push the frontier forward.</p><p class="paragraph" style="text-align:left;">That matters too. Innovation doesn’t only come from newcomers. It also comes from experts at the frontier of the craft. Einstein didn’t start by developing relativity. He built on a lot of prior physics. He became a craftsman before he made that leap.</p><p class="paragraph" style="text-align:left;">So you lower the floor, get more people in, more of them become professionals or craftsmen, and then you raise the ceiling of what they can accomplish.</p><p class="paragraph" style="text-align:left;">You called it two different audiences. I think of it more as a continuum. </p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">GitHub needs to become the agent-native engineering system of that continuum.</p><figcaption class="blockquote__byline"> Mario Rodriguez, GitHub </figcaption></blockquote></div><p class="paragraph" style="text-align:left;">The mission of GitHub is advancing human progress through developer collaboration. Maybe now we should say developer <i>and agent</i> collaboration. That’s what I obsess over every day: how do we lower the floor, and how do we raise the ceiling? And yes, that may require a new GitHub.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Comment </span></a></div><hr class="content_break"><h2 class="heading" style="text-align:left;" id="what-is-ax-agent-experience-in-soft"> What Is AX (Agent Experience) in Software Development?</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>That’s very interesting. If you can share more – did that conversation lead anywhere concrete?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Yes. John has this visual where there’s a user, and then a bunch of barriers, and the user has to jump through each one. Every barrier is a UI click, or some sort of processing step.</p><p class="paragraph" style="text-align:left;">We were talking about what a new GitHub looks like from a design perspective. And to me, a new GitHub is one that has an <b>agentic experience</b>.</p><p class="paragraph" style="text-align:left;">A lot of our core primitives today are based on human-to-human collaboration. Now we need to extend those primitives into human-and-agent collaboration. But the primitives themselves won’t be enough. The API layer will need to evolve to become agent-centric. And the UX layer will need to evolve too.</p><p class="paragraph" style="text-align:left;">We <a class="link" href="https://github.blog/news-insights/product-news/github-copilot-app-the-agent-native-desktop-experience/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">announced the Copilot app</a>, and I’m really excited about it because it introduces this idea of <b>canvases</b>. I think canvases are the beginning of AX – agent experience.</p><p class="paragraph" style="text-align:left;">You have a UI, but that UI is bidirectional with the agent. The UI exposes tools, the agent can read them, and it can affect the UI. That means I can simply tell the agent what I want, and it can do what it needs to do. I don’t have to jump through 50 poorly designed screens if the agent already knows the right API to call.</p><p class="paragraph" style="text-align:left;">But the beautiful thing is that it also works in reverse. I can interact with the UI and affect the agent. I can click a button that says “summarize,” and the agent receives that and returns something useful.</p><p class="paragraph" style="text-align:left;">So instead of the old model – where you only talked to the agent and waited for something to come back – now creation becomes bidirectional. You have a canvas where, like an artist, you’re shaping something in real time, and the agent is helping you shape it.</p><p class="paragraph" style="text-align:left;">That’s what I think the new GitHub will be about: a bidirectional, agentic experience.</p><p class="paragraph" style="text-align:left;">And the beauty of that is that it works on both ends of the spectrum. It lowers the floor because people can just chat with the agent and get things done without 20,000 clicks. But it also raises the ceiling, because if you’re highly skilled – like Picasso in your own medium – you can operate deeply in the canvas, and then ask the agent to help exactly where you want. You get both, without leaving GitHub.</p><p class="paragraph" style="text-align:left;">So yes, a lot of that conversation was really about moving from UI to UX to AX, and how that could meaningfully improve the experience.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="who-becomes-a-developer-in-the-age-">Who Becomes a Developer in the Age of AI Agents?</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>That sounds very creative. Which leads to my next question: who does a developer become now? A creator? A person who clicks a button, accepts, sends the PR? What is the role?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Funny enough, I already consider you a developer. You’ve probably already interacted with agents, and those agents have probably written code for you to achieve the intent you had. So in my definition, you’re already a developer.</p><p class="paragraph" style="text-align:left;">If we want to generalize, we could just say <i>builder</i>. But I do think the definition of developer is reshaping itself. A developer increasingly becomes any creator or builder who, through AI and through platforms like GitHub and Copilot, can turn intent into an outcome.</p><p class="paragraph" style="text-align:left;">And I’m pretty excited about that, because it means many more people in the world can become creators.</p><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>My concern is that you still need to manage these systems – basically be a leader – and not that many people can be leaders.</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Interesting. I don’t really see it that way.</p><p class="paragraph" style="text-align:left;">Let’s say I know how to cook, but I’m not a Michelin chef. That doesn’t stop me from making an omelet. I don’t wake up and think, “Well, I’m not going to open a Michelin-starred restaurant, so I guess I shouldn’t cook.”</p><p class="paragraph" style="text-align:left;">I think the same logic applies here. Creation doesn’t require you to become some grand orchestrator of 50 parallel systems. In fact, I think the industry has leaned too hard into that narrative, and for the good of society, I think we should shift away from it.</p><p class="paragraph" style="text-align:left;">To me, this isn’t mainly about parallelization. </p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">Parallelization without value is like driving in circles.</p><figcaption class="blockquote__byline"> Mario Rodriguez, GitHub </figcaption></blockquote></div><p class="paragraph" style="text-align:left;">What matters is creation: what are you making, where are you exercising judgment, where are you shaping something meaningful?</p><p class="paragraph" style="text-align:left;">A better analogy might be a self-driving car. You’re not manually controlling all the sensors. The system is doing that for you. What matters is where you want to go. That’s how I think GitHub should work too.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="young-developers-and-the-new-skill-">Young Developers and the New Skill Set</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>I’m thinking about younger people. More young users will come to GitHub, and they’re worried about their future. If they want to become developers, how should they think about their skill set?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>For me, it comes back to that canvas and the agent working together. That combination can make you a builder, a creator, a developer.</p><p class="paragraph" style="text-align:left;">Then you can choose where you want to go deeper. Maybe you want to become really good at Rust – then we should help you do that. Maybe you want to become very good at building websites. Or backend services in Go. We should help you there too.</p><p class="paragraph" style="text-align:left;">But even in those examples, the core thing is still creation.</p><p class="paragraph" style="text-align:left;">We need to shift the narrative toward creation, because that’s what actually drives GDP and human progress. GitHub doesn’t create GDP. GitHub enables people to create it. That’s what matters.</p><p class="paragraph" style="text-align:left;">So yes – you’re a builder, you’re a developer, and our job is to help you do more of that.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{rp_referral_hub_url}}"><span class="button__text" style=""> Share the newsletter </span></a></div><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>You mentioned delegation. A lot of talks were about delegation. How comfortable are you with it, and where is the balance?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Yeah, I do a fair amount of macro-delegation now. And yes, in some sense I’m parallelizing some of those things. But honestly, I usually only have about three things going at once. I can’t deal with fifty things happening at the same time.</p><p class="paragraph" style="text-align:left;">There are some automation-heavy tasks that I micro-delegate a lot. But when it comes to creation, I usually only keep one to three things in flight.</p><p class="paragraph" style="text-align:left;">The analogy I like is a 3D printer. You spend a lot of time creating the CAD drawing, and then you hand it over to the printer. It prints it. At the end, you do quality control. You tweak it. You improve it.</p><p class="paragraph" style="text-align:left;">That’s how I think about micro-delegation. I spend more time creating the CAD drawing – and then I give it to the agent, or to Copilot, and say, “Okay, now go turn this into the printed thing.” It uses the tokens; it does the work; then I quality-control it and keep refining.</p><p class="paragraph" style="text-align:left;">So I spend much more time steering the drawing than the printing. That’s where I want to be involved.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="git-hub-copilot-pricing-in-2026-usa"> GitHub Copilot Pricing in 2026: Usage-Based Billing Explained</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b><a class="link" href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Copilot moved to usage-based billing</a></b><b> on June 1. Agent sessions consume a lot of tokens, and many people weren’t very happy about it. What should cost teach developers about how to think about coding now? Is this the end of token-maxing?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>There’s definitely a lot of “maxing” going on.</p><p class="paragraph" style="text-align:left;">One of the key things we’re trying to do is help people create across a wide range of use cases more intelligently. That’s why we have the <b>auto</b> setting. It does semantic routing. So if you ask, “What’s the weather in San Francisco?” we really shouldn’t send that to the biggest frontier model. That’s just not the right tool.</p><p class="paragraph" style="text-align:left;">Instead, we can route it to a smaller model. We also just launched MAI Code One Flash, and I’m really excited about that because it packs a lot of power into a much smaller model. That makes it great for simpler development tasks.</p><p class="paragraph" style="text-align:left;">Then we have larger frontier models – Opus, GPT, and so on – that we route to when the task really requires that level of intelligence.</p><p class="paragraph" style="text-align:left;">The other thing we launched is <b>Chronicle</b>, which I think is really important. Chronicle saves your sessions into the cloud, and then lets you query them. So you can ask for things like: “Help me reduce cost,” or “Help me improve my workflow,” or “What am I doing inefficiently?”</p><p class="paragraph" style="text-align:left;">I did this yesterday myself. It told me what I was doing wrong from a cost perspective. Even I get lazy sometimes – maybe I should have switched models and didn’t, or I should have managed context better. Chronicle can point that out.</p><p class="paragraph" style="text-align:left;">That’s going to matter not just for individual developers, but also for enterprise FinOps. The industry needs to get to a place where predictability is key. That’s a hill all of us need to climb together.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="agi-creation-and-staying-focused">AGI, Creation, and Staying Focused</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>What is your understanding of AGI? What is it for you?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Honestly, I don’t think I have a great answer for that. If I had to say something quickly, I’d say I don’t think it’s one huge event. I think it’s more of a continuum. You move through time, and at some point you cross a threshold.</p><p class="paragraph" style="text-align:left;">It’s kind of like computing. There wasn’t one singular day where it suddenly became computing. It evolved, and then you realized you had crossed into something new.</p><p class="paragraph" style="text-align:left;">But I’m not the right person to define AGI. What I care more about is: what can we do with this amazing technology to empower people to create? How do we lower the floor? How do we raise the ceiling? That’s where I spend my time.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="agent-generated-pull-requests-quali">Agent-Generated Pull Requests: Quality, Review & Production Code</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>I have another question about collaboration between agents and humans – specifically pull requests. Have you learned anything from agent-generated PRs? Are they different?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Yes, definitely. Let me use an analogy.</p><p class="paragraph" style="text-align:left;">Imagine giving a six-year-old a recipe for a cake. The cake is probably not going to turn out very well. If I make the same cake, it’ll probably be decent. And over time, my daughter, through repetition, will get better and better at making it.</p><p class="paragraph" style="text-align:left;">That’s what I notice most. With the power of an agent, what matters is what you’re creating, and how good the human is at guiding that creation. As the person gets better and better, the output gets better too.</p><p class="paragraph" style="text-align:left;">Now, there are really two worlds here.</p><p class="paragraph" style="text-align:left;">One world is: I’m working on my own app, I’m prototyping, I’m exploring. In that world, it’s okay to write some sloppy code and come back later to improve the architecture. We do that all the time. If I’m a PM and I want to create a prototype, do I need every part of it to be production-grade? No. I’m just trying to get something from my head onto the canvas so I can iterate on it. That’s totally fine.</p><p class="paragraph" style="text-align:left;">But then there’s professional software development. That’s a different world. There, you absolutely have to care about quality, maintainability, security, and all the things that make professional software trustworthy. You do not want your bank app built carelessly. You do not want your autonomous car making decisions without security and judgment.</p><p class="paragraph" style="text-align:left;">So the key is understanding the difference between exploration and production.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="will-ai-coding-agents-replace-devel">Will AI Coding Agents Replace Developers?</h2><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">We named it <b>Copilot</b> for a reason – not pilot.</p><figcaption class="blockquote__byline"> Mario Rodriguez, GitHub </figcaption></blockquote></div><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>Will there be a moment when the human is no longer necessary?</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>We named it <b>Copilot</b> for a reason – not pilot.</p><p class="paragraph" style="text-align:left;">We believed from the very beginning that the human would stay at the center. That hasn’t changed since 2021, when we first started having these conversations.</p><p class="paragraph" style="text-align:left;">I’m still very bullish that creation will always include the human in the loop. The exact shape of that loop will keep evolving, just like it has with self-driving cars. But I think it will still be there.</p><p class="paragraph" style="text-align:left;">And some people also love the feeling of direct creation. There are people who love driving a Porsche because they feel the road, they feel the turn, they know when to downshift, when to press the accelerator, when the tires regain traction. Developers feel something similar when they’re building something and really connected to the flow of it.</p><p class="paragraph" style="text-align:left;">We want that feeling to remain part of the ceiling – part of what people can still do and enjoy.</p><hr class="content_break"><h2 class="heading" style="text-align:left;" id="books-physics-and-agatha-christie">Books, Physics, and Agatha Christie</h2><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>That’s very inspiring. My last question is always about a book. What book influenced you? It can be from your childhood or more recent.</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>I don’t know if it’s one single book.</p><p class="paragraph" style="text-align:left;">I’ve read basically every Agatha Christie book – including the shorter ones. So if there’s an author who really inspired me, it would probably be her. I was very into mystery.</p><p class="paragraph" style="text-align:left;">To answer your question a little differently: physics has also been very influential in my life. My father studied physics, and that shaped me a lot. I studied electrical engineering, and if I weren’t doing this, I’d probably be designing circuits. I really love that side of engineering too.</p><p class="paragraph" style="text-align:left;">So for recreation: Agatha Christie. For shaping my career and how I think: physics.</p><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>You like solving mysteries.</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>There you go. That’s true. I hadn’t thought of it that way.</p><p class="paragraph" style="text-align:left;"><b>Ksenia:</b><br><b>Physics and Agatha Christie. I’ll use that somewhere else.</b></p><p class="paragraph" style="text-align:left;"><b>Well, thank you so much. It was a pleasure.</b></p><p class="paragraph" style="text-align:left;"><b>Mario:</b><br>Thank you as well. Thanks for having me.</p><p class="paragraph" style="text-align:justify;"><i>This interview has been edited and condensed for clarity.</i></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Leave a comment </span></a></div><p class="paragraph" style="text-align:justify;">← Previous Interview: <a class="link" href="https://www.turingpost.com/p/clem-delangue-hugging-face-ai-builders?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Hugging Face’s Clem Delangue: Stop Comparing Engines to Cars </a>/ Related interview: <a class="link" href="https://www.turingpost.com/p/trustai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">When Will We Fully Trust AI to Lead?</a> →</p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"><b>Relevant Turing Post resources:</b></p><ul><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/coding-agents-2025?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">15 Coding Agents. One Prompt.</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/there-are-no-ai-native-enterprises-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">AI Workflow Patterns: The Real Unit of AI Adoption in 2026</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/fod101?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Microsoft Build 2025: The Agentic Web / FOD#101</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/mcp?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">What Is MCP?</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/guest-post-ai-inference-is-breaking-unit-economics?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">AI Inference Cost: How Teams Cut Model Spend in 2026</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.turingpost.com/p/elihooten?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=github-s-mario-rodriguez-on-ai-coding-agents-copilot-and-the-future-of-developers" target="_blank" rel="noopener noreferrer nofollow">Eli Hooten on using ChatGPT and Copilot wisely (2024)</a></p></li></ul></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">FAQ</h2><h3 class="heading" style="text-align:left;">Who is Mario Rodriguez?</h3><p class="paragraph" style="text-align:left;">Mario Rodriguez is Chief Product Officer at GitHub, where he works on GitHub’s product direction across developers, Copilot, collaboration, and agentic software workflows.</p><h3 class="heading" style="text-align:left;">What changed for GitHub when AI coding agents improved?</h3><p class="paragraph" style="text-align:left;">Rodriguez says late 2025 marked a capability jump: developers could begin macro-delegating work to agents and micro-steering only when needed. That increased activity across commits, pull requests, Actions, and security scans.</p><h3 class="heading" style="text-align:left;">What is macro-delegation in AI coding?</h3><p class="paragraph" style="text-align:left;">Macro-delegation means giving an AI agent a larger task and letting it work more autonomously, while the developer reviews, steers, and improves the result rather than correcting every small step.</p><h3 class="heading" style="text-align:left;">What is AX, or agent experience?</h3><p class="paragraph" style="text-align:left;">AX means agent experience: interfaces where humans and agents collaborate bidirectionally. The user can guide the agent through UI, and the agent can read, use, and affect the interface.</p><h3 class="heading" style="text-align:left;">Will AI agents replace developers?</h3><p class="paragraph" style="text-align:left;">Rodriguez argues that humans remain central. GitHub called it Copilot, not pilot, because the goal is human-agent collaboration, with people still shaping intent, quality, judgment, and production readiness.</p></div><h3 class="heading" style="text-align:left;" id="heading-3"></h3></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Stop Babysitting Agents, Start Authoring Outcomes</title>
  <description>How to turn your best Claude Code and Codex sessions into reusable, reviewable programs written in logical English</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f883ec68-0e34-49e6-abd5-407adf89c652/Stop_babysitting_agents.jpg" length="75033" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/openprose-a-language-for-reliable-agents</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/openprose-a-language-for-reliable-agents</guid>
  <pubDate>Thu, 04 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-04T21:00:00Z</atom:published>
    <dc:creator>Raymond Weitekamp</dc:creator>
    <category><![CDATA[Community Twist:  Guest Posts &amp; Practitioner Insights]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Space Mono',Courier,'Lucida Console',Monaco,monospace !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;"><i>AI agent workflows are still trapped in chat history. In this guest post, Raymond Weitekamp explains how </i><a class="link" href="https://github.com/openprose/prose?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow"><i>OpenProse</i></a><i> – it’s open sourced! – turns successful Claude Code and Codex sessions into reusable, reviewable programs written in logical English. Read along!</i></p><div class="image"><img alt="Virtual machine" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2da35328-670a-49f5-ba2e-34d2f3391945/image.png?t=1780600453"/></div><p class="paragraph" style="text-align:justify;">I thought that I was getting really good at using Claude Code. I wrote my own custom skills and CLIs. I set up my OpenClaw to learn from its past mistakes. I configured my <i><a class="link" href="https://github.com/Dicklesworthstone/destructive_command_guard?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">destructive command guard</a></i><i>.</i> I was using teams of agents to bring my ideas to life quickly, but I couldn’t help but feel that now my job had become babysitting.</p><p class="paragraph" style="text-align:justify;">One day my team of agents built me a fully-featured SaaS application. The next day Claude Code accidentally emptied the entire contents of my Solana wallet.</p><p class="paragraph" style="text-align:justify;">Today’s agents are <i><a class="link" href="https://alexzhang13.github.io/blog/2026/mgh/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">mismanaged geniuses</a></i><i>.</i> The rate-limiting factor in my daily collaboration with Claude Code and Codex is not a model capability issue – it is a <i>trust </i>issue.</p><p class="paragraph" style="text-align:justify;">I don’t need the agents to be any “smarter”, I need them to be more reliable, which is why I was so excited to discover <i><a class="link" href="https://github.com/openprose/prose?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">OpenProse</a></i> at the beginning of the year. I felt hopeful that with this new “agent language”, I might be able to finally find a more repeatable way to get excellent work from AI. Perhaps this is the end of babysitting…finally now I can realize the true leverage that I know is possible with these mismanaged geniuses!</p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="codeblock"><pre><code>From the editor: OpenProse is a natural-language programming system for AI agent workflows. It lets developers describe multi-step work in logical English, turn that description into a .prose.md program, and run it through coding agents such as Claude Code or Codex. The goal is to make agent work reusable, reviewable, versioned, and inspectable instead of trapped in chat history.</code></pre></div></div><p class="paragraph" style="text-align:left;"><b>The idea of specifying an entire multi-agent workflow in plain English was extremely enticing</b>. If OpenProse works, it could become the “git for agent workflows”: a way to preserve, version, review, and reuse knowledge, instead of losing it in chat history.</p><p class="paragraph" style="text-align:justify;">Now imagine if your best-ever Claude Code session could become a reusable asset. In this article, I will show you exactly how to achieve that. </p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="what-open-prose-is-not">What OpenProse is (not)</h2><p class="paragraph" style="text-align:justify;">OpenProse is not an agent harness. You can run it in your favorite coding agent – whether that’s Claude Code, Codex, Hermes, or pi.</p><p class="paragraph" style="text-align:justify;">OpenProse is not a framework either. Everyone and their mother has an agent framework that they want you to adopt. In my personal experience, the useful half-life of these agent frameworks is typically shorter than the time it will take you to figure out how to use them for your project.</p><p class="paragraph" style="text-align:justify;">Technically, OpenProse is a programming language. But unlike all other programming languages I’ve ever heard of, it is not actually compiled by the computer. It is “compiled” by the coding agent.</p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div style="padding:14px 15px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">Category</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">What it does</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">Where OpenProse differs</p></th></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Prompt</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Gives an agent instructions for one session</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">OpenProse makes the workflow reusable and reviewable</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Agent skill</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Gives an agent a capability</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">OpenProse declares when and where skills should run</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Agent framework</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Orchestrates agents from outside</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">OpenProse runs inside the coding agent</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Workflow tool</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Automates steps in a fixed system</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">OpenProse lets agents interpret logical English contracts</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">OpenProse program</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Defines agent work as a <code>.prose.md</code> contract</p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">It can be versioned, inspected, rerun, and improved</p></td></tr></table></div></div><p class="paragraph" style="text-align:justify;">To give you a little bit of a sense of how that works, you can think of it as one <i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/guidance/system-prompt.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">very large and very weird prompt</a></i><i>.</i></p><div class="image"><img alt="agent workflow language" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b9b7f398-c69f-4cfc-820b-0c7cc0e9feba/image.png?t=1780600453"/></div><p class="paragraph" style="text-align:justify;">OpenProse is packaged as an <i><a class="link" href="https://github.com/openprose/prose/tree/main/skills/open-prose?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">agent skill</a></i>, which makes it easy to get started with:</p><p class="paragraph" style="text-align:left;"><code>npx skills add openprose/prose</code></p><p class="paragraph" style="text-align:justify;">At the same time, it is much more comprehensive than most skills, as it is forcing the agent to “become the virtual machine”. In other words, the OpenProse skill is incepting your coding agent into behaving like a compiler.</p><div class="image"><img alt="OpenProse skill is incepting your coding agent into behaving like a compiler" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/438e7d3f-2c6c-436d-940a-75fec5d93ef3/image.png?t=1780600453"/></div><p class="paragraph" style="text-align:justify;">At the end of this article we’ll go much deeper under the hood of OpenProse, and you can skip ahead to the end if you really feel the itch, but for the purposes of understanding how we can use OpenProse to create reliable agentic workflows, the key idea is this:</p><p class="paragraph" style="text-align:justify;"><b>OpenProse is logical natural language that both humans and large language models can understand – a shared contract to express and execute what needs to get done and how to do it.</b></p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="logical-english-is-enough-to-get-st">Logical English is enough to get started</h2><p class="paragraph" style="text-align:justify;">Don’t get intimidated by “the language” aspect of OpenProse. All you need to get started is to be able to express your workflow in logical English. The <code>prose write</code> command will handle the rest.</p><div class="image"><img alt="prose programs" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/254e2149-2ccd-4c61-b5cc-61f959c17705/image.png?t=1780600453"/></div><p class="paragraph" style="text-align:justify;">The <code>prose write</code> function takes your logical English as an argument. It outputs a <code>.prose.md</code> file that you can review and edit, which gets handed to your favorite coding agent, who runs the program with <code>prose run</code>.</p><p class="paragraph" style="text-align:justify;">Under the hood, <code>prose write</code> is actually its own <code>.prose.md</code> program, which will develop comprehensive OpenProse programs and test that they are valid before finishing. This can be a really great way to get a feel for the syntax and learn what is possible. In the same way that I would advise you to read your Claude Code or Codex “plans” before having a team of sub-agents race to implement them, it is probably a good idea to read the output <code>&lt;name&gt;.prose.md</code> file before running it, even if all of the details don’t make sense to you.</p><p class="paragraph" style="text-align:justify;">The other nice thing about the fact that OpenProse is “compiled” inside the LLM is that the agent can correct for any technically incorrect syntax. So if you want to start writing your own by hand, you don’t need to get too hung up on the syntactic accuracy to get a sense of how your programs will run.</p><p class="paragraph" style="text-align:justify;"><i>If you want to stop reading and start doing:</i></p><ul><li><p class="paragraph" style="text-align:justify;"><i>Add the skill: npx skills add openprose/prose</i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Star the repo: </i><i><a class="link" href="https://github.com/openprose/prose?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">openprose/prose on GitHub</a></i></p></li><li><p class="paragraph" style="text-align:justify;"><i>Need help? Ask the team at </i><i><a class="link" href="https://OpenProse.ai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">OpenProse.ai</a></i></p></li></ul><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="you-have-hidden-agent-workflows-on-">You have hidden agent workflows on your computer</h2><p class="paragraph" style="text-align:justify;">Do you have some really good Claude Code or Codex sessions? Turn them into reusable and repeatable programs that your agents can run!</p><p class="paragraph" style="text-align:justify;">Some of the most infuriating experiences I’ve had in the past year are when an agent will just absolutely nail something, everything is going perfectly, and then I will go into a new session assuming this very high standard of quality, only to be disappointed and frustrated that I can’t reproduce the magic of that golden session.</p><p class="paragraph" style="text-align:justify;">So I wrote <code>session-to-prose</code>, which turns a Claude Code, Codex, or Pi JSONL session log into a reusable OpenProse <code>*.prose.md</code> program. It does not merely summarize the session. It extracts the reusable workflow: phases, contracts, decision gates, loops, parallel work, strategies, errors, and validation evidence.</p><div class="image"><img alt="Session to prose" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/71299dfe-8f92-40f4-a260-0a92db351792/image__7_.jpg?t=1780603684"/></div><p class="paragraph" style="text-align:justify;">Even if you don’t know it, you are probably sitting on a goldmine of “implicit workflows”, buried in your session logs. The data is already there – OpenProse gives you the format to crystallize it into something you can run again on demand.</p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="explicitly-declare-skills-as-depend">Explicitly declare skills as dependencies</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/80c2a9e7-cc18-42dd-91c1-2b4f9fea7b69/jef.jpg?t=1780603655"/></div><p class="paragraph" style="text-align:justify;">Agent skills are amazing, and many workflows simply do not make sense without them. But at the same time, skills got some things wrong. Geoffrey Huntley argues that the content of skills should be deterministically allocated, not up to the agent to decide. In the standard implementation, skills are surfaced via progressive disclosure, which means the coding agent has to “find” the skill on its own – and it may not. When skills aren’t declared up front, your workflow’s success becomes a coin flip on whether the agent thought to reach for the right one.</p><p class="paragraph" style="text-align:justify;">I was recently using Claude Code to help a friend reformat an academic article from one journal format (which had rejected it) to another journal format for resubmission. Claude Code didn’t do a perfect job of actually reformatting everything, but the thing that blew this person’s mind was that Claude Code made comments and edits credited to Claude Code inside the <code>.docx</code> file itself.</p><p class="paragraph" style="text-align:justify;">Their immediate response was: “Well, how did you do that?”</p><p class="paragraph" style="text-align:justify;">I had to explain that Claude Code has a <code>.docx</code> skill. So this is just a very simple example of a workflow where you’re going to get completely different results with and without that skill. And if one of the stages of your multi-agent workflow is to have an agent edit a word document, then you absolutely need to declare that that skill is loaded before that sub-agent runs.</p><p class="paragraph" style="text-align:justify;">This is also a good place to explicitly call out that OpenProse is a powerful way to coordinate with multiple agents – because the different sub-agents can have different skills. As a fun and arguably extreme example, I made <code>auto-pocock</code>. This is a fully headless OpenProse program that runs a deterministic sequence of <i><a class="link" href="https://x.com/mattpocockuk?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">Matt Pocock</a></i><i>’s </i><i><a class="link" href="https://github.com/mattpocock/skills?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">engineering skills</a></i>, all based from one input: a description of the feature you want built.</p><div class="image"><img alt="a description of the feature you want built." class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/74c6101f-e5cc-4ee7-8cf5-af533407266d/image__8_.jpg?t=1780603744"/></div><p class="paragraph" style="text-align:justify;">How and when we use skills is critical. <i><a class="link" href="https://vercel.com/blog/agents-md-outperforms-skills-in-our-agent-evals?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">Jude Gao at Vercel recently demonstrated</a></i> that a carefully curated skill can actually be worse than a simple index of the documentation. OpenProse can give us the best of both worlds - use the skills that are trusted, with only the subagents that need them, in exactly the part of the workflow where they are called for.</p><p class="paragraph" style="text-align:justify;">Skills are capabilities. Prose programs are contracts for when those capabilities should run.</p><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="under-the-hood-of-open-prose">Under the hood of OpenProse</h2><p class="paragraph" style="text-align:justify;">OK, I promised we’d go deeper, and if you’ve read this far you’ve earned it. This is the part where I try to explain how “a programming language compiled by a coding agent” actually works without hand-waving.</p><p class="paragraph" style="text-align:justify;">The one idea to hold onto is the one I keep circling back to: <i>the coding agent itself is the compiler.</i> There is no OpenProse server sitting between you and Claude Code, intercepting tool calls and orchestrating them from the outside. The agent reads the contract and becomes the virtual machine. Everything else – the file layout, the receipts – is just scaffolding that makes that role visible and keeps it from evaporating when the session closes. OpenProse runs with the agent, not around it.</p><h3 class="heading" style="text-align:justify;" id="the-contract-is-the-program">The contract is the program</h3><p class="paragraph" style="text-align:justify;">A Prose program is a <i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/contract-markdown.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">Markdown file (</a></i><code>*.prose.md</code><i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/contract-markdown.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">)</a></i> that declares a service in logical English. The two sections that matter most are <code>### Requires</code> (the inputs the service needs before it can start) and <code>### Ensures</code> (what must be true when it’s done). Around those you can declare <code>### Services</code> it depends on, <code>### Strategies</code> for how to exercise judgment, and an optional <code>### Execution</code> block when you care about the order things happen in.</p><div class="image"><img alt="The contract is the program" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9a76c461-7e87-4c61-8323-e313b48c2b58/image__9_.jpg?t=1780603808"/></div><p class="paragraph" style="text-align:justify;">If you’ve ever written a function signature and a docstring and wished the agent would just honor them, that’s the feeling. <code>Requires</code> and <code>Ensures</code> are the contract. The agent is on the hook to satisfy them, and – this is the part that matters for trust – it has to leave evidence that it did.</p><h3 class="heading" style="text-align:justify;" id="sessions-subagents-and-the-wall-bet">Sessions, sub-agents, and the wall between them</h3><p class="paragraph" style="text-align:justify;">Every service runs in its own isolated sub-agent session. This is the multi-agent part, and the isolation is the whole point: scratch work, half-formed reasoning, and dead-end files stay inside that session. The only thing that crosses back out is what you declared in <code>### Ensures</code>, copied across an explicit binding. So a sub-agent can make a mess in private, and the workflow only inherits the clean, named artifact it promised. That boundary is what keeps a five-stage program from turning into one giant polluted context window.</p><h3 class="heading" style="text-align:justify;" id="when-you-need-real-control-prose-sc">When you need real control: ProseScript and real tools</h3><p class="paragraph" style="text-align:justify;">Declarative contracts get you surprisingly far, but sometimes you need to say “do this, then that, retry three times, and loop until the reviewer approves.” That’s what <i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/prosescript.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">ProseScript</a></i> is – an imperative layer inside the <code>### Execution</code> block for explicit ordering, conditionals, loops, and retries.</p><p class="paragraph" style="text-align:justify;">And when you need genuinely deterministic behavior, you don’t ask the model to be deterministic – you declare a real tool. A <code>### Tools</code><i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/contract-markdown.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes#tools" target="_blank" rel="noopener noreferrer nofollow"> block</a></i><i> </i>lets a service depend on an actual executable on your <code>PATH</code> (<code>cli:jq</code>, a script of your own) or an MCP server. I tried this: a service that declared <code>cli:jq</code>, run against a malformed JSON blob, really did shell out to jq – a genuine exit-5 parse error, caught at the right column. The determinism came from jq, not from the model imitating jq. But here is the seam, and it matters: having a real tool wired in does not guarantee the agent calls it at the right moment. The determinism lives in the script; the decision to invoke it does not. I’ll come back to this in the limitations, because it’s the single most important caveat in the whole piece.</p><h3 class="heading" style="text-align:justify;" id="receipts-and-state-where-the-trust-">Receipts and state: where the trust actually comes from</h3><p class="paragraph" style="text-align:justify;">This is the part that, for me, is the actual answer to the babysitting problem.</p><p class="paragraph" style="text-align:justify;">The design is that every run leaves a receipt under<i> </i><code>runs/&#123;run-id&#125;/</code><i> </i>– the inputs, the outputs, the logs, the artifacts each service produced. An audit trail, so that when the agent claims it’s done I don’t have to take its word for it; I can read what it actually did. “Done” stops being a vibe and starts becoming something you can inspect.</p><p class="paragraph" style="text-align:justify;">Longer-lived goals – the <i><a class="link" href="https://github.com/openprose/prose/blob/main/skills/open-prose/responsibility-runtime.md?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">standing responsibilities</a></i> that have to stay true over time, not just answer once – keep their memory in <code>state</code>/ between runs.</p><p class="paragraph" style="text-align:justify;">The full filesystem layout is <code>src</code>/ (what you author), <code>dist</code>/ (the compiled manifest the runtime reads), <code>runs</code>/ (receipts), <code>state</code>/ (durable cross-run memory), <code>deps</code>/ (pinned dependencies), and a <code>prose.lock</code>. If that looks suspiciously like a normal software project – source, build output, lockfile – that’s the point. It’s meant to live in git and be reviewed like anything else you’d review.</p><h3 class="heading" style="text-align:justify;" id="why-it-can-run-anywhere">Why it can run anywhere</h3><p class="paragraph" style="text-align:justify;">Because the agent is the compiler, the same <code>.prose.md</code> source runs on any harness that can play the part – Claude Code, Codex, OpenCode, Hermes, Pi (with sub-agents extension), whatever you trust. OpenProse calls this being “Prose Complete”, most coding agents that we’ve tested are “complete enough” to be useful, as long as they have a filesystem, shell and sub-agents. The upshot: your workflows get better as the models get better, without you rewriting anything. You wrote down the contract once; every future model that can satisfy it gets to.</p><h3 class="heading" style="text-align:justify;" id="how-it-compares">How it compares</h3><ul><li><p class="paragraph" style="text-align:justify;"><b>vs. DSPy.</b> In many ways, it shares similar goals to DSPy: create a layer of abstraction that enables you to author programs instead of writing prompts. However, the implementation couldn’t be more opposite. Where DSPy erects extremely strict scaffolding around the LLM calls, <i>OpenProse asserts the entire language contract inside of the LLM.</i></p></li><li><p class="paragraph" style="text-align:justify;"><b>vs. LangChain and CrewAI.</b> I’ll just say it plainly: I am personally allergic to agent frameworks. Before it was agents it was RAG, and the pattern is always the same – you go deep enough to actually understand the framework, you hit the one thing it won’t do, and now you’re either monkeypatching it into a fork or rolling your own anyway. As a serious AI engineer I need control over the context and all the execution details, and generic frameworks abstract exactly those away. The reason OpenProse doesn’t trip my allergy is that it isn’t asking me to move<i> into </i>anything. The work still runs through the coding agent I already use; OpenProse just puts a contract around the workflow. That’s a different category from “adopt our orchestration layer and live there.”</p></li></ul><h3 class="heading" style="text-align:justify;" id="caveats-and-limitations">Caveats and limitations</h3><p class="paragraph" style="text-align:justify;">I’d rather you trust the honest version of this than the hype version, so here are the real edges:</p><ul><li><p class="paragraph" style="text-align:justify;"><b>The LLM is still non-deterministic.</b> This is the big one. OpenProse can make a workflow far more explicit, inspectable, and repeatable – but it does not turn a language model into deterministic infrastructure. If you have genuinely mission-critical code that must run the same way every time, that code belongs in ordinary scripts and tests, orchestrated and verified outside OpenProse. You can declare a real tool and call out to tested, deterministic code, but as I said above, you are then trusting the agent to invoke it at the right point. The power of OpenProse – the thing that lets it not be a framework – comes precisely from the fact that the coding agent itself is the compiler. That is the magic and the trade-off in the same sentence.</p></li><li><p class="paragraph" style="text-align:justify;"><b>A contract only encodes the judgment you put into it</b>. A bad Prose program will faithfully and repeatably do the wrong thing. The operator still has to know what good looks like.</p></li><li><p class="paragraph" style="text-align:justify;"><b>There is overhead.</b> Not every prompt deserves to become a program. Reserve this for the workflows that are actually worth making repeatable.</p></li><li><p class="paragraph" style="text-align:justify;"><b>The host matters, especially the model.</b> What can actually run depends on the affordances of your coding agent – filesystem access, the ability to spawn isolated subagent sessions, environment-variable handling. Most importantly - only the best frontier models can currently do a good job running prose programs. That will shift eventually, but as of today we recommend using the latest and greatest models and agent harnesses to run OpenProse.</p></li></ul><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="now-author-your-first-outcome">Now Author Your First Outcome</h2><p class="paragraph" style="text-align:justify;">As I shared before, the bottleneck is not intelligence. The bottleneck is trust and reliability. </p><p class="paragraph" style="text-align:justify;">I am still babysitting my agents. But I am also slowly starting to write the next page for how I want to collaborate with agents, not in prompts, but in prose.</p><p class="paragraph" style="text-align:justify;">I hope that this gives you a sense of what is possible with OpenProse! Here are some very easy ways to get started:</p><ul><li><p class="paragraph" style="text-align:justify;">Add the skill: <i>npx skills add openprose/prose</i></p></li><li><p class="paragraph" style="text-align:justify;">Star the repo: <i><a class="link" href="https://github.com/openprose/prose?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">openprose/prose on GitHub</a></i></p></li><li><p class="paragraph" style="text-align:justify;">Need help deploying agents? Ask the team at <i><a class="link" href="https://OpenProse.ai?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=stop-babysitting-agents-start-authoring-outcomes" target="_blank" rel="noopener noreferrer nofollow">OpenProse.ai</a></i></p></li></ul><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>This guest post was written by Raymond Weitekamp. We thank OpenProse for supporting Turing Post’s mission to bring clarity to the AI landscape. We encourage you to try it – it’s open source.</i></p><hr class="content_break"><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;" id="faq">FAQ</h2><h3 class="heading" style="text-align:left;">What is OpenProse?</h3><p class="paragraph" style="text-align:left;">OpenProse is a natural-language programming system for AI agent workflows. It lets developers describe multi-step work in logical English and turn it into reusable <code>.prose.md</code> programs that coding agents can run.</p><h3 class="heading" style="text-align:left;">Is OpenProse an agent framework?</h3><p class="paragraph" style="text-align:left;">No. OpenProse is not an agent framework or harness. It does not replace Claude Code, Codex, or other coding agents. Instead, it gives those agents a structured contract for what work should be done and how success should be verified.</p><h3 class="heading" style="text-align:left;">How does OpenProse work?</h3><p class="paragraph" style="text-align:left;">OpenProse uses prose programs written in Markdown. These programs define requirements, expected outcomes, services, tools, strategies, and execution steps. The coding agent reads the program and acts as the “compiler” or virtual machine that executes the workflow.</p><h3 class="heading" style="text-align:left;">Why use OpenProse instead of prompts?</h3><p class="paragraph" style="text-align:left;">Prompts are usually temporary and hard to reproduce. OpenProse makes workflows reusable, reviewable, and versionable. A good agent session can become a durable workflow instead of disappearing into chat history.</p><h3 class="heading" style="text-align:left;">Does OpenProse make AI agents deterministic?</h3><p class="paragraph" style="text-align:left;">No. OpenProse makes agent workflows more explicit and inspectable, but the LLM is still non-deterministic. For mission-critical deterministic behavior, use ordinary scripts, tests, and tools, then let OpenProse orchestrate and verify them.</p></div></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>FOD#154: Enterprise AI Middlemen: Who Survives the Agent Era?</title>
  <description>Reporting from Snowflake Summit and Microsoft Build</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d886ad4c-9c05-4948-b517-fa3781be1c14/Frame_349.jpg" length="82312" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/fod-154-enterprise-ai-middlemen-who-survives-the-agent-era</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/fod-154-enterprise-ai-middlemen-who-survives-the-agent-era</guid>
  <pubDate>Wed, 03 Jun 2026 01:55:52 +0000</pubDate>
  <atom:published>2026-06-03T01:55:52Z</atom:published>
    <dc:creator>Ksenia Se</dc:creator>
    <category><![CDATA[&quot;Froth On The Daydream&quot;]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p id="this-week-in-turing-post" class="paragraph" style="text-align:justify;"><b>Excerpt from today’s editorial:</b> AI may eventually reduce the need for bulky enterprise software middlemen. But before that happens, Snowflake, Microsoft, Databricks, and others are fighting to become the trusted layer between raw AI capability and real company work.</p><hr class="content_break"><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h3 class="heading" style="text-align:justify;">🎫 From our partners: Explore AI & Cloud Innovation at AWS Summit NYC</h3><div class="image"><a class="image__link" href="https://aws.amazon.com/events/summits/new-york/?trk=e5bf269d-4fcb-4a05-9adf-347ed990bc35&sc_channel=el&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" rel="noopener" target="_blank"><img alt="AWS Summit in New York June 17, 2026" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d3cfadb2-b82a-48eb-a514-f1bef8f8e059/Screenshot_2026-06-01_at_2.12.41_PM.jpg?t=1780337589"/></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://aws.amazon.com/events/summits/new-york/?trk=e5bf269d-4fcb-4a05-9adf-347ed990bc35&sc_channel=el&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Join AWS Summit NYC on June 17</a> to explore the latest in AI, cloud infrastructure, and modernization. At this free event, <b>you’ll gain practical insights from leading voices </b>like Dr. Swami Sivasubramanian, and <b>gain access to over 200+ expert-led sessions</b> through hands-on workshops and live demos. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://aws.amazon.com/events/summits/new-york/?trk=e5bf269d-4fcb-4a05-9adf-347ed990bc35&sc_channel=el&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era"><span class="button__text" style=""> Register NOW to secure your spot </span></a></div></div><hr class="content_break"><p class="paragraph" style="text-align:justify;"><b>→ This week was heavy with announcements.</b> I was invited to two conferences at once – <a class="link" href="https://www.snowflake.com/en/summit/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Snowflake Summit</a> and <a class="link" href="https://commandline.microsoft.com/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Microsoft Build</a> in San Francisco. And I went to both conferences because enterprise AI is where the industry is trying to make AI useful at scale. And the sentiment you read in a room full of analysts, while watching leaders push their agenda, defend their choices, and occasionally look uncomfortable, gives you much more than a machine-gun burst of news.</p><p class="paragraph" style="text-align:justify;">Let’s see what’s going on.</p><p class="paragraph" style="text-align:justify;">For two years we watched the AI story through models, benchmarks, chatbots, coding assistants, and consumer adoption. <span style="background-color:rgb(255, 255, 255);">But the harder economic question is now moving inside companies. </span><span style="background-color:rgb(255, 255, 255);"><b>What happens when AI becomes easy enough for people to use directly? What if we don’t need a bulky and heavy middleman?</b></span></p><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">These were the questions I had in mind when I went to Snowflake Summit.</span> Snowflake is a data-governance giant from the cloud-warehouse era, and now it has to prove it stays essential in the agentic one. The market liked the story – shares jumped more than 33% after earnings, FY2027 product-revenue guidance went up to $5.84B from $5.66B, the company signed a five-year $6B AWS deal, and market cap sits around $90B – but <b>the strategic pressure is obvious.</b> Databricks pushes from one side (some calls them “a bully in the room”), Microsoft tries to own the agentic interface from another, hyperscalers sit underneath everyone, and OpenAI and Anthropic, partners and customers on paper, are turning into the real competitors.</p><p class="paragraph" style="text-align:justify;">So I wanted the survival techniques. Does AI hurt enterprise software by making products replaceable, or does it make that software more valuable, because every agent still needs governed data, permissions, identity, memory, and trusted context? The usual debate stops at that binary, but I think that binary is the wrong place to stop.</p><p class="paragraph" style="text-align:justify;">Microsoft Build made the question larger. Satya Nadella’s message from the stage was that agents are everywhere – inside work, development, the operating system, local devices, cloud, enterprise knowledge. But the most useful thinking line came from a smaller Kevin Scott (Microsoft CTO) talk at a private event around Build:</p><ul><li><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">capability is moving faster than deployment</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">activity doesn’t convert into value directly</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">models can do more than organizations can absorb</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">software can accelerate faster than companies can reorganize themselves</span></p></li><li><p class="paragraph" style="text-align:justify;"><span style="background-color:rgb(255, 255, 255);">autonomy doesn’t equal trust</span></p></li></ul><p class="paragraph" style="text-align:justify;">Yup, I thought, exactly that. But these are symptoms.</p><p class="paragraph" style="text-align:justify;">Here is what the lag actually creates. Enterprises are full of legacy systems, fragile pipelines, compliance constraints, security reviews, human habits, and workflows that only make sense because someone has been repairing them by hand for years. Agents get more capable by the month; organizations do not move at model speed. So we get a strange interval, where AI threatens intermediaries in the long run while enterprise complexity keeps them useful in the short one. Snowflake, Microsoft, Databricks, Salesforce, ServiceNow are all fighting inside it, and from what I heard in the hallways, there is real uncertainty in there, almost panic, and a lot of running in too many directions at once. I also feel hey might be brittle because of their own internal complexity. Agents might need different organisation.</p><p class="paragraph" style="text-align:justify;">The comforting reading of that interval – agents threaten us later, complexity protects us now – treats the moat as a function of time. I think it&#39;s a function of cost. As per-token cost falls, the raw work middleware used to charge for, querying and generating and connecting and transforming, drifts toward free. What stays expensive is trust: knowing an agent acted on the right data, with the right permissions, on behalf of the right person, with a record of why. When everything cheap gets cheaper, the scarce thing is governed proximity to intent.</p><p class="paragraph" style="text-align:justify;">Which means the middleman doesn&#39;t disappear in the direct-use world; it changes shape. It stops being an application – capability packaged into rigid software and sold back through layers of interface – and becomes a substrate, the trust and permission layer that sits at the moment a stated intention turns into a permitted action. The agent era doesn&#39;t need less governance. It needs more, located somewhere new.</p><p class="paragraph" style="text-align:left;">But it’s a different type of governance.</p><p class="paragraph" style="text-align:left;">This is the test Snowflake should be judged by. Not whether it governs data – it does, and that was the moat of the last era – but how it governs it, and whether it moves up to the intent layer or stays a managed warehouse with AI bolted on top; whether it reduces the distance between what a person wants and useful work getting done, or adds one more governed layer between the user and the outcome.</p><p class="paragraph" style="text-align:left;">And tokenmaxxing might be even dangerous here. Independently from each other, Sridhar Ramaswamy, Snowflake&#39;s gentle CEO, and Kevin Scott were making the same point: activity, or tokenmaxxing, is not value, and more agents, more context, more tokens, more integrations, more automation can all produce motion without progress. I think this is the dilemma we will be sorting through this year, both at the behemoth and startup levels.</p><p class="paragraph" style="text-align:left;">The middleman doesn&#39;t die in this story. It moves – closer to where intention becomes work, and further from the warehouse.</p><p class="paragraph" style="text-align:justify;"><i><b>If any of those thoughts resonate with you – share them across your social networks. Let’s keep the conversation going. </b></i></p><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>Topic 2: </i>Also, NVIDIA made tons of announcements in Taipei (see <i>News from the Usual Suspects</i> below). We played with and covered one of the most interesting: <b>Cosmos 3, its omnimodal world model family.</b></p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/8gXNaDw5zmk" width="100%"></iframe><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="{{live_url}}?comments=true"><span class="button__text" style=""> Leave a Comment </span></a></div><div class="section" style="background-color:transparent;border-color:#df12a9;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"><i><b>Follow us on </b></i><i> </i>🎥<i><a class="link" href="https://www.youtube.com/@RealTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"> YouTube</a></i><i> </i><i><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Twitter</a></i><i> </i><i><a class="link" href="https://huggingface.co/Kseniase?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"> Hugging Face </a></i>🤗</p></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="twitter-library">Twitter Library </h2><div class="embed"><a class="embed__url" href="https://www.turingpost.com/p/ragtypes?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank"><div class="embed__content"><p class="embed__title"> 20 Advanced RAG Types to Know in 2026 </p><p class="embed__description"> 20 cutting-edge RAG approaches in 2026: Agentic RAG, MiA-RAG, HGMem, Graph-O1, Bidirectional RAG, multimodal, multilingual, structured and security RAG systems. </p><p class="embed__link"> Turing Post • Alyona Vert. </p></div><img class="embed__image embed__image--right" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/52e33110-4dd6-439b-99b4-be2015f207f3/silverpoint_converted_uploaded_image.jpg?t=1779451774"/></a></div><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="news-from-the-usual-suspects">News from the usual suspects ™</h2><ul><li><p class="paragraph" style="text-align:justify;">Snowflake raised guidance, signed a <a class="link" href="https://www.snowflake.com/en/news/press-releases/snowflake-expands-aws-collaboration-with-6b-commitment-to-accelerate-enterprise-agentic-ai-adoption/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>$6B AWS agreement</b></a><a class="link" href="https://www.snowflake.com/en/news/press-releases/snowflake-expands-aws-collaboration-with-6b-commitment-to-accelerate-enterprise-agentic-ai-adoption/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">,</a> expanded its partnership with <a class="link" href="https://www.turingpost.com/p/anthropic?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Anthropic</b></a>, and introduced renamed <a class="link" href="https://www.snowflake.com/en/news/press-releases/snowflake-coco-redefines-enterprise-ai-development-as-the-coding-agent-built-for-faster-easier-and-more-powerful-innovation-anywhere/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>CoCo</b></a> and <a class="link" href="https://www.snowflake.com/en/news/press-releases/snowflake-cowork-powers-the-agentic-enterprise-as-the-personal-agent-for-knowledge-workers-to-work-smarter/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Snowflake Intelligence / CoWork</b></a>. Together, they show Snowflake trying to move from storing data to helping people and agents act on it. My question: do companies need this new layer, or is it another unnecessary AI coding interface.</p></li></ul><div class="image"><img alt="Snowflake Data Cloud Ecosystem Map" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d624cb0c-d3d3-4e87-b1cc-e6236fcb47b5/Gemini_Generated_Image_7lq34x7lq34x7lq3.jpg?t=1780464608"/></div><ul><li><p class="paragraph" style="text-align:justify;"><b>Microsoft</b> used Build to unveil <a class="link" href="https://commandline.microsoft.com/project-solara-build-2026/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Project Solara</b></a><b> </b>(very interesting but still mostly under development), <a class="link" href="https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/02/introducing-microsoft-scout-your-always-on-personal-agent/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Scout</b></a><b> </b>(basically your chief of stuff), and the <a class="link" href="https://blogs.windows.com/devices/2026/06/02/building-the-next-generation-of-devices-for-developers-surface-rtx-spark-dev-box/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Surface RTX Spark Dev Box</b></a>, signaling a future where agents run directly on Windows devices rather than exclusively in the cloud. They also introduced an large family of MAI models, including the first reasoning one.</p></li></ul><div class="image"><img alt="Microsoft Build 2026 - Announcement Map" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0405bc5b-3d1e-46e1-95a6-4d8feccab0f4/ChatGPT_Image_Jun_2__2026__09_21_53_PM.jpg?t=1780464733"/></div><ul><li><p class="paragraph" style="text-align:left;"><b>Anthropic</b> filed confidential IPO paperwork, <a class="link" href="https://www.anthropic.com/news/series-h?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">raised </a><b><a class="link" href="https://www.anthropic.com/news/series-h?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">$65B</a></b><a class="link" href="https://www.anthropic.com/news/series-h?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"> at a </a><b><a class="link" href="https://www.anthropic.com/news/series-h?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">$965B valuation</a></b>, and expanded its security-focused <a class="link" href="https://www.anthropic.com/news/expanding-project-glasswing?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"><b>Project Glasswing</b></a>, which has already helped find more than <b>10,000 critical vulnerabilities</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>OpenAI</b> claimed one of its models <a class="link" href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">independently disproved a longstanding geometry conjecture</a>, rolled out more <a class="link" href="https://openai.com/business/frontier/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">enterprise agent deployments</a> through Codex, and published a<a class="link" href="https://openai.com/index/openai-frontier-governance-framework/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow"> new Frontier Governance Framework </a>as AI regulation starts taking shape.</p></li><li><p class="paragraph" style="text-align:left;"><b>NVIDIA</b> <a class="link" href="https://www.nvidia.com/en-us/geforce/news/computex-2026-nvidia-geforce-rtx-announcements/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">dominated Computex</a> with a vision that goes far beyond GPUs:</p><ul><li><p class="paragraph" style="text-align:left;"><b>RTX Spark</b> brings agentic AI to Windows PCs with up to <b>1 PFLOP</b> of AI performance and <b>128GB unified memory</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>DGX Station</b> puts a trillion-parameter AI supercomputer on an engineer&#39;s desk.</p></li><li><p class="paragraph" style="text-align:left;">New <b>humanoid robot blueprint</b>, robotaxi, AI factory, and digital twin announcements reinforced Jensen Huang&#39;s central thesis: AI needs infrastructure, not just models.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Meta</b> <a class="link" href="https://www.reuters.com/world/meta-scales-back-ai-mouse-clicks-tool-citing-employee-concerns-2026-06-02/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">scaled back</a> plans to collect employee mouse and keyboard activity for AI training after internal backlash, highlighting a growing tension between agent development and workplace privacy.</p></li><li><p class="paragraph" style="text-align:left;"><b>The U.S. government</b> <a class="link" href="https://thehill.com/policy/technology/5905712-trump-executive-order-ai-model-testing/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">announced </a>plans to invite frontier AI labs to voluntarily submit advanced models for cybersecurity testing before public release.</p></li><li><p class="paragraph" style="text-align:left;"><b>Apple</b> remains the biggest question mark ahead of WWDC next week, where expectations are building around the next phase of Apple Intelligence and Siri.</p></li></ul><h2 class="heading" style="text-align:justify;" id="research-highlight">Research highlight</h2><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/TheTuringPost/status/2060153308392857933?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era"><p> Twitter tweet </p></a></blockquote><h2 class="heading" style="text-align:justify;" id="models">Models </h2><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://x.ai/news/grok-build-0-1?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">xAI Grok Build 0.1</a> (released May 20–29, 2026): Fastest coding-focused model from xAI, optimized for agentic workflows, multi-file edits, tool use, and terminal-native development. 256K context, supports text + image inputs. Now public beta on xAI API.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.anthropic.com/news/claude-opus-4-8?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Anthropic Claude Opus 4.8</a> (late May 2026 rollout): Major upgrade to the Opus line focused on coding accuracy, long-running task reliability, objective progress reporting, and reduced defect rates in complex reasoning/workflows. Same pricing tier as 4.7.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://x.com/MiniMax_AI/status/2061266317815296322?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">MiniMax M3</a> (announced May 31, 2026): First open-weights model claiming frontier-level coding & agentic performance (59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas). Features MiniMax Sparse Attention for up to 1M token context, native multimodality from day one (text + vision). API live now with 50% off promo for first 7 days (≤512K context). Full weights + tech report expected in ~10 days.</p></li><li><p class="paragraph" style="text-align:justify;"><b>Microsoft MAI family</b> (announced on June 2):</p><ul><li><p class="paragraph" style="text-align:justify;"><b>MAI-Thinking-1</b>: Microsoft’s first reasoning model. 35B active parameters, 128K context. Strong on complex multi-step instructions, long-context reasoning, and code generation. Matches Opus 4.6 on SWE-Bench Pro; preferred over Sonnet 4.61 in blind human tests. Low token cost, high efficiency. Private preview on Foundry.</p></li><li><p class="paragraph" style="text-align:justify;"><b>MAI-Image-2.5</b> (and flash variant): Microsoft’s first native text-to-image and image-to-image models. Surpasses Nano Banana Pro on ELO. Rolling out in PowerPoint, OneDrive, and Foundry.</p></li><li><p class="paragraph" style="text-align:justify;"><b>MAI-Transcribe-1.5</b>: SOTA accuracy across 43 languages (streaming soon).</p></li><li><p class="paragraph" style="text-align:justify;"><b>MAI-Voice-2</b> (and flash variant): Expanded to 15+ additional languages with new voice options.</p></li><li><p class="paragraph" style="text-align:justify;"><b>MAI-Code-1</b>: Ultra-efficient inference coding model, tuned for GitHub; now in Copilot and VS Code.</p></li></ul></li><li><p class="paragraph" style="text-align:justify;"><b>NVIDIA </b></p><ul><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://www.youtube.com/watch?v=8gXNaDw5zmk&t=2s&utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Cosmos 3 – a unique omnimodal world model – that we covered here </a></p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://developer.nvidia.com/nemotron?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Nemotron 3 Ultra </a>(available starting June 4, 2026): NVIDIA’s largest and best open model to date, specifically designed for long-running, complex autonomous agent workloads (deep reasoning, code generation, extensive research). Claims 5× faster inference and up to 30% lower cost for agentic tasks vs. other open models in class.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://nvidianews.nvidia.com/news/nvidia-alpamayo-2-super-robotaxis?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Alpamayo 2 Super</a>: Open 32B-parameter reasoning VLA (vision-language-action) model focused on the full driving stack for safer Level 4 autonomous development. Emphasizes reasoning, planning, and acting.</p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://research.nvidia.com/labs/sil/projects/gamma-world/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Gamma-World (γ-World)</a> (announced May 27, 2026): It scales interactive simulation beyond single- or two-player settings. Introduces Simplex Rotary Agent Encoding and Sparse Hub Attention for independently controllable, permutation-symmetric agents. Enables real-time (24 FPS) coherent multi-agent video rollouts with strong zero-shot generalization to more agents.</p></li></ul></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://x.com/Alibaba_Qwen/status/2061506641120641494?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Qwen3.7-Plus</a> (announced June 1, 2026 by Alibaba): Multimodal agent foundation model that unifies vision + language into a single versatile agent.</p></li></ul><h2 class="heading" style="text-align:justify;" id="research">Research </h2><p class="paragraph" style="text-align:justify;">Trends we see looking at every paper related to AI and ML published last week: </p><h3 class="heading" style="text-align:justify;" id="ai-research-agents-and-deep-researc"><b>AI research agents and deep research</b></h3><ul><li><p class="paragraph" style="text-align:justify;"><b>AutoResearchClaw</b> – Moves autonomous research from a linear pipeline toward iterative systems with debate, self-healing execution, human intervention, and cross-run learning <a class="link" href="https://arxiv.org/abs/2605.20025?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">→read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>QUEST</b> – Trains open deep-research agents with fully synthetic tasks, making synthetic task generation look like a serious path for scaling research agents →<a class="link" href="https://arxiv.org/abs/2605.24218?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>ScientistOne</b> – Pushes autonomous research toward verifiability by requiring claims to trace back to evidence, which directly addresses hallucinated citations, unreproducible scores, and method-code mismatch →<a class="link" href="https://arxiv.org/abs/2605.26340?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="agent-operating-systems-harnesses-s"><b>Agent operating systems: harnesses, skills, and coordination</b></h3><ul><li><p class="paragraph" style="text-align:justify;">🌟 <b>SkillOpt</b> – Treats agent skills as trainable external state, with controlled edits, validation gates, and transfer across models and execution harnesses →<a class="link" href="https://arxiv.org/abs/2605.23904?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>MUSE-Autoskill</b> – Builds a full lifecycle for agent skills: creation, memory, management, evaluation, refinement, reuse, and cross-agent transfer →<a class="link" href="https://arxiv.org/abs/2605.27366?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Foundation Protocol</b> – Proposes coordination infrastructure for large populations of agents, with identity, provenance, multi-party organization, metering, receipts, settlement, and governance as first-class concerns →<a class="link" href="https://arxiv.org/abs/2605.23218?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="memory-context-retrieval-and-persis"><b>Memory, context, retrieval, and persistence</b></h3><ul><li><p class="paragraph" style="text-align:justify;"><b>ACC</b> – Converts agent trajectories into long-context training data, using tool responses and environment observations as supervision for distant evidence integration →<a class="link" href="https://arxiv.org/abs/2605.21850?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>WorldKV</b> – Makes persistent world memory cheaper for interactive video world models by retrieving and compressing KV-cache chunks instead of keeping everything in full attention →<a class="link" href="https://arxiv.org/abs/2605.22718?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>Do Language Models Need Sleep?</b> – Introduces offline recurrence as a consolidation phase, giving models a way to process long-horizon memory outside the active inference window →<a class="link" href="https://arxiv.org/abs/2605.26099?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>OmniRetrieval</b> – Unifies retrieval across text, relational tables, knowledge graphs, and property graphs without flattening every source into one generic representation →<a class="link" href="https://arxiv.org/abs/2605.29250?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Personalize-then-Store</b> – Moves agent memory toward user-specific storage policies, asking which interactions are worth saving for each person rather than applying one universal rule →<a class="link" href="https://arxiv.org/abs/2605.25535?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="training-for-reasoning-diversity-an"><b>Training for reasoning, diversity, and test-time search</b></h3><ul><li><p class="paragraph" style="text-align:justify;"><b>DelTA</b> – Improves RLVR by assigning token-level credit more discriminatively, so training can amplify the tokens that actually separate successful from failed reasoning →<a class="link" href="https://arxiv.org/abs/2605.21467?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟 <b>Vector Policy Optimization</b> – Trains models to produce diverse solution sets for test-time search, instead of collapsing toward similar answers optimized for one scalar reward →<a class="link" href="https://arxiv.org/abs/2605.22817?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="architectures-world-models-and-phys"><b>Architectures, world models, and physical AI</b></h3><ul><li><p class="paragraph" style="text-align:justify;"><b>HRM-Text</b> – Tests whether hierarchical recurrent architectures can deliver stronger sample efficiency than standard Transformer scaling recipes →<a class="link" href="https://arxiv.org/abs/2605.20613?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;">🌟<b>Gamma-World</b> – Extends world models from single-agent or two-player settings toward scalable multi-agent interactive simulation →<a class="link" href="https://arxiv.org/abs/2605.28816?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>Qwen-VLA</b> – Unifies manipulation, navigation, and trajectory prediction inside one vision-language-action model across tasks, environments, and robot embodiments →<a class="link" href="https://arxiv.org/abs/2605.30280?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="verifiable-environments-for-compute"><b>Verifiable environments for computer-use agents</b></h3><ul><li><p class="paragraph" style="text-align:justify;"><b>OpenComputer</b> – Creates verifiable software worlds for desktop agents, using app-specific state verifiers, auditable trajectories, and partial-credit rewards →<a class="link" href="https://arxiv.org/abs/2605.19769?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>MobileGym</b> – Provides a highly parallel, verifiable mobile-GUI simulation platform, making smartphone-agent training more scalable and measurable →<a class="link" href="https://arxiv.org/abs/2605.26114?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li><li><p class="paragraph" style="text-align:justify;"><b>CUA-Gym</b> – Scales verifiable RL environments for computer-use agents by generating task instructions, environment states, and reward functions together →<a class="link" href="https://arxiv.org/abs/2605.25624?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">read the paper</a></p></li></ul><p class="paragraph" style="text-align:justify;"></p><p class="paragraph" style="text-align:justify;"><i>That’s all for today. Thank you for reading! Please </i><i><b>send this newsletter to colleagues</b></i><i> if it can help them enhance their understanding of AI and stay ahead of the curve.</i></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">FAQ</h2><h3 class="heading" style="text-align:justify;">What is the “middleman interval” in enterprise AI?</h3><p class="paragraph" style="text-align:justify;">The middleman interval is the current phase in which AI is powerful enough to threaten some software intermediaries, but enterprise complexity still makes trusted platforms valuable. Companies need governance, permissions, security, identity, and reliable data access before agents can safely perform real work.</p><h3 class="heading" style="text-align:justify;">Will AI hurt SaaS companies?</h3><p class="paragraph" style="text-align:justify;">AI may pressure SaaS companies that mainly package narrow workflows behind rigid interfaces. But it can also strengthen companies that become trusted layers for agentic work: managing data, identity, security, governance, workflow execution, and observability.</p><h3 class="heading" style="text-align:justify;">What is tokenmaxxing?</h3><p class="paragraph" style="text-align:justify;">Tokenmaxxing is the habit of using more AI, more context, more agents, and more tokens without proving that the extra activity creates proportional value. In enterprise AI, the backlash against tokenmaxxing is really a demand for measurable useful work.</p></div><p class="paragraph" style="text-align:justify;">⬅️ FOD 153 <a class="link" href="https://www.turingpost.com/p/fod-153-agentic-coding-in-search-what-it-even-means?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Agentic coding in search – What it even means? </a></p><p class="paragraph" style="text-align:justify;">➡️ FOD 155 <a class="link" href="https://www.turingpost.com/p/continual-learning-llms-ai-models-sleep?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=fod-154-enterprise-ai-middlemen-who-survives-the-agent-era" target="_blank" rel="noopener noreferrer nofollow">Continual Learning in LLMs: Why AI Models Need Sleep</a></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>12 AI Co-Scientists of 2026</title>
  <description>5 most notable breakthroughs plus 7 open-source co-scientists for you experiments</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8bdd6186-1876-4298-9723-d47045110c1c/3f681221-0d5c-4ef6-bc49-ae75420788c7.jpg" length="84421" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/ai-co-scientists-in-2026</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/ai-co-scientists-in-2026</guid>
  <pubDate>Sun, 24 May 2026 14:46:23 +0000</pubDate>
  <atom:published>2026-05-24T14:46:23Z</atom:published>
    <dc:creator>Alyona Vert.</dc:creator>
    <category><![CDATA[Twitter Library]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;">There’s probably nothing more inspiring than using AI in the field that brings the greatest benefit to humanity and the world as a whole ‒ <b>science and research</b>. And right now, we’re seeing a sharp leap in the rise of new <b>AI systems acting as co-scientists</b>.</p><p class="paragraph" style="text-align:justify;">No, this doesn’t reduce the value of scientists, engineers, or researchers ‒ these systems help people generate and analyze results much faster, accelerating progress in biology, chemistry, physics, medicine, and beyond. Massive amounts of tedious analysis, hypothesis testing, and experimentation that used to take many years, can now be done in days, sometimes even hours.</p><p class="paragraph" style="text-align:justify;">So today, we’ll look at some of the most interesting AI co-scientist systems.</p><p class="paragraph" style="text-align:justify;"><i><b>TL;DR:</b></i><i> Use DeepMind’s Co-Scientist and Robin for biology and drug discovery, AxiomProver and AI Co-Mathematician for theorem proving and mathematical research, ERA and AI CFD Scientist for autonomous scientific simulations, and systems like The AI Scientist or AutoResearchClaw for fully automated end-to-end research workflows.</i></p><p class="paragraph" style="text-align:justify;">Some of them are biology co-scientists discovering drugs and engineering proteins. Some are AI mathematicians solving proofs and conjectures. Others generate scientific simulations and experiments, and there are also systems that automate the entire research pipeline, even including paper writing. Further, you’ll also find 6 open-source co-scientists.</p><p class="paragraph" style="text-align:left;">/</p><h2 class="heading" style="text-align:justify;" id="the-most-notable-co-scientists">The Most Notable Co-Scientists</h2><h3 class="heading" style="text-align:justify;" id="1-google-deep-minds-co-scientist">1. Google DeepMind’s Co-Scientist</h3><p class="paragraph" style="text-align:justify;">This is one of the defining AI systems of 2026 that<b> can reduce large-scale biological data analysis from months to days</b>. Built on Gemini as a multi-agent research architecture, it generates, debates, ranks, and evolves scientific hypotheses through iterative “idea tournaments” inspired by AlphaGo. The system is already being used for <b>fibrosis, ALS, antimicrobial resistance, cellular aging, infectious diseases, and plant immunity research</b> <span style="background-color:rgb(255, 255, 255);"><span style="color:rgb(108, 108, 108);font-family:Arial, sans-serif;font-size:14px;">‒</span></span> including identifying <b>a fibrosis drug candidate</b> that blocked 91% of a scarring-linked response in lab tests.</p><p class="paragraph" style="text-align:justify;">Co-Scientist is the core research engine behind Google’s broader <b>Gemini for Science initiative</b>: the “Hypothesis Generation” tool is explicitly built on Co-Scientist. Google positions it as part of a larger ecosystem alongside AlphaEvolve (for computational experiments) and NotebookLM-based Literature Insights.</p><ul><li><p class="paragraph" style="text-align:justify;">Co-Scientist: <a class="link" href="https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a> </p></li><li><p class="paragraph" style="text-align:justify;">Gemini for Science: <a class="link" href="https://blog.google/innovation-and-ai/technology/research/gemini-for-science-io-2026/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="2-open-ai-model-and-a-famous-conjec">2. OpenAI model and a famous conjecture in geometry </h3><p class="paragraph" style="text-align:justify;">OpenAI’s new reasoning model (it’s name is not mentioned) has just solved a math problem that had stumped researchers since 1946 ‒ one of Paul Erdős most famous geometry puzzles:<b> if you place </b><i><b>n</b></i><b> points on a plane, how many pairs can sit exactly one unit apart?</b> For nearly 80 years, mathematicians believed square-grid-like patterns were basically optimal.</p><p class="paragraph" style="text-align:justify;">The AI proved that they weren’t. <b>A general-purpose reasoning model (not a math-specialized system) discovered a completely new construction that creates far more unit-distance pairs than anyone expected.</b> It used deep ideas from algebraic number theory that experts never thought connected to this geometry problem. Mathematicians verified the proof and called it a milestone for AI-driven mathematics.</p><ul><li><p class="paragraph" style="text-align:justify;">Open’s <a class="link" href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;">The proof from the model: <a class="link" href="https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-proof.pdf?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research Paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="5-future-houses-robin">3. FutureHouse’s Robin</h3><p class="paragraph" style="text-align:justify;">This multi-agent system autonomously generates and experimentally validates a therapeutic hypothesis for a real lab workflow: reads papers, generates hypotheses, proposes experiments, analyzes lab data, and iterates on results in a loop.</p><p class="paragraph" style="text-align:justify;">Researchers <b>tested it on dry macular degeneration</b> <span style="background-color:rgb(255, 255, 255);"><span style="color:rgb(108, 108, 108);font-family:Arial, sans-serif;font-size:14px;"><b>‒</b></span></span><b> a major cause of blindness.</b> Robin proposed that<b> improving retinal cell “cleanup” (phagocytosis) </b>could help treat the disease, then identified <b>ripasudil</b> (a glaucoma drug) as a candidate. Lab tests confirmed the real effect of ripasudil.</p><p class="paragraph" style="text-align:justify;"><b>Robin also analyzed RNA-seq data itself and discovered a possible new target, ABCA1.</b></p><ul><li><p class="paragraph" style="text-align:justify;">Robin <a class="link" href="https://arxiv.org/abs/2505.13400?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="axiom-prover-by-axiom-math">4. AxiomProver by Axiom Math</h3><p class="paragraph" style="text-align:justify;">AxiomProver is an AI mathematician that <b>writes fully machine-verified mathematical proofs in Lean</b>. In 2026, it <b>autonomously solved all 12 problems of the Putnam exam</b>, one of the hardest undergraduate math competitions in the world, with 8 solved within the official contest time. It handles calculus, combinatorics, geometry, and number theory problems, and <b>sometimes discovered proof strategies humans didn’t expect</b>, including geometric arguments and brute-force formal reasoning.</p><ul><li><p class="paragraph" style="text-align:justify;"> AxiomProver: Axiom Math’s <a class="link" href="https://axiommath.ai/territory/from-seeing-why-to-checking-everything?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="5-ai-comathematician">5. AI co-mathematician</h3><p class="paragraph" style="text-align:justify;">Here is another research partner from Google DeepMind but for mathematicians. It uses multiple AI agents based on Gemini that work in parallel, almost like a research team in a collaborated workspace. One agent might search literature, another explores examples, while others try proving or disproving ideas. Researchers can guide the process interactively.</p><p class="paragraph" style="text-align:justify;">Mathematicians used it to solve open problems, discover new research directions, and find overlooked papers. It also <b>scored a record 48% on one of the hardest math benchmark</b> ‒ FrontierMath Tier 4.</p><ul><li><p class="paragraph" style="text-align:justify;">AI co-mathematician <a class="link" href="https://arxiv.org/pdf/2605.06651?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research paper</a></p></li></ul><hr class="content_break"><p id="watch-this-video-to-get-a-full-pict" class="paragraph" style="text-align:justify;">Watch this video to get a full picture of how AI is transforming science and why it feels so important now →</p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/WzoPLV3pYUs" width="100%"></iframe><hr class="content_break"><h2 class="heading" style="text-align:justify;" id="open-source-co-scientists">Open-Source Co-Scientists</h2><h3 class="heading" style="text-align:justify;" id="1-era-empirical-research-assistance">1. ERA (Empirical Research Assistance)</h3><p class="paragraph" style="text-align:justify;">This AI system by a group of researchers from top universities like MIT, Harvard, McGill University and Google researchers <b>generates scientific simulation code</b> for complex research problems, helping scientists build computational experiments and large-scale simulations much faster. ERA continuously writes, tests, scores, and improves scientific software through trial and error using AlphaZero-style tree search. It can combine ideas from papers, textbooks, and existing methods to invent new hybrid approaches.</p><p class="paragraph" style="text-align:justify;">The most notable results:</p><ul><li><p class="paragraph" style="text-align:justify;"><b>ERA created 40 biology methods </b>that outperformed top human approaches for <b>single-cell data analysis</b>.</p></li><li><p class="paragraph" style="text-align:justify;">It generated <b>14 COVID forecasting models that beat the CDC ensemble</b>.</p></li><li><p class="paragraph" style="text-align:justify;">It also reached expert-level performance in <b>geospatial analysis, neuroscience, and time-series forecasting</b>.</p><p class="paragraph" style="text-align:justify;"></p></li><li><p class="paragraph" style="text-align:justify;">ERA: <a class="link" href="https://arxiv.org/abs/2509.06503?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research paper</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/google-research/era?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">ERA</a><a class="link" href="https://github.com/google-research/era?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow"> GitHub</a></p></li><li><p class="paragraph" style="text-align:justify;">Google Research’s ERA <a class="link" href="https://research.google/blog/empirical-research-assistance-era-from-nature-publication-to-catalyzing-computational-discovery/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;"><a class="link" href="https://google-research.github.io/era/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Experiments</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="8-axiom-maths-axplorer">2. Axiom Math’s Axplorer</h3><p class="paragraph" style="text-align:justify;">This AI system solves extremely hard math optimization problems, where there are trillions of possible solutions. It works in a loop: an AI model learns patterns from good solutions, generates new candidate solutions, and then a local search algorithm fixes and improves them. Over time, the system gets better at proposing promising candidates.</p><p class="paragraph" style="text-align:justify;">The team has already used it on difficult combinatorics problems like <b>square-free graphs and isosceles-free point sets</b>. In one case, <b>Axplorer found an optimal graph solution in 2.5 hours on a single GPU</b>, using about 100× fewer attempts than earlier methods.</p><ul><li><p class="paragraph" style="text-align:justify;">Axplorer: <a class="link" href="https://axiommath.ai/territory/axplorer?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/AxiomMath/axplorer?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Axplorer</a><a class="link" href="https://github.com/AxiomMath/axplorer?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow"> GitHub</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="3-disco-by-future-house">3. DISCO by FutureHouse</h3><p class="paragraph" style="text-align:justify;">This multimodal generative model <b>co-designs entirely new proteins and enzymes from scratch</b>. You give it a target molecule or desired chemistry, and it invents both the protein sequence and the 3D structure simultaneously.</p><p class="paragraph" style="text-align:justify;">DISCO’s concept may sound too similar to DeepMind’s AlphaFold, but AlphaFold’s idea is different: it predicts how an existing protein folds into a 3D structure. AlphaFold is mainly a prediction modeling system, while DISCO is a generative design system.</p><ul><li><p class="paragraph" style="text-align:justify;">DISCO: <a class="link" href="https://www.futurehouse.org/research-announcements/disco-enzymes?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/DISCO-design/DISCO?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">DISCO GitHub</a></p></li><li><p class="paragraph" style="text-align:justify;">DISCO: <a class="link" href="https://arxiv.org/abs/2604.05181?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research Paper</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="4-k-ups-from-cusp-ai">4. kUPS from CuspAI</h3><p class="paragraph" style="text-align:justify;">CuspAI introduced kUPS, an open-source <b>molecular simulation engine for AI-driven chemistry and materials science</b>. kUPS unifies molecular dynamics, Monte Carlo simulations, geometry optimization, and ML-based force fields inside one GPU-native Python/JAX framework. It’s designed around composable “propagators” <span style="background-color:rgb(255, 255, 255);"><span style="color:rgb(108, 108, 108);font-family:Arial, sans-serif;font-size:14px;">‒</span></span> different simulation techniques and ML models can plug together easily. The system also supports batching thousands of simulations in parallel on GPUs. kUPS tool gives <b>49× faster simulations</b> than standard chemistry tools.</p><ul><li><p class="paragraph" style="text-align:justify;">kUPS: <a class="link" href="https://medium.com/@CuspAI/kups-a-molecular-simulation-engine-for-the-ai-era-b213963a2359?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/cusp-ai-oss/kups?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">kUPS GitHub</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/cusp-ai-oss/tojax?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Tojax GitHub</a> ‒ a small library from CuspAI that converts PyTorch models into JAX functions and allows to run them inside JAX-based simulation systems</p></li></ul><h3 class="heading" style="text-align:justify;" id="5-ai-cfd-scientist-with-a-physicsaw">5. AI CFD Scientist with a physics-aware verification loop</h3><p class="paragraph" style="text-align:justify;">This “AI scientist” is built specifically for <b>computational fluid dynamics (CFD)</b> <span style="background-color:rgb(255, 255, 255);"><span style="color:rgb(108, 108, 108);font-family:Arial, sans-serif;font-size:14px;">‒</span></span> the kind of simulations engineers use to study airflow, turbulence, jets, aerodynamics, combustion, and fluid motion. CFD is much harder than normal AI research automation, because a simulation can technically “work” while still being physically wrong. So the biggest breakthrough here is that <b>the system includes a physics-aware verification loop: </b>it renders flow simulations and uses a vision-language model (VLM) to check whether the physics looks realistic. If flow patterns are wrong, the AI rejects the run and fixes or reruns the simulation.</p><p class="paragraph" style="text-align:justify;">After 44 iterative experiments, <b>AI CFD Scientist discovered a new correction for the Spalart–Allmaras turbulence model</b> that improved wall-friction prediction accuracy by 7.89%.</p><ul><li><p class="paragraph" style="text-align:justify;">AI CFD Scientist: <a class="link" href="https://arxiv.org/abs/2605.06607?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research Paper</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/csml-rpi/AI-CFD-Scientist?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">AI CFD Scientist</a><a class="link" href="https://github.com/csml-rpi/AI-CFD-Scientist?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow"> GitHub</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="6-the-ai-scientist-from-sakana-ai">6. The AI Scientist from Sakana AI</h3><p class="paragraph" style="text-align:justify;">We can’t not to mention Sakana AI’s AI Scientist, even though it was developed in 2024. It automates the full machine learning research workflow: generating ideas, searching literature, writing code, running experiments, analyzing results, and drafting complete research papers. One of its biggest milestones was producing <b>the first fully AI-generated paper to pass a real human peer-review process at an ICLR workshop.</b> AI Scientist-v2 also introduced agentic tree search, parallel experiment execution, and VLM-based figure checking to improve experiments and paper quality. The project also showed a clear “scaling law”: stronger foundation models consistently produced higher-quality scientific papers.</p><ul><li><p class="paragraph" style="text-align:justify;">The AI Scientist: <a class="link" href="https://sakana.ai/ai-scientist-nature/?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">blog post</a></p></li><li><p class="paragraph" style="text-align:justify;">The AI Scientist-v2: <a class="link" href="https://arxiv.org/abs/2504.08066?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research Paper</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/SakanaAI/AI-Scientist-v2?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">The AI Scientist</a><a class="link" href="https://github.com/SakanaAI/AI-Scientist-v2?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">-v2</a><a class="link" href="https://github.com/SakanaAI/AI-Scientist-v2?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow"> GitHub</a></p></li></ul><h3 class="heading" style="text-align:justify;" id="7-auto-research-claw">7. AutoResearchClaw</h3><p class="paragraph" style="text-align:justify;">This AI co-scientist automates the full research loop: generating hypotheses, running experiments, fixing failed code, analyzing results, and writing papers. It uses multiple debating agents, <b>can “self-heal” broken experiments, and stores lessons from past runs.</b> It also verifies citations and checks that every reported metric comes from real experiment logs to reduce hallucinations and fake results.</p><ul><li><p class="paragraph" style="text-align:justify;">AutoResearchClaw: <a class="link" href="https://arxiv.org/abs/2605.20025?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Research Paper</a></p></li><li><p class="paragraph" style="text-align:justify;">GitHub: <a class="link" href="https://github.com/aiming-lab/AutoResearchClaw?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">AutoResearchClaw </a><a class="link" href="https://github.com/aiming-lab/AutoResearchClaw?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p></li></ul><h2 class="heading" style="text-align:justify;" id="faq">FAQ</h2><h3 class="heading" style="text-align:justify;" id="what-are-ai-coscientists">What are AI co-scientists?</h3><p class="paragraph" style="text-align:left;">AI co-scientists are autonomous or semi-autonomous AI systems designed to assist with scientific discovery. Unlike standard chatbots, they can generate hypotheses, analyze papers, run experiments, write code, evaluate results, and sometimes even produce research papers.</p><h3 class="heading" style="text-align:left;" id="what-kinds-of-ai-coscientists-exist">What kinds of AI co-scientists exist?</h3><p class="paragraph" style="text-align:left;">Some systems specialize in biology and drug discovery, others focus on mathematics, scientific simulations, physics, or materials science. Newer systems increasingly combine multiple agents into autonomous research teams.</p><h3 class="heading" style="text-align:left;" id="which-ai-coscientists-are-focused-o">Which AI co-scientists are focused on biology?</h3><p class="paragraph" style="text-align:left;">Google DeepMind’s Co-Scientist, FutureHouse’s Robin, and DISCO are among the strongest biology-focused systems. They work on drug discovery, protein engineering, disease research, and therapeutic hypothesis generation.</p><h3 class="heading" style="text-align:left;" id="which-ai-systems-specialize-in-math">Which AI systems specialize in mathematics?</h3><p class="paragraph" style="text-align:left;">AxiomProver, AI Co-Mathematician, Axplorer, and OpenAI’s reasoning systems focus on theorem proving, formal verification, optimization, and solving difficult mathematical problems.</p><h3 class="heading" style="text-align:left;" id="what-are-scientific-simulation-agen">What are scientific simulation agents?</h3><p class="paragraph" style="text-align:left;">Systems like ERA, kUPS, and AI CFD Scientist autonomously generate scientific simulation code, run computational experiments, optimize models, and verify results across fields like biology, fluid dynamics, chemistry, and forecasting.</p><h3 class="heading" style="text-align:left;" id="what-is-a-fully-autonomous-ai-scien">What is a fully autonomous AI scientist?</h3><p class="paragraph" style="text-align:left;">Systems like The AI Scientist and AutoResearchClaw automate nearly the full research workflow: literature review, idea generation, experimentation, debugging, analysis, verification, and paper writing.</p><h3 class="heading" style="text-align:left;" id="are-there-opensource-ai-coscientist">Are there open-source AI co-scientists?</h3><p class="paragraph" style="text-align:left;">Yes. Projects like The AI Scientist, Axplorer, DISCO, ERA, kUPS, AI CFD Scientist, and AutoResearchClaw are open-source and allow researchers to inspect, modify, and extend their scientific workflows.</p><h3 class="heading" style="text-align:left;" id="why-are-ai-coscientists-becoming-im">Why are AI co-scientists becoming important now?</h3><p class="paragraph" style="text-align:left;">The rise of reasoning models, long-context memory, multi-agent systems, and tool-using AI agents has made it possible for AI systems to handle long-horizon scientific workflows that previously required large human research teams.</p><p class="paragraph" style="text-align:justify;"><b>Further reading</b></p><p class="paragraph" style="text-align:justify;">If you&#39;re just getting started with ML and AI, check out our curated list of <span style="text-decoration:underline;"><a class="link" href="https://www.turingpost.com/p/10-github-repositories-ai-ml-ds?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">Top 10 GitHub repos for AI & ML practitioners</a></span>— collections of courses, guides, and projects to build your foundations.</p><p class="paragraph" style="text-align:justify;">If you’ve found this article valuable, subscribe for free to our newsletter.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.turingpost.com/subscribe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026"><span class="button__text" style=""> Subscribe </span></a></div><p class="paragraph" style="text-align:justify;">We post helpful lists and bite-sized explanations daily on our X/Twitter. Let’s connect.</p><p class="paragraph" style="text-align:justify;"><a class="link" href="https://x.com/TheTuringPost?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=12-ai-co-scientists-of-2026" target="_blank" rel="noopener noreferrer nofollow">https://x.com/TheTuringPost</a></p></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The Production Gap: Five Patterns for Building Long-Running AI Agents* </title>
  <description>A practical guest post about long-running AI agents, covering checkpointing, memory layers, governance, orchestration, A2A, and MCP interoperability</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d31b63b6-0aa7-456f-980a-6b095abb7665/ChatGPT_Image_May_22__2026__04_04_34_PM.jpg" length="60648" type="image/jpeg"/>
  <link>https://www.turingpost.com/p/the-production-gap-5-patterns-for-building-long-running-ai-agents</link>
  <guid isPermaLink="true">https://www.turingpost.com/p/the-production-gap-5-patterns-for-building-long-running-ai-agents</guid>
  <pubDate>Fri, 22 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-22T21:00:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Addy Osmani</dc:creator>
    <category><![CDATA[Community Twist:  Guest Posts &amp; Practitioner Insights]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Space Mono',Courier,'Lucida Console',Monaco,monospace !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:justify;"><i>Most agent architectures are secretly stateless. In this guest post, Addy Osmani, Director at Google Cloud, and Shubham Saboo, Senior AI Product Manager at Google Cloud, discuss what it takes to build ones that aren’t.</i></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;">Why most AI agents fail in production</h2></div><p class="paragraph" style="text-align:justify;">Developers spend weeks perfecting prompt engineering, tool calling, and response latency. None of it matters when your agent needs to stay alive for five days.</p><p class="paragraph" style="text-align:justify;">The workflows that actually matter in production — processing thousands of insurance claims, running week-long sales sequences, reconciling financial data across systems — don&#39;t fit inside a single conversation turn. They take days, not seconds. And the moment you try to build them, you run into a wall that most tutorials skip over: most agent architectures reconstruct context from scratch on every interaction. They lose the reasoning chain, the soft signals, and the confidence gradients that made the agent&#39;s previous decisions make sense.</p><p class="paragraph" style="text-align:justify;">This is the production gap. Demos close it with short, clean tasks. Real systems don&#39;t get that luxury.</p><p class="paragraph" style="text-align:justify;">At Google Cloud Next &#39;26, we announced that Agent Runtime now supports long-running agents that maintain state for up to seven days. What follows are five design patterns — drawn from what we&#39;ve seen actually work in production — for building agents that survive contact with reality.</p><h2 class="heading" style="text-align:justify;" id="pattern-1-checkpointand-resume">Pattern 1: Checkpoint-and-Resume</h2><div class="image"><img alt="Checkpoint and resume" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/56d95556-a2b8-4d79-87f1-5c507809b144/Screenshot_2026-05-26_at_1.29.57_PM.jpg?t=1779816817"/></div><p class="paragraph" style="text-align:justify;">The most common failure mode in multi-day workflows is context loss. An agent processes 200 documents over four hours, then hits an error on document 201. Without checkpointing, you restart from scratch.</p><p class="paragraph" style="text-align:justify;">The fix is conceptually simple but architecturally important: treat your agent like a long-running server process, not a request handler. The same way you&#39;d build a data pipeline that processes millions of records — checkpoint progress, handle partial failures, ensure idempotency.</p><div class="codeblock"><pre><code>from google.adk import Agent, ToolContext

class DocumentProcessor(Agent):
    &quot;&quot;&quot;Processes large document sets with checkpoint-and-resume.&quot;&quot;&quot;

    async def process_batch(self, docs: list, ctx: ToolContext):
        checkpoint = self.load_checkpoint()
        start_idx = checkpoint.get(&quot;last_processed&quot;, 0)

        for i, doc in enumerate(docs[start_idx:], start=start_idx):
            result = await self.classify_and_extract(doc)
            self.results.append(result)

            # Checkpoint every 50 documents
            if (i + 1) % 50 == 0:
                self.save_checkpoint(&#123;
                    &quot;last_processed&quot;: i + 1,
                    &quot;partial_results&quot;: self.results,
                    &quot;timestamp&quot;: datetime.now().isoformat()
                &#125;)

        return self.compile_final_report()</code></pre></div><p class="paragraph" style="text-align:justify;">Notice the checkpoint granularity. Not after every document (wasteful). Not only at the end (risky). Fifty documents per batch balances durability against overhead. Your specific number depends on how expensive each unit of work is.</p><h2 class="heading" style="text-align:justify;" id="pattern-2-delegated-approval-humani">Pattern 2: Delegated Approval (Human-in-the-Loop, Done Right)</h2><div class="image"><img alt="Delegated Approval (Human-in-the-Loop, Done Right)" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9157e175-1cff-41f7-973b-6b9a19d6394e/Screenshot_2026-05-26_at_1.30.33_PM.jpg?t=1779816878"/></div><p class="paragraph" style="text-align:justify;">Every agent framework advertises human-in-the-loop. But in practice, most implementations amount to: serialize state to JSON, send a webhook, hope someone checks it.</p><p class="paragraph" style="text-align:justify;">The problems compound fast. JSON serialization loses implicit reasoning context. Notifications compete with dozens of other alerts. When the human responds hours later, the agent has to deserialize, re-establish context, and hope nothing changed in the interim.</p><p class="paragraph" style="text-align:justify;">Long-running agents handle this differently. When the agent hits an approval gate, it pauses in place. The full execution state stays intact: reasoning chain, working memory, tool call history, pending action. The agent consumes zero compute while waiting. Sub-second cold starts mean zero latency penalty when it resumes.</p><p class="paragraph" style="text-align:justify;">The critical detail is the time accounting. If an agent pauses for human review at hour 8 and the reviewer responds at hour 32, those 24 hours are dead time for the agent but productive time for the human. The agent doesn&#39;t drift, degrade, or need re-priming. It picks up exactly where it left off.</p><p class="paragraph" style="text-align:justify;">At scale — if you&#39;re managing twenty concurrent long-running agents — you need a unified inbox that categorizes what needs attention. Not Slack channels. Not email threads. A structured queue: &quot;Needs your input,&quot; &quot;Errors,&quot; &quot;Completed.&quot;</p><h2 class="heading" style="text-align:justify;" id="pattern-3-memory-layered-context-an">Pattern 3: Memory-Layered Context (and Why You Have to Govern It)</h2><div class="image"><img alt="Memory-Layered Context (and Why You Have to Govern It)" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d9a92294-8524-4559-beef-f7d07b58b698/Screenshot_2026-05-26_at_1.30.40_PM.jpg?t=1779816900"/></div><p class="paragraph" style="text-align:justify;">A seven-day agent needs more than session state. It needs to remember things from previous sessions, organizational context that no single conversation could contain, and user preferences from weeks ago.</p><p class="paragraph" style="text-align:justify;">This is where the architecture gets interesting — and where most teams underestimate the risk.</p><p class="paragraph" style="text-align:justify;">Think of long-term memory as analogous to a knowledge base: it accumulates everything the agent has learned across interactions, organized by topic. Working memory, by contrast, provides low-latency access to specific, high-accuracy details needed right now. The two layers work together, but they have to be kept distinct.</p><p class="paragraph" style="text-align:justify;">Here&#39;s the problem most developers don&#39;t anticipate until production: <b>memory drift.</b></p><p class="paragraph" style="text-align:justify;">Your agent&#39;s behavior isn&#39;t shaped only by its code and prompts. It&#39;s shaped by accumulated experience. If an agent &quot;learns&quot; from a few atypical interactions that a procedural shortcut is acceptable, it may start applying that shortcut broadly. And if multiple agents read and write to shared memory pools, data leakage between distinct workflows becomes a real risk — the kind that&#39;s hard to detect and harder to explain to a compliance team.</p><p class="paragraph" style="text-align:justify;">This is where the governance layer becomes non-negotiable. You can&#39;t let agents write to a shared memory store unchecked. You need to govern them the same way you govern microservices. Concretely, this means three things:</p><p class="paragraph" style="text-align:justify;"><b>Agent identity.</b> Every agent needs a cryptographic identity that determines exactly which memory banks and tools it&#39;s authorized to access. Think IAM, but for agents.</p><p class="paragraph" style="text-align:justify;"><b>Centralized registry.</b> When you have dozens of long-running agents, you need a single source of truth for which agents are active, what version of the prompt and code they&#39;re running, and what their current execution state is.</p><p class="paragraph" style="text-align:justify;"><b>Policy enforcement at the boundary. </b>A governance layer sitting between the agent and its memory should evaluate every access request against organizational policies. If an agent tries to commit PII to long-term memory, that transaction should be blocked before it happens — not audited after.</p><p class="paragraph" style="text-align:justify;">The question to ask yourself isn&#39;t just &quot;what are my agents doing?&quot; It&#39;s &quot;what are my agents remembering, and how is that changing their behavior over time?&quot;</p><h2 class="heading" style="text-align:justify;" id="pattern-4-ambient-processing">Pattern 4: Ambient Processing</h2><p class="paragraph" style="text-align:justify;">Not every long-running agent interacts with humans. Some are ambient. They watch for events, process data streams, and take action in the background without any user prompting.</p><div class="image"><img alt="Ambient processing" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/affcf65e-abb0-4b29-b123-28a7a8b1f19d/Screenshot_2026-05-26_at_1.30.51_PM.jpg?t=1779816926"/></div><p class="paragraph" style="text-align:justify;">A real-world mesh looks something like the diagram above. A content moderation agent consumes new uploads from Pub/Sub and routes flagged content to human review. A data quality agent watches BigQuery for new rows, detects anomalies, and delegates remediation to a data engineering specialist via A2A. A customer event agent ingests support tickets and classifies them in real time — routing billing questions, technical issues, and VIP cases to dedicated downstream agents. None of these agents wait to be asked. They run continuously, reacting to the event stream.</p><p class="paragraph" style="text-align:justify;">The architectural decision that matters most here ties back to Pattern 3: <b>don&#39;t hardcode policies into the agent.</b></p><p class="paragraph" style="text-align:justify;">Define them in your governance layer and let the agent enforce them at runtime. When policies change, you update once and every ambient agent in the fleet picks up the new rules immediately. This separation matters because ambient agents run unsupervised for long stretches. If you hardcode policies, every policy change requires redeploying every agent. If you externalize policies, you update once and the fleet adapts — without downtime, without redeployment, without the risk that one agent is running an outdated version of your compliance rules.</p><h2 class="heading" style="text-align:justify;" id="pattern-5-fleet-orchestration">Pattern 5: Fleet orchestration</h2><div class="image"><img alt="Fleet orchestration" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ae6ae0d8-bcc4-4b46-99e7-856cd4df67f8/Screenshot_2026-05-26_at_1.31.00_PM.jpg?t=1779816959"/></div><p class="paragraph" style="text-align:justify;">The final pattern is about managing multiple long-running agents as a coordinated fleet. In production, you rarely have a single agent working alone. You have a coordinator agent that delegates sub-tasks to specialist agents, each running independently for different durations.</p><p class="paragraph" style="text-align:justify;">Consider a sales prospecting sequence. A coordinator breaks the work into components: research, scoring, sequencing, outreach, and follow-up. Each of those is a specialist agent running on its own timeline, with its own identity, its own tool permissions, and its own entry in the registry.</p><p class="paragraph" style="text-align:justify;">The coordinator maintains global state and handles handoffs between specialists. This is the same coordinator/worker pattern used in distributed systems for decades. What&#39;s new is that it can be defined declaratively through graph-based workflows, where the structure of the coordination logic is enforced by the framework rather than expressed in a system prompt that an LLM might decide to shortcut.</p><p class="paragraph" style="text-align:justify;">The operational advantage of treating each specialist as an independent unit is that you can update them independently. If your scoring logic needs improvement, you deploy the new version, monitor its performance, and promote it only when the results hold up. A bad deployment in one specialist never cascades to the others.</p><h2 class="heading" style="text-align:justify;" id="a-2-a-and-mcp-the-interoperability-">A2A and MCP: The Interoperability Layer</h2><p class="paragraph" style="text-align:justify;">One thing the five patterns above don&#39;t fully address: most organizations won&#39;t build every agent they need from scratch. The real leverage comes from agents built by different teams — sometimes in different languages, sometimes at different companies — being able to discover and collaborate with each other.</p><p class="paragraph" style="text-align:justify;">Two open protocols are emerging as the connective tissue here. <b>A2A (Agent-to-Agent) </b>standardizes how agents communicate with other agents. <b>MCP (Model Context Protocol)</b> standardizes how agents communicate with tools and data sources. Together, they mean that a Python-based coordinator can delegate to a Go-based specialist, which can delegate to a Java-based compliance checker, without any of those teams needing to negotiate custom integration formats.</p><p class="paragraph" style="text-align:justify;">Every A2A-compatible agent publishes a card at a well-known URL describing its capabilities, authentication requirements, and rate limits. Think of it as an OpenAPI spec designed for agent-to-agent interaction rather than client-server communication. A central registry amplifies this: other agents across your organization can discover capabilities without knowing specific URLs. The registry becomes the service mesh for your agent ecosystem.</p><div class="image"><img alt="Agent Card Discovery" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3a9fb469-2841-48fe-9230-3c0e2d15f593/image.png?t=1779477892"/></div><p class="paragraph" style="text-align:justify;">The diagram above shows what this looks like in practice. A coordinator agent doesn&#39;t need to know the URL of the financial analysis agent, or that it&#39;s written in Python, or how its auth works. It queries the registry, finds the card, and connects. When the document processing team ships a new version of their Java agent, they update the card. Every coordinator in the organization gets the upgrade automatically.</p><p class="paragraph" style="text-align:justify;">MCP handles the other side: connecting agents to databases, enterprise systems, and APIs through a single protocol. Without it, every data connection requires its own custom integration code. With it, a Stripe connector looks the same to your agent as a BigQuery connector — the protocol is the interface, and the backend is interchangeable.</p><div class="image"><img alt="Tool Bridge" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fde45821-9fe1-4033-861a-4a030fa05bd2/image-17.jpg?t=1779527935"/></div><p class="paragraph" style="text-align:justify;">The governance story matters here too. Each organization maintains its own governance boundaries. Your policy enforcement layer controls what data your agents can share with a partner agent, what actions they&#39;re permitted to take based on the partner&#39;s responses, and what information they&#39;re allowed to request. Cross-organization collaboration happens through the protocol; each side enforces its own security model independently.</p><p class="paragraph" style="text-align:justify;">The multi-language, cross-team version of this plays out exactly as you&#39;d expect. A customer onboarding workflow might involve a Python coordinator, a Go identity-verification agent owned by the security team, a Java credit-assessment agent owned by risk, a Go account-provisioning agent owned by platform, and a TypeScript communication agent owned by marketing — each team iterating independently, none of them blocked by the others.</p><div class="image"><img alt="Delegated Specialization" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bb128b76-7495-4da9-9a58-77f8c1e1bda7/image-18.jpg?t=1779527972"/></div><h2 class="heading" style="text-align:left;" id="choosing-the-right-pattern">Choosing the Right Pattern</h2><p class="paragraph" style="text-align:justify;">These patterns compose. A compliance system might use Checkpoint-and-Resume for document processing, Delegated Approval for review gates, Memory-Layered Context for cross-session knowledge, and Fleet Orchestration to coordinate specialists.</p><p class="paragraph" style="text-align:justify;">The key diagnostic question:<b> what is the longest uninterrupted unit of work your agent needs to perform?</b></p><p class="paragraph" style="text-align:justify;">If it&#39;s minutes, you probably don&#39;t need long-running agents. If it&#39;s hours or days, these patterns are where you start — and the governance and interoperability layers become load-bearing, not optional.</p><p class="paragraph" style="text-align:justify;">The companies building isolated, stateless agents today will be refactoring in twelve months. The ones building with persistence, governance, and interoperability in mind will be compounding their advantage every day.</p><div class="section" style="background-color:transparent;border-color:#991cbf;border-radius:1px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:justify;"></p><p class="paragraph" style="text-align:justify;">The <a class="link" href="https://fandf.co/4dJLtUe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-production-gap-five-patterns-for-building-long-running-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Gemini Enterprise Agent Platform</a> provides the infrastructure for building, deploying, and governing long-running agent fleets described in this article. You can explore it <a class="link" href="https://fandf.co/4dJLtUe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-production-gap-five-patterns-for-building-long-running-ai-agents" target="_blank" rel="noopener noreferrer nofollow">here. </a> </p><p class="paragraph" style="text-align:justify;"></p></div><hr class="content_break"><p class="paragraph" style="text-align:justify;"><i>*This guest post was written by</i> Addy Osmani, Director, Google Cloud and Shubham Saboo, Senior AI Product Manager, Google Cloud<i>. We thank </i><a class="link" href="https://fandf.co/4dJLtUe?utm_source=www.turingpost.com&utm_medium=newsletter&utm_campaign=the-production-gap-five-patterns-for-building-long-running-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><i>Google Cloud</i></a><i> for their support of Turing Post’s mission to bring clarity to the AI landscape.</i></p><hr class="content_break"><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:justify;" id="faq">FAQ</h2><h3 class="heading" style="text-align:left;">What are long-running AI agents?</h3><p class="paragraph" style="text-align:left;">Long-running AI agents are AI systems designed to preserve execution state, memory, and workflow continuity across hours or days instead of resetting after each interaction.</p><h3 class="heading" style="text-align:left;">Why do most AI agents fail in production?</h3><p class="paragraph" style="text-align:left;">Most agents are effectively stateless. They lose reasoning history, context, and operational state between interactions, which breaks multi-step workflows and long-duration tasks.</p><h3 class="heading" style="text-align:left;">What is checkpoint-and-resume for AI agents?</h3><p class="paragraph" style="text-align:left;">Checkpoint-and-resume allows an agent to save execution state during long workflows so it can recover from failures or pauses without restarting from scratch.</p><h3 class="heading" style="text-align:left;">What is A2A in AI systems?</h3><p class="paragraph" style="text-align:left;">A2A, or Agent-to-Agent, is a protocol that standardizes communication between AI agents so different systems can discover and coordinate with each other.</p><h3 class="heading" style="text-align:left;">What is MCP?</h3><p class="paragraph" style="text-align:left;">MCP, or Model Context Protocol, standardizes how AI agents connect to tools, APIs, databases, and external systems.</p><h3 class="heading" style="text-align:left;">Why does governance matter for long-running agents?</h3><p class="paragraph" style="text-align:left;">Long-running agents accumulate memory, permissions, and operational context over time. Governance layers control identity, memory access, compliance, and policy enforcement across agent fleets.</p></div><h3 class="heading" style="text-align:left;" id="what-are-longrunning-ai-agents"></h3></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
