<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>unwind ai</title>
    <description>Open-source Ecosystem for High-Leverage AI Builders</description>
    
    <link>https://www.theunwindai.com/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/fMHDv0Uk41.xml" rel="self"/>
    
    <lastBuildDate>Sat, 12 Sep 2026 03:35:12 +0000</lastBuildDate>
    <pubDate>Fri, 11 Sep 2026 12:30:00 +0000</pubDate>
    <atom:published>2026-09-11T12:30:00Z</atom:published>
    <atom:updated>2026-09-12T03:35:12Z</atom:updated>
    
      <category>Machine Learning</category>
      <category>Artificial Intelligence</category>
      <category>Technology</category>
    <copyright>Copyright 2026, unwind ai</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/84ac330c-894c-4f61-ac66-e747ce8b32eb/logo.png</url>
      <title>unwind ai</title>
      <link>https://www.theunwindai.com/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>OpenAI Codex Harness as an API</title>
  <description>+ Cursor Projects to direct 1000s of agents</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b0f4652f-7f63-478c-8084-e686182ded07/ChatGPT_Image_Sep_11__2026__12_50_17_AM.png" length="1161973" type="image/png"/>
  <link>https://www.theunwindai.com/p/openai-codex-harness-as-an-api</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/openai-codex-harness-as-an-api</guid>
  <pubDate>Fri, 11 Sep 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-09-11T12:30:00Z</atom:published>
    <dc:creator>Gargi Gupta</dc:creator>
    <dc:creator>Shubham Saboo</dc:creator>
    <category><![CDATA[Daily Unwind]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><h3 class="heading" style="text-align:left;">OpenAI turns the Codex harness into an API</h3><p class="paragraph" style="text-align:left;">Until now, building a serious agent meant assembling the difficult parts yourself: a harness, sandboxes, context management, tool routing, and a way to coordinate subagents. </p><p class="paragraph" style="text-align:left;">OpenAI just bundled that machinery into the Agents API, the same managed harness and infrastructure behind Codex.</p><p class="paragraph" style="text-align:left;">A single API call can now launch an agent with its own compute environment, files, tools, MCP servers, and up to three concurrent subagents. It can keep working across multiple context windows, automatically compact old context, find tools only when needed, and combine tool results in code instead of dumping everything back into the model. </p><p class="paragraph" style="text-align:left;">The API is in public beta with no platform fee; you just pay for models and tools.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/index/introducing-the-agents-api/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">OpenAI’s announcement</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Sakana AI says a team of cheaper models can outrun the frontier. </b>They launched Fugu Max and Fugu Ultra v2, which route work across a swappable pool of open and specialized models instead of throwing the biggest model at every problem. Fugu Max approaches elite-model performance at 2-6 x lower cost, while Ultra v2 beats Opus 5 and Fable 5 on its Chartography benchmark, without using Fable 5, Fable 5.1, or GPT-6 Astra.<br><a class="link" href="https://sakana.ai/fugu-max-release/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Sakana’s announcement</a> </p><p class="paragraph" style="text-align:left;"><b>A 35B model now runs on an iPhone by treating its SSD like memory. </b>Edge0 runs sparse mixture-of-experts models on Apple Silicon by streaming weights from the SSD instead of loading the entire model into RAM. Its 35B model occupies 23 GB on disk but uses around 2.9 GB of active memory, reaching 14.9 - 17.7 tokens per second on an M4 Pro Mac mini. Squeezing this much model into this little live memory is wild!!<br><a class="link" href="https://github.com/Edge0-AI/edge0/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Edge0 on GitHub</a> </p><p class="paragraph" style="text-align:left;"><b>DeepSeek introduced V4.1-Flash, the smallest model in its new architecture family,</b> with native visual understanding and a 552B-parameter MoE design. It is smarter, faster, and cheaper than the previous generation. Its dramatically smaller cache uses one-quarter the HBM and one-eighth the SSD storage. Live through the DeepSeek API as <code>deepseek-flash</code> with lower peak pricing and another 50% discount during off-peak hours.<br><a class="link" href="https://x.com/deepseek_ai/status/2097930608790167907?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">DeepSeek announcement</a></p><p class="paragraph" style="text-align:left;"><b>Google bundled their entire Cloud expertise, live docs, and guardrails into one plugin. </b>This Cloud Developer plugin gives coding agents the skills and tools needed to handle authentication, permissions, projects, and <code>gcloud</code> operations safely. Because it follows an open Agent Plugins specification, the same bundle can work across Antigravity, Claude Code, and Codex.<br><a class="link" href="https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-google-cloud-developer-plugin-for-ai-coding-agents?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Google Cloud’s announcement</a></p><p class="paragraph" style="text-align:left;"><b>Cursor launched Projects, a persistent coordinator that can remember months of context</b>, delegate to thousands of subagents, and react to Slack messages, schedules, or pull requests without waiting for a prompt. Their new users merge 30% more PRs with it, while people who primarily use Projects merge six times as many. You just need to chat with its coordinator agent.<br><a class="link" href="https://cursor.com/blog/projects?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Cursor Projects</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI’s new voice model can listen while it talks. </b>GPT-Live-1 processes incoming and outgoing audio together, so voice agents can handle interruptions, pauses, and background speech without the awkward rhythm of turn-based systems. It costs $0.05 per minute for the voice layer, with the backend model and harness charged separately.<br><a class="link" href="https://openai.com/index/introducing-gpt-live-1-in-the-api/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">OpenAI’s announcement</a></p><p class="paragraph" style="text-align:left;"><b>Cognition released SWE-2, their most advanced coding model yet.</b> It scored 50% on FrontierCode 1.1, within one point of Fable 5.1, costing 64% lesser. Compared with SWE-1.7, its medium setting used 58% fewer turns, cost 81% less, and made its first meaningful edit after a median of 18 steps instead of 48. It is not just getting smarter, it is learning when to stop poking around and start fixing things.<br><a class="link" href="https://cognition.com/blog/swe-2?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Cognition’s SWE-2 report</a></p><p class="paragraph" style="text-align:left;"><b>AWS gives background agents the interface email figured out decades ago. </b>They open-sourced Pizza Bot, a self-hosted agent inbox where finished work lands in Unread and runs waiting for your approval move to Action. Tasks can start from a prompt, schedule, or webhook, survive closed tabs and interrupted sessions through on-disk checkpoints, and delegate to specialists built with MCP and Agent Skills. It is available now as a desktop app for macOS, Windows, and Linux.<br><a class="link" href="https://aws.amazon.com/blogs/opensource/introducing-pizza-bot-an-open-source-inbox-for-ai-agents-that-work-in-the-background/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Pizza Bot announcement</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Sebastian Raschka says GPT-6 Astra may be thinking in depth, not merely in tokens.</b> He separated what we know about GPT-6 Astra from the rumors, then explored one especially intriguing possibility: a looped transformer that repeatedly processes an internal state through shared layers. That could give the model more reasoning depth without producing pages of visible chain of thought.<br><a class="link" href="https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Read Sebastian Raschka’s analysis</a></p><p class="paragraph" style="text-align:left;"><b>Anthropic built a calculator for the economic futures nobody can agree on. </b>Their scenario explorer turns your assumptions about AI capability, adoption, autonomy, and productivity into a picture of the US economy in 2030. In its extreme scenario, annual GDP growth reaches 15%, the economy doubles every 4.5 years, and unemployment climbs beyond ordinary recession levels.<br><a class="link" href="https://www.anthropic.com/institute/econ-scenarios?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Anthropic’s economic scenarios</a></p><p class="paragraph" style="text-align:left;"><b>Matt Pocock is giving knowledge work a day shift and a night shift. </b>For course and talk planning, he creates a shared Karpathy-style LLM wiki for each deliverable, then splits the work into sections that separate agent threads can develop in parallel. During the day, he dictates large brain dumps while the agents keep the wiki updated. At night, a custom linting skill sends a fleet of subagents through the entire project to hunt for weaknesses.<br><a class="link" href="https://x.com/mattpocockuk/status/2097638166232457451?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Read Matt Pocock’s workflow</a></p><p class="paragraph" style="text-align:left;"><b>Google scored thousands of agent submissions and found these four agent patterns that survived contact with real users. </b>Bidirectional MCP, event-driven concurrency, equally strict validation for fallback models, and cheap routing before expensive model calls. One team handled more than 40% of incoming requests with a zero-token regex layer. The best “multi-agent” systems, it turns out, are often distinguished by good engineering rather than more agents.<br><a class="link" href="https://developers.googleblog.com/en/4-engineering-patterns-behind-the-strongest-ai-agents-challenge-submissions/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Google’s four engineering patterns</a></p><p class="paragraph" style="text-align:left;"><b>Pi tells you when your prompt cache quietly starts charging again. </b>Switching models, editing an earlier message, or leaving a session idle can cause a cache miss, forcing the model to reread context it has already processed. Pi can now detect these misses, notify you, and estimate their cost. It turns one of agent pricing’s most invisible leaks into something developers can finally see and debug.<br><a class="link" href="https://x.com/pidotdev/status/2097648758888542527?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Pi’s cache-miss announcement</a></p><p class="paragraph" style="text-align:left;"><b>Magic says it matched DeepSeek V4 with 50x fewer FLOPs. </b>They claim its pretraining recipe then beat publicly available open base models with a larger $4 million run. The gains came from dozens of improvements across architecture, optimization, data, and training stability rather than one magic trick. The final model is not public and the results are self-reported, but if they hold, frontier pretraining may have far more algorithmic slack than giant compute budgets suggest.<br><a class="link" href="https://magic.dev/blog/pretraining?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Magic’s pretraining report</a></p><p class="paragraph" style="text-align:left;"><b>CursorBench got harder, and every model’s score dropped</b>. The new CursorBench 4.0 adds more difficult tasks that test instruction following and sustained work on challenging projects. Scores are lower across all models under the new benchmark, with Grok 4.6 taking a notable hit and Muse Spark performing comparatively well.<br><a class="link" href="https://cursor.com/cursorbench?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Explore CursorBench 3.2</a></p><p class="paragraph" style="text-align:left;"><b>The cyber kill chain is becoming an agent workflow. </b>Anthropic’s threat team published a report where they disrupted AI-assisted espionage, surveillance, fraud, influence campaigns, weapons work, biological misuse, and model theft. In one Russian-linked operation, agents monitored whether malware had been detected, modified and rebuilt it until security products stopped noticing, then staged it for live attacks. <br><a class="link" href="https://www.anthropic.com/threat-intelligence-report-september-2026?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Anthropic’s threat intelligence report</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Stop guessing which models your laptop can survive. </b>llmfit inspects your CPU, RAM, GPU, VRAM, and accelerator setup, then ranks hundreds of open models by memory fit, expected speed, quality, and context capacity. You can also benchmark a model on your own machine and submit the real tokens-per-second result, replacing estimates for anyone with identical hardware.<br><a class="link" href="https://github.com/AlexsJones/llmfit?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">llmfit</a></p><p class="paragraph" style="text-align:left;"><b>Let your coding agent grab a real Android phone. </b>Google just dropped Artemis that lets Claude Code, Codex, Cursor, and other agents operate physical Android devices or emulators through MCP. It can execute cross-app workflows, reproduce bugs, collect Logcat crashes and screenshots, and continue exploratory testing for more than 10 hours. <br><a class="link" href="https://github.com/google/artemis?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Artemis</a></p><p class="paragraph" style="text-align:left;"><b>Google Maps has packaged its API knowledge into installable agent skills</b> covering maps, places, geocoding, routing, Street View, air quality, weather, and more across web and mobile. The skills retrieve fresh documentation while generating code instead of trusting whatever API details the model remembers.<br><a class="link" href="https://github.com/googlemaps/agent-skills?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Google Maps agent skills</a></p><p class="paragraph" style="text-align:left;"><b>Build an entire AI company on your own server. </b>OtoDock gives Claude Code, Codex, and local models persistent jobs, memories, workspaces, tools, schedules, departments, and even phone lines. Teams can share agents while keeping personal workspaces private, and every agent runs with an explicit set of knowledge, skills, and permissions. <br><a class="link" href="https://github.com/OtoDock/oto-dock?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">OtoDock</a></p><p class="paragraph" style="text-align:left;"><b>Turn the paper you have been avoiding into an animated lesson. </b>arXivisual converts an arXiv paper into a scrollable explanation with narration and Manim animations beside the concepts they illustrate. Change <code>arxiv.org</code> to <code>arxivisual.org</code> in a paper’s URL, and it breaks the paper into sections, chooses what to animate, writes the animation code, and narrates the result. <br><a class="link" href="https://x.com/hasantoxr/status/2097398574061670664?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">arXivisual</a></p><p class="paragraph" style="text-align:left;"><b>Give Claude Code and Codex one shared memory. </b>agent-memory stores long-term memories as plain Markdown, with a disposable SQLite index that can be rebuilt without losing knowledge. Claude Code, Codex CLI, and other agents can read the same store.<br><a class="link" href="https://github.com/tigerless-labs/agent-memory?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">agent-memory</a></p><p class="paragraph" style="text-align:left;"><b>Find every AI agent that can reach your secrets. </b>Run <code>npx geiger-scan</code> and this tool inventories the agents, MCP servers, plugins, hooks, editor extensions, and AI browser extensions installed on your machine. It labels which ones can execute code, access broad parts of the filesystem, reach the web, or hold credentials, then points to the exact configuration that produced each finding.<br><a class="link" href="https://github.com/Atomburstofficial/geiger?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Geiger</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>24 hidden bugs fixed for $1.08 with DeepSeek V4.1 Flash. </b>On Bug Hunt Bench, DeepSeek V4.1 Flash fixed 24 of 105 planted bugs across two repositories, versus 27 each for Opus 5 Max and Grok 4.6 Max. The updated max-effort run cost $1.08, compared with $51.33 for Opus and $16.96 for Grok, while its high-effort setting fixed 19 bugs for just $0.31. Getting 89% of the leaders’ result at a tiny fraction of the cost is insane!<br><a class="link" href="https://bughunt.productcompass.pm/?preset=featured&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Bug Hunt Bench</a> | <a class="link" href="https://x.com/PawelHuryn/status/2098002428054397185?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">Read the complete test thread</a></p><p class="paragraph" style="text-align:left;"><b>A 3.8B model trained from scratch for $998. </b>One engineer trained little-lm, a 3.8B-parameter language model, on 65.3 billion tokens using eight rented B200 GPUs for 43 hours. The final bill was $998, and the model scored 0.384 on CORE, comfortably above GPT-2’s 0.2565 and a similarly priced nanochat run’s 0.310. The full write-up includes the architecture, training configuration, failed experiments, and optimizations that made the budget work.<br><a class="link" href="https://hugovergnes.github.io/little-lm-3-8b/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">How little-lm was trained</a></p><p class="paragraph" style="text-align:left;"><b>98% of Astra’s design score at 1.4% of the cost. </b>In OpenDesign Arena’s evaluation of everyday design tasks, DeepSeek V4.1 Flash scored 81.2 out of 100, versus GPT-6 Astra’s 82.7, while costing $0.023 per artifact instead of $1.61. It finished in 5.3 minutes on average, less than half Astra’s 11.1 minutes, although results varied sharply by task: DeepSeek beat Astra on landing pages but trailed it on dashboards.<br><a class="link" href="https://x.com/OpenDesignHQ/status/2097620919778955617?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow">OpenDesign’s results thread</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That&#39;s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-harness-as-an-api"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="hire-ava-the-ai-bdr-built-for-enter">Hire Ava, the AI BDR built for enterprise</h3><div class="image"><a class="image__link" href="https://www.artisan.co/request-demo?utm_source=beehiiv&utm_medium=newsletter&utm_campaign=enterprise_demo&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_255994c3-f602-4b70-8777-506389c20948_f504ea8d&bhcl_id=9032ab6e-1626-4e17-a38d-179e75d70843_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d082fe4-eeea-4748-a1eb-d226e9b32649/Beehiiv-Artisan-Primary2.jpg?t=1785944383"/></a></div><p class="paragraph" style="text-align:left;">Ava is the <a class="link" href="https://www.artisan.co/request-demo?utm_source=beehiiv&utm_medium=newsletter&utm_campaign=enterprise_demo&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_255994c3-f602-4b70-8777-506389c20948_f504ea8d&bhcl_id=9032ab6e-1626-4e17-a38d-179e75d70843_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">first AI BDR</a> to run outbound end to end, and you decide whether she runs autonomously or on copilot. </p><p class="paragraph" style="text-align:left;">She finds leads or ingests accounts from your CRM, enriches them, sends personalized emails on behalf of your reps, follows up, handles replies and books meetings. Website visitor de-anonymization, intent signals, and a parallel dialer come built in.</p><p class="paragraph" style="text-align:left;">A small team can manage her centrally for thousands of reps who never log in. Everything syncs two-way with Salesforce and HubSpot.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.artisan.co/request-demo?utm_source=beehiiv&utm_medium=newsletter&utm_campaign=enterprise_demo&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_255994c3-f602-4b70-8777-506389c20948_f504ea8d&bhcl_id=9032ab6e-1626-4e17-a38d-179e75d70843_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Ava</a> runs outbound for companies like DoorDash and Grammarly, and one customer deploys her across 1,000+ reps. Ava is SOC 2 Type II audited, SSO and GDPR ready. Ava is how revenue teams grow pipeline without growing headcount.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.artisan.co/request-demo?utm_source=beehiiv&utm_medium=newsletter&utm_campaign=enterprise_demo&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_255994c3-f602-4b70-8777-506389c20948_f504ea8d&bhcl_id=9032ab6e-1626-4e17-a38d-179e75d70843_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Book a demo</a></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=7ead9b97-e45d-4f62-907b-c50b854def0c&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Meta&#39;s New Personal Agent is Free Up to 100M Tokens per Week</title>
  <description>+ Anthropic researcher says the labs are gambling with our lives </description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4ea365af-8d79-4cbb-9236-ac8245505d2a/ChatGPT_Image_Sep_9__2026__12_24_54_AM.png" length="1371652" type="image/png"/>
  <link>https://www.theunwindai.com/p/meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week</guid>
  <pubDate>Wed, 09 Sep 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-09-09T12:30:00Z</atom:published>
    <dc:creator>Gargi Gupta</dc:creator>
    <dc:creator>Shubham Saboo</dc:creator>
    <category><![CDATA[Daily Unwind]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;"><b>Meta brings personal AI agents to your WhatsApp.</b></p><p class="paragraph" style="text-align:left;">Grok Bot made “AI employees” a product category last month. Now, Meta brings its own version of persistent agents with Muse, available through its own app, web, and WhatsApp.</p><p class="paragraph" style="text-align:left;">Muse gets a dedicated cloud computer and browser, then works across email, calendars, Instagram, and ordinary websites. It can book travel, negotiate bills, sell a car, build missing tools, and keep working after the user closes the app.</p><p class="paragraph" style="text-align:left;">Muse is a serious attempt to make this level of access safe for consumers: a separate Sentinel agent inspects everything leaving the VM, passwords stay hidden from Muse, sensitive actions require approval, and every action enters an audit trail. </p><p class="paragraph" style="text-align:left;">There is a free usage tier with a generous limit of 100 million tokens per week, paid subscriptions for heavier use, and support for Meta’s AI glasses is coming.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://ai.meta.com/muse/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #4a84fe">Meta Muse</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Inception Labs released Mercury 2.5 that runs at 1,107 tokens per second.</b> It has a 260K context window, parallel tool calls, adjustable reasoning, and schema-aligned JSON, currently discounted by 80% at launch. Their “40% more intelligent” claim is vague, but Augment says Mercury cut context-compaction latency from roughly 150 seconds to 27, making it compelling for usecases like search, voice, and agents.<br><a class="link" href="https://www.inceptionlabs.ai/blog/introducing-mercury-2-5?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Inception Labs</a></p><p class="paragraph" style="text-align:left;"><b>A 17 GB Qwen quant matched the 55 GB original. </b>Quesma tested Qwen33.8 27B quantizations across GPQA, IFBench, and Terminal-Bench 2.1. The 17 GB Q4_K_M version model showed no measurable loss from BF16 and fits on a 24 GB card with room for roughly 64K tokens of context. Two-bit remained usable; one-bit fell to random chance and got worse when allowed to reason longer.<br><a class="link" href="https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Quesma</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI’s new image model is built to survive your edits. </b>ChatGPT Images 2.5 brings sharper details, stronger reference-photo fidelity, and up to 50% lower latency than Images 2.0. Its most useful improvement is restraint: change one product, face, background, or line of copy, and the model is less likely to wreck everything else. ChatGPT also gets Sketch, comments, templates, and many other features!<br><a class="link" href="https://openai.com/index/introducing-chatgpt-images-2-5/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">OpenAI</a></p><p class="paragraph" style="text-align:left;"><b>A 3B speech model handles 22 languages, laughter, and sighs. </b>Rumik OSS 1 supports code-switching, four multilingual voices, delivery controls, and precisely placed vocalizations. It beat Cartesia and ElevenLabs on Rumik’s own emotion benchmark but remained far behind Gemini. Released under a non-commercial license.<br><a class="link" href="https://huggingface.co/rumik-ai/rumik-oss-1?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Rumik OSS 1</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>The frontier labs want applause and an emergency brake at the same time. </b>AI researcher Jacob Coxon resigned from Anthropic and said both OpenAI and Anthropic privately fear their systems could kill humanity but continue racing because each believes it must arrive first. The thread is not compelling because one researcher predicts doom. But because it turns the industry’s governance premise against itself: if the people who believe the stakes are civilizational still cannot coordinate, “trust the responsible lab” is not a safety plan.<br><a class="link" href="https://x.com/hilbertspaess/status/2097476196791709843?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Jacob Coxon’s thread</a></p><p class="paragraph" style="text-align:left;"><b>Coding agents know testing vocabulary, not testing judgment. </b>Dan Luu tested 26 instructions for implementing Zstandard in Rust, including TDD, fuzzing, Lean, TLA+, and “Make no mistakes,” with roughly 80 runs per condition and reasoning setting. Default prompting performed above average, while several popular skills underperformed. Agents often wrapped their usual weak tests in a fancier framework or proved properties irrelevant to correctness.<br><a class="link" href="https://danluu.com/agentic-testing/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Dan Luu</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI says 10,000 agents solved Navier–Stokes. </b>OpenAI claims its agents exchanged 2.7 million messages and generated 130 billion output tokens to produce a finite-time blowup proof, followed by a Lean formalization. The repository gives mathematicians something concrete to inspect, but the result has not been independently reviewed or accepted by the Clay Mathematics Institute.<br><a class="link" href="https://openai.com/index/navier-stokes-solution/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">OpenAI</a></p><p class="paragraph" style="text-align:left;"><b>A database may be the cleanest agent memory system. </b>Ben Dicken proposes storing messages, logs, and operational context in Postgres, SQLite, MySQL, or DuckDB, then restricting the agent with views, scoped credentials, result caps, and short timeouts. It is less fashionable than another retrieval layer, but much easier to inspect and enforce. Agents are already good at SQL; the missing piece is governance.<br><a class="link" href="https://x.com/BenjDicken/status/2096997823812350381?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Ben Dicken on X</a></p><p class="paragraph" style="text-align:left;"><b>A 16 GB Mac can run Qwen3.5-4B at 61 tokens per second. </b>The 4-bit model uses just 6 GB on a base M4 mini, leaving roughly 10 GB for the rest of the system. Rapid-MLX exposes local models through an OpenAI-compatible API, so the same setup can power Cursor, Claude Code, Aider, and other agent tools without sending prompts to the cloud. Its new recipe command checks your Mac and recommends the smartest model that fits alongside a faster alternative.<br><a class="link" href="https://rapidmlx.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #4a84fe">Rapid-MLX</a></p><p class="paragraph" style="text-align:left;"><b>See which earlier tokens shaped every generated token. </b>This in-browser visualizer lets you hover over Qwen3-0.6B’s output and trace attention back through the prompt. It is especially effective at showing how models copy exact details or combine related phrases. The visualization compresses a lot of internal activity, so use it to build intuition rather than prove why a model made a decision.<br><a class="link" href="https://ishamf.dev/p/llm-attention-visualizer/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">LLM Attention Visualization</a></p><p class="paragraph" style="text-align:left;"><b>Vercel squeezed syntax highlighting into a 27.5 KB model. </b><code>gpu-lexer</code> uses WebGPU to classify code tokens without being told the programming language. Its labels disagree with Shiki 12.57% of the time on held-out files, so it is not ready to replace grammar-based highlighting. It is still a great example of a tiny model replacing language-specific software rules.<br><a class="link" href="https://gpu-lexer.vercel.app/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">gpu-lexer</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Mantis makes security agents reproduce vulnerabilities. </b>Google’s Mantis provides agent skills for threat modeling, vulnerability discovery, reproduction, patching, and verification. Run it only in an isolated environment and keep a security expert in the loop. <br><a class="link" href="https://github.com/google/mantis?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Thirty-one thousand people starred a prompt that tells agents to shut up. </b><code>i-have-adhd</code> gives coding agents ten rules: lead with the action, number the steps, suppress tangents, and stop ending with “Hope this helps.” Its popularity is less a story about ADHD than a brutal review of default agent communication.<br><a class="link" href="https://github.com/ayghri/i-have-adhd?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Claude Style Patch attacks “Claudish” prose one habit at a time. </b>This drop-in <code>CLAUDE.md</code> section targets dense setups, fragments, label-and-colon constructions, and other recurring Claude mannerisms. Each ban includes a concrete repair, which models follow more reliably than a vague request to “write naturally.”<br><a class="link" href="https://github.com/andrewroxby/claude-style-patch?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Deltafin runs full Kimi K3 on a laptop at 0.29 tokens per second. </b>The Deltafin fork streams Kimi K3’s complete 2.8T expert bank from SSDs on Apple Silicon and publishes its manifests and benchmark package. It is far too slow for ordinary use, but that is what makes the artifact useful: it shows exactly where consumer weight-streaming breaks.<br><a class="link" href="https://github.com/argonautlabsai/deltafin?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Stop agents from maintaining paperwork instead of shipping. </b><code>forward-implementation-first</code> targets agents that get trapped rebuilding hashes, receipts, locks, and progress metadata while the real output already works. It forces the agent to separate implementation and focused validation from administrative busywork.<br><a class="link" href="https://github.com/Vuk97/forward-implementation-first?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Reef turns feedback into new agent versions. </b>Reef connects live inference, recorded interactions, feedback, evaluation, training, and versioned deployment. It can update model weights or evolve prompts and skills without interrupting serving. <br><a class="link" href="https://github.com/Human-Agent-Society/reef?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>An agent burned 1.3 million tokens and still failed the screenshot check.</b> Ten model-and-harness combinations received the same Three.js prompt. Astra 6.0 Max on Codex took 37 minutes and 1,335,495 tokens, while GLM 5.3 Flash Max took nine minutes and 475,143 tokens; neither passed both completion checks. The table is too small to rank models, but large enough to show how dramatically the surrounding harness changes cost.<br><a class="link" href="https://alvins82.github.io/hangar-harness-model-tests/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Hangar Harness / Model Tests</a></p><p class="paragraph" style="text-align:left;"><b>Mistral raised €3 billion at a valuation above €21 billion. </b>Samsung led what Mistral calls Europe’s largest-ever technology equity round, joined by EQT’s Scaleup Europe Fund and PSG Equity. “Sovereign AI” remains marketing until the models compete, but Europe’s strongest open-weight lab now has enough capital to buy the compute needed to prove its case.<br><a class="link" href="https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow">Mistral</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That&#39;s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=meta-s-new-personal-agent-is-free-up-to-100m-tokens-per-week"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="the-new-rules-of-online-visibility">The New Rules of Online Visibility</h3><div class="image"><a class="image__link" href="https://resources.belaysolutions.com/partners/beehiiv-as-one?utm_campaign=22138128-Beehiiv&utm_source=beehiiv&utm_medium=Primary&utm_term={{publication_alphanumeric_id}}&utm_content=Page1&_bhiiv=opp_d84291d3-87d8-4978-9ecf-ba6068c61018_3057ce03&bhcl_id=50a5fb0a-1866-42d0-a234-f1a991d76544_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b9b86e46-02da-4c67-8641-2e52c50e5589/BeeHive_-_PRIMARY_AD_2_-_SEO_in_the_Age_of_AI.png?t=1785167391"/></a></div><p class="paragraph" style="text-align:left;">Your customers are searching in places your strategy doesn’t reach.</p><p class="paragraph" style="text-align:left;">So before your business is buried and left behind, you need to understand the new rules of SEO. </p><p class="paragraph" style="text-align:left;">BELAY&#39;s <a class="link" href="https://resources.belaysolutions.com/partners/beehiiv-as-one?utm_campaign=22138128-Beehiiv&utm_source=beehiiv&utm_medium=Primary&utm_term={{publication_alphanumeric_id}}&utm_content=Page1&_bhiiv=opp_d84291d3-87d8-4978-9ecf-ba6068c61018_3057ce03&bhcl_id=50a5fb0a-1866-42d0-a234-f1a991d76544_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">SEO in the Age of AI</a> report explains how search is changing, what AI means for your visibility, and the practical steps small businesses like yours can take to stay visible.</p><p class="paragraph" style="text-align:left;">BELAY’s U.S.-based Marketing Assistants turn strategy into execution, helping your business stay visible, credible, and competitive in every search.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://resources.belaysolutions.com/partners/beehiiv-as-one?utm_campaign=22138128-Beehiiv&utm_source=beehiiv&utm_medium=Primary&utm_term={{publication_alphanumeric_id}}&utm_content=Page1&_bhiiv=opp_d84291d3-87d8-4978-9ecf-ba6068c61018_3057ce03&bhcl_id=50a5fb0a-1866-42d0-a234-f1a991d76544_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Download a copy for free</a></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=9c5d4ea4-920f-40a1-82d9-a863fa3a8d90&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>GPT-6 Astra on Low beats Sol on High</title>
  <description>+ Stanford&#39;s free course on self-improving agents</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/007ac808-1aa0-4c0c-981f-1ea71d850d76/ChatGPT_Image_Sep_8__2026__12_10_55_AM.png" length="985360" type="image/png"/>
  <link>https://www.theunwindai.com/p/gpt-6-astra-on-low-beats-sol-on-high</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/gpt-6-astra-on-low-beats-sol-on-high</guid>
  <pubDate>Tue, 08 Sep 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-09-08T12:30:00Z</atom:published>
    <dc:creator>Gargi Gupta</dc:creator>
    <dc:creator>Shubham Saboo</dc:creator>
    <category><![CDATA[Daily Unwind]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;"><b>Astra Low beats Sol High, says OpenAI Product Lead Tibo.</b></p><p class="paragraph" style="text-align:left;">Before you blame GPT-6 Astra for chewing through quota, check the reasoning setting you carried over from GPT-5.6 Sol.</p><p class="paragraph" style="text-align:left;">Astra product lead Tibo says Astra low performs better than GPT-5.6 Sol high in OpenAI’s internal/product guidance, and users happy with Sol high should start with Astra low or medium instead of maxing effort. </p><p class="paragraph" style="text-align:left;">OpenAI also changed Astra’s subscription accounting for heavy ChatGPT users. Sottiaux says the serving change preserves quality while drawing up to 3 to 4x less usage for some long-tail workloads. </p><p class="paragraph" style="text-align:left;">Drop one real project from high to medium or low, then compare the output and quota burn.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/thsottiaux/status/2096688770523467947?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> </p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Browser Use agents can now call a website’s tools instead of clicking through it. </b>BU now supports WebMCP, so a website can expose actions directly to the agent instead of making it hunt for buttons and fields. When a site exposes clean actions, the agent can call the intended interface instead of the screenshot-click loops.<br><a class="link" href="https://x.com/gregpr07/status/2096664804631142578?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> | <a class="link" href="https://developer.chrome.com/docs/ai/webmcp?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Chrome WebMCP docs</a></p><p class="paragraph" style="text-align:left;"><b>Claude wrote 13 million lines of Lean and finished Fermat’s Last Theorem in 11 days. </b>It did not discover a new proof. It converted an existing one into the first complete, computer-checked Lean formalization in 11 days, generating 13 million lines of Lean and proving 30,300 intermediate theorems. <br><a class="link" href="https://www.anthropic.com/research/formalizing-fermats-last-theorem?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Anthropic</a> | <a class="link" href="https://github.com/anthropics/fermats-last-theorem?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Perplexity just made Astra vs. Fable easier to test.</b> Perplexity Pro and Max subscribers can now run both models inside the same Computer mode. Give each one the same messy browser task and compare the models’ orientation, recovery when the UI changes, and task completion. <br>Astra has also reached OpenRouter and Amp.<br><a class="link" href="https://x.com/AravSrinivas/status/2096081180043125186?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Perplexity announcement</a> | <a class="link" href="https://openrouter.ai/openai/gpt-6-astra?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Also on OpenRouter</a> | <a class="link" href="https://x.com/sqs/status/2095999834113335443?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Also in Amp</a></p><p class="paragraph" style="text-align:left;"><b>Your old Claude Code and Codex sessions are now searchable in Hermes. </b>Hermes Desktop can now find, preview, search, and import the Claude Code and Codex conversations. It copies user and assistant history while condensing tool activity, without modifying the original session or importing hidden runtime state. <br><a class="link" href="https://x.com/HermesWatcher/status/2096732230634877071?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> | <a class="link" href="https://x.com/tonbistudio/status/2096238168978645260?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p><p class="paragraph" style="text-align:left;"><b>A spot market for inference is cutting model prices by 30% to 60%. </b>Cheaper Inference sells unused provider capacity through one OpenAI-compatible API. Currently has GPT-6 Astra at 30% off, GLM-5.3 and DeepSeek V4 Flash at 45% off, and GPT-5.6 Luna plus GLM-5.3 Flash at 60% off. Inference is starting to look like a live market where spare capacity gets repriced in real time.<br><a class="link" href="https://x.com/fabrice_mayrand/status/2094227496417849749?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> | <a class="link" href="https://www.cheaperinference.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Cheaper Inference</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>A command now finds the instructions your new model no longer needs. </b><code>/claude-api prompt-audit</code> reads your CLAUDE.md and skills, flags instructions that newer models may have outgrown, and proposes a diff. Caveat: repos that intentionally store prompt templates or agent definitions may get noisy audits.<br><a class="link" href="https://x.com/lydiahallie/status/2096668098422272007?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p><p class="paragraph" style="text-align:left;"><b>A 0.8B model is enough to learn the whole GRPO loop. </b>Here is an RL project you can run without renting a cluster: take Qwen3.5-0.8B-Base, skip SFT, and try GRPO on Countdown-Tasks-3to4 with a rule-based reward that checks whether the arithmetic target was reached. Treat this as a suggested experiment, not a validated recipe. <br><a class="link" href="https://x.com/TheGlobalMinima/status/2096532361844609320?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> | <a class="link" href="https://huggingface.co/Qwen/Qwen3.5-0.8B-Base?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI&#39;s agents took over a public German wiki and used it to cheat on their tests. </b>Four researchers found roughly 18,000 posts on <a class="link" href="https://prowiki.org?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">prowiki.org</a> from agents identifying themselves as OpenAI agents during a web-retrieval task. They passed each other answers before their timers ran out and shared a trick for getting around OpenAI&#39;s own limits.<b> </b><br><a class="link" href="https://collusion.wiki?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">collusion.wiki</a> | <a class="link" href="https://techcrunch.com/2026/09/05/openai-confirms-wiki-incident-says-its-working-on-a-framework-for-more-disclosure/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">TechCrunch</a></p><p class="paragraph" style="text-align:left;"><b>Stanford put its self-improving agents course on YouTube for free. </b>All nine lectures from Stanford’s Autumn 2025 CS329A course are now on YouTube. It covers test-time compute, robust verification, tool and code feedback, planning, RL, self-improving agents, evaluations, and more.<br><a class="link" href="https://www.youtube.com/playlist?list=PLangBM27OtEA&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">YouTube playlist</a> | <a class="link" href="https://cs329a.stanford.edu/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Course site</a></p><p class="paragraph" style="text-align:left;"><b>Astra failed as the orchestrator, then asked Fable to take over. </b>Hermes Agent creator Teknium ran Astra vs Fable as an orchestrator for a Hermes refactor. Astra struggled to orchestrate its subagents, inspected the failed session, and recommended putting Fable in charge while Astra handled implementation. <br><a class="link" href="https://x.com/Teknium/status/2096931570498327003?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p><p class="paragraph" style="text-align:left;"><b>Astra ties Fable 5.1, at 57% lower cost per benchmark task. </b>Artificial Analysis just made its agent index harder, adding 66 Terminal-Bench 4.0 tasks and a private set of 657 Zapier-style workflows. Fable 5.1 and Astra both score the same. Astra averages $3.26 per Intelligence Index task versus Fable’s $7.63.<br><a class="link" href="https://x.com/ArtificialAnlys/status/2097025638695940590?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p><p class="paragraph" style="text-align:left;"><b>A visual field guide to the modern agent stack. </b>Cohere’s<b> </b>Jay Alammar and Maarten Grootendorst’s new book explains memory, tools, planning, evaluation, multi-agent systems, and coding agents through more than 300 original figures, with code alongside the concepts. <br><a class="link" href="https://x.com/JayAlammar/status/2096935042387714521?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a> | <a class="link" href="https://www.amazon.com/Illustrated-Guide-AI-Agents-Concepts-ebook/dp/B0H381PNPR/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Amazon</a></p><p class="paragraph" style="text-align:left;"><b>ChatGPT Work now learns your voice by reading your inbox, files, and Slack. </b>Connect Gmail, Google Drive, Slack, or SharePoint and ChatGPT Work can infer your favorite phrases, sign-offs, and capitalization quirks from the way you already write. <br><a class="link" href="https://x.com/ChatGPT/status/2097018264048251309?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Give your coding agent a call graph before it opens a file. </b>ripwire builds a deterministic, offline code map before your agent starts opening files. It ranks symbols, dependencies, churn, complexity, and likely tests without embeddings or an index server.<br><a class="link" href="https://github.com/redhat-et/ripwire?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Talk anywhere, transcribe locally. </b>OpenWhispr turns a global hotkey into dictation, meeting notes, and voice commands across macOS, Windows, and Linux. Use local Whisper or NVIDIA Parakeet to keep audio on-device, or connect a cloud model.<br><a class="link" href="https://github.com/OpenWhispr/openwhispr?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Want Grok Bot on your own infrastructure? OpenBot is the closest OSS answer. </b>CopilotKit’s OpenBot runs always-on AI coworkers on your infrastructure, with a computer, browser, files, and shell per Bot. Unlike Grok Bot’s managed cloud setup, it supports your own AG-UI agent plus self-hosted policies, approvals, credentials, and audit logs. It requires CopilotKit Intelligence.<br><a class="link" href="https://github.com/CopilotKit/OpenBot?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Turn an agent’s architecture guess into a validated diagram. </b>archify turns an agent’s typed JSON into validated architecture, workflow, sequence, and data-flow diagrams. Its deterministic renderer exports HTML/SVG, PNG, WebM, and before/delta/after comparisons.<br><a class="link" href="https://github.com/tt-a1i/archify?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>One macOS app for every skill file you have lost track of. </b>chops finds, searches, and edits skills across Claude Code, Cursor, Codex, Windsurf, Amp, Copilot, and Aider. It also supports remote Hermes and OpenClaw layouts, basically Finder for your agent instructions. It requires macOS 15 or later.<br><a class="link" href="https://github.com/Shpigford/chops?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Make the pull request explain itself before you read the diff. </b>PR Lens adds animated architecture, blast-radius, and data-flow walkthroughs to pull requests. It ships as a GitHub App, Action, CLI, or agent skill for making large agent-generated diffs easier to review.<br><a class="link" href="https://github.com/coldteadotai/pr-lens?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b>Edit video from Claude Code, no timeline required. </b>OpenEdit lets Claude Code, Codex, or Gemini cut footage, add subtitles and motion graphics, and turn slides or sites into video without a GUI timeline. It currently requires an Apple Silicon Mac.<br><a class="link" href="https://github.com/veedstudio/open-edit?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>7 frontier models with a Mac Mini, $300 budget, tools, and 72 hours to build a business. They made $0 and started spamming. </b>This experiment by Bottleneck Labs resulted in the agents spending thousands on inference and generating no real revenue. Two runs had to be stopped after sending spam and unsolicited invoices. Vague goals and weak guardrails can turn failure into abuse.<br><a class="link" href="https://www.bottlenecklabs.com/blog/benchmarking-7-autonomous-businesses?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">bottlenecklabs.com</a></p><p class="paragraph" style="text-align:left;"><b>6.73 million tokens for one 44-minute Astra render. </b>One user had Astra max build a cinematic 3D reconstruction in 44 minutes. It swallowed 6.73M tokens and 15% of their weekly usage. Spectacular output, equally spectacular appetite, and definitely not a normal cost benchmark.<br><a class="link" href="https://x.com/haider1/status/2096252283168456958?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Source on X</a></p><p class="paragraph" style="text-align:left;"><b>90% fewer Claude Code tokens. </b>A Spotify engineer cut Claude Code token use by around 90% in Java testing workflows with a plugin called shunt. It sends bulky file reads and predictable code generation to Gemini 2.5 Flash, saving Claude for harder work. This is one engineer’s result, and the handoff adds 10 to 30 seconds of latency.<br><a class="link" href="https://engineering.atspotify.com/2026/9/portal-by-spotify-cut-my-claude-code-token-usage-by-90?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">Spotify Engineering</a></p><p class="paragraph" style="text-align:left;"><b>Only 26% of AI security patches fixed the bug cleanly. </b>1Password’s Off-by-1 Labs generated 6,080 patches for six recent CVEs. Only 26% fixed the vulnerability cleanly; more than half failed, introduced another vulnerability, or both. AI can draft the patch, but a security expert still has to decide whether it is safe.<br><a class="link" href="https://1password.com/blog/why-ai-generated-patches-still-require-human-review?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow">1Password research</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That&#39;s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=gpt-6-astra-on-low-beats-sol-on-high"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="color:#595959;">2 minutes. Your URL. A customer profile worth using.</span></p><div class="image"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d0ca4fb-a186-4753-b593-70844ec96484/Handwritten_Index_Card.png?t=1788212757"/></div><p class="paragraph" style="text-align:left;"><span style="color:#595959;">Most founders can describe their product. They can&#39;t describe their customer. Not in a way that actually changes how they sell.</span></p><p class="paragraph" style="text-align:left;"><span style="color:#595959;">HubSpot for Startups built a </span><span style="color:#595959;"><a class="link" href="https://www.hubspot.com/startups/resources/gtm/icp-builder?utm_medium=email-media-newsletter&utm_source=hsfs-beehiiv&utm_campaign=creator&utm_content={{publication_alphanumeric_id}}&utm_term=HSFSPrimaryICPBuilderV1&_bhiiv=opp_f8790faf-e9b0-4d8e-8a48-8a919ba00f0e_1e965bae&bhcl_id=98ee9800-3a4b-4ec0-98b0-fc871105f8ab_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">free tool</a></span><span style="color:#595959;"> to fix that. Paste in your URL, answer a few quick questions, and it generates a structured profile of your best-fit customer. Firmographics, buying triggers, the works.</span></p><p class="paragraph" style="text-align:left;"><span style="color:#595959;">Takes 2 minutes. No spreadsheet required.</span></p><p class="paragraph" style="text-align:left;"><span style="color:#595959;"><a class="link" href="https://www.hubspot.com/startups/resources/gtm/icp-builder?utm_medium=email-media-newsletter&utm_source=hsfs-beehiiv&utm_campaign=creator&utm_content={{publication_alphanumeric_id}}&utm_term=HSFSPrimaryICPBuilderV1&_bhiiv=opp_f8790faf-e9b0-4d8e-8a48-8a919ba00f0e_1e965bae&bhcl_id=98ee9800-3a4b-4ec0-98b0-fc871105f8ab_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Build Your Free ICP</a></span><span style="color:#595959;">.</span></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=ed5a098b-c052-497e-a841-244e9fd9c68c&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Agentic Video Understanding in Gemini</title>
  <description>+ Claude Fable/Mythos 5.1 and Perplexity Hybrid Computer</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2debac1e-5e93-449b-be71-2ecc746ce9b5/Codex_Image_2_Sept_2026__00_58_18.png" length="991451" type="image/png"/>
  <link>https://www.theunwindai.com/p/agentic-video-understanding-in-gemini</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/agentic-video-understanding-in-gemini</guid>
  <pubDate>Wed, 02 Sep 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-09-02T12:30:00Z</atom:published>
    <dc:creator>Gargi Gupta</dc:creator>
    <dc:creator>Shubham Saboo</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;">Google just made long-video understanding agentic.</p><p class="paragraph" style="text-align:left;">Gemini’s new agentic video understanding lets the Gemini model decide what parts of a video to inspect through frames, audio, or transcripts. That is very different from sampling a video at a fixed FPS and hoping the important moment made it into context.</p><p class="paragraph" style="text-align:left;">It means the model can skim, search, zoom in, rewatch, and pull evidence only when the question needs it.</p><p class="paragraph" style="text-align:left;">This feature cuts token usage by up to 88%, lowers cost by up to 66%, and improves accuracy by up to 7%. Now available through the Gemini API in AI Studio and Gemini Enterprise Agent Platform.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Introducing agentic video understanding with Gemini</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Anthropic launched Claude Fable 5.1 and Mythos 5.1.</b> Fable 5.1 is the generally available model for coding and knowledge work, while Mythos 5.1 is the restricted version for trusted cybersecurity and life-sciences programs. Anthropic is also cutting cache-read pricing by 75%, which matters for long agent runs that keep reusing the same repo, docs, and tool context.</p><p class="paragraph" style="text-align:left;">Cursor already added Fable 5.1 and says it is their strongest model on CursorBench 3.2, scoring 73.4% at max effort.<br><a class="link" href="https://www.anthropic.com/claude-fable-and-mythos-5-1?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Claude Fable 5.1 and Mythos 5.1</a> | <a class="link" href="https://x.com/cursor_ai/status/2094852929282879596?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Cursor announcement</a></p><p class="paragraph" style="text-align:left;"><b>Dr. Fei-Fei Li’s World Labs introduced Atlas, a world model that cares about camera control</b>. Atlas works across text, images, video, and 3D, and can generate video with precise camera paths instead of vague “pan left” prompting. It can also reconstruct scenes from sparse images or video into point clouds and Gaussian splats, which is the part to watch for robotics, VFX, and game tooling.<br><a class="link" href="https://www.worldlabs.ai/blog/atlas?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Atlas by World Labs</a></p><p class="paragraph" style="text-align:left;"><b>Perplexity’s Hybrid Compute is a great application of “local when it matters, cloud when it helps.”</b> Hybrid Compute lets the app split a task between cloud models and a local model on your Mac. Web research can go to the cloud, but private files, app steps, and PII-heavy work can stay on-device, with an open-sourced classifier deciding what leaves the machine.<br><a class="link" href="https://www.perplexity.ai/hub/products/hybrid-compute?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Perplexity Hybrid Compute</a></p><p class="paragraph" style="text-align:left;"><b>Shopify open-sourced Tangle for building ML pipelines visually.</b> It gives teams a drag-and-drop editor for experiments, reusable pipeline components, collaborative runs, and caching so repeated steps don’t waste compute. The nice part: it makes the loop inspectable, so agents are not just “trying stuff” in a chat window.<br><a class="link" href="https://tangleml.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Tangle</a></p><p class="paragraph" style="text-align:left;"><b>Alibaba refreshed Qwen3.8-Max for bigger coding and agent runs</b>. <code>qwen3.8-max-0902</code> is a 2.4T-parameter model with a 1M-token context window, stronger coding/cowork post-training, and better tool orchestration for long-horizon projects. It is live on QwenCloud at $2/M input and $6/M output tokens, with cheaper cache reads.<br><a class="link" href="https://www.qwencloud.com/models/qwen3.8-max-0902?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Qwen3.8-Max-0902</a></p><p class="paragraph" style="text-align:left;"><b>GojiberryAI open-sourced a Sales OS for Grok Bot</b>. It turns Grok into a 13-agent outbound team for finding intent signals, checking ICP fit, researching accounts, drafting LinkedIn messages, handling replies, and qualifying meetings. It can run autonomously, but starts in “show me the list before anyone is contacted” mode, which is exactly where sales agents should start.<br><a class="link" href="https://github.com/romangojiberryAI/gojiberryai-sales-os?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">GojiberryAI Sales OS</a></p><p class="paragraph" style="text-align:left;"><b>GitHub CLI can now attach screenshots and videos without opening GitHub</b>. Their latest release adds a repeatable <code>--attach</code> flag for issues, PRs, and comments. This is a small CLI change that fixes a very real agent annoyance: stop describing the broken UI, attach the screenshot.<br><a class="link" href="https://github.blog/changelog/2026-09-01-github-cli-media-in-issues-pull-requests-and-comments/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">GitHub CLI media in issues, pull requests, and comments</a></p><p class="paragraph" style="text-align:left;"><b>Reducto shipped r-1, a cheaper parser for ugly documents.</b> The new model is built for the stuff that breaks other OCRs like dense tables, strikethroughs, watermarks, citations, and weird visual layouts. Reducto says r-1 preview cuts error rate 20% versus its legacy agentic OCR models and costs 1 cent per page.<br><a class="link" href="https://reducto.ai/blog/parse-r-1-model?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Introducing r-1</a></p><p class="paragraph" style="text-align:left;"><b>Hermes Agent v0.21.0 is basically agent society plumbing. </b>The Pantheon release adds bot-to-bot DMs across profiles and gateways, cron jobs with continuity and monitor-mode suppression, live steering for delegated subagents, JSON schema validation, and a much better MCP command center. If you are running multiple agents already, this is the release where handoffs start feeling less like duct tape.<br><a class="link" href="https://github.com/NousResearch/hermes-agent/releases/tag/v2026.8.31?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Hermes Agent v0.21.0</a></p><p class="paragraph" style="text-align:left;"><b>OpenClaw 2.0 is incredibly easy to use.</b> Cleaner setup that can reuse your existing ChatGPT/Claude/API/local model access, and a rebuilt browser app that opens straight into a working agent workspace. Other interesting things include shared cloud sessions, so someone else can join or take over live agent work without losing the context. A lot of other QoL updates!<br><a class="link" href="https://openclaw.ai/blog/openclaw-2-accidentally?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">OpenClaw 2.0, Accidentally</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>This is one of the best agentic engineering setups you’ll read today</b> for working with multiple coding agents. David Ondrej is running agents across bb, cmux, Ghostty, Herdr, VPS boxes, and subscription plans, then prioritizing finished work with Corral. You probably do not need the whole stack, but if your current workflow is “check five agent windows randomly,” this is worth skimming.<br><a class="link" href="https://x.com/DavidOndrej1/status/2094424967345496191?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">David Ondrej’s agentic engineering setup</a></p><p class="paragraph" style="text-align:left;"><b>Poteto published the workflow stack she uses to keep Grok Bot coding runs from turning into slop</b>. Lauren Tan works on Grok Bot, and her pstack plugin for Grok Bot gives your agents reusable skills for verification, feature maps, playbooks, and multi-model review. If you are trying to run more coding agents without drowning in sloppy diffs, this is worth checking.<br><a class="link" href="https://x.ai/bot/plugin/9717366?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">pstack for Grok Bot</a> | <a class="link" href="https://github.com/cursor/plugins/tree/main/pstack?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">pstack repo</a></p><p class="paragraph" style="text-align:left;"><b>A new paper takes aim at the laziest safety phrase: “human in the loop.” </b>Researchers argue that human oversight is not a checkbox, especially when agents hide intermediate steps, produce too much to review, and slowly train users to stop checking. <br><a class="link" href="https://arxiv.org/abs/2608.23642?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">AI Agents Push Humans Out of the Loop</a></p><p class="paragraph" style="text-align:left;"><b>Simon Willison wrote the missing manual for ChatGPT Work.</b> According to him “it&#39;s a deeply confusing but extremely powerful tool with a whole lot of useful features that aren&#39;t available in regular ChatGPT”<br><a class="link" href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Understanding ChatGPT Work</a></p><p class="paragraph" style="text-align:left;"><b>TinyFish gave DeepSeek Harness free Search and Fetch</b>. Install the TinyFish CLI and DSH picks it up automatically, which means agents can search the web and fetch clean page text for free when stale context is not enough.<br><a class="link" href="https://x.com/Tiny_Fish/status/2094838868499714441?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">TinyFish announcement</a> | <a class="link" href="https://docs.tinyfish.ai/cli?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">TinyFish CLI</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Diffusion Studio is an open-source video editor built for agents</b>. The simplest way to think about it: an IDE, but it renders a video canvas instead of text. Every edit is code, so an agent can cut, tweak, and reuse video edits without handing you one frozen MP4.<br><a class="link" href="https://github.com/diffusionstudio/editor?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Diffusion Studio Editor</a></p><p class="paragraph" style="text-align:left;"><b>Obscura is a headless browser for agents</b>. It gives agents a browser they can drive through CLI or MCP, with tools for clicks, forms, screenshots, PDFs, JavaScript, network logs, and scraping. Worth trying if your agent keeps getting stuck on websites that plain fetch cannot handle.<br><a class="link" href="https://github.com/h4ckf0r0day/obscura?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Obscura</a></p><p class="paragraph" style="text-align:left;"><b>Rakazo is an open-source Grok Bot-style app</b>. You get persistent AI teammates with their own memory, routines, browser, terminal, files, and computer access. You can bring your own models and run the stack yourself instead of depending on one hosted agent app.<br><a class="link" href="https://github.com/elie222/rakazo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Rakazo</a></p><p class="paragraph" style="text-align:left;"><b>ECC is a giant repo of agent skills, configs, and workflows</b>. It has setup files for Claude Code, Codex, OpenCode, Cursor, Gemini, Hermes, Qwen, and more. Would treat it as an awesome reference library first, not something to drop blindly into a real repo.<br><a class="link" href="https://github.com/affaan-m/ECC?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">ECC</a></p><p class="paragraph" style="text-align:left;"><b>headcount turns Claude Code into an agent org chart</b>. Instead of one giant coding agent, it gives you departments like engineering, security, finance, legal, growth, and support, each with its own skills. <br><a class="link" href="https://github.com/cbrock84/headcount?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">headcount</a></p><p class="paragraph" style="text-align:left;"><b>VoiceStudio is a local-first ElevenLabs alternative.</b> It runs voice cloning, dubbing, dictation, transcription, and audiobook workflows across Mac, Windows, and Linux. No account or API key needed for the core workflow<br><a class="link" href="https://github.com/debpalash/VoiceStudio?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">VoiceStudio</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Together AI cut dedicated H100 inference from $5.49/hr to $3.99/hr for September.</b> The discount applies automatically to new and existingdedicated inference deployments, so teams already running H100 endpoints get the lower bill without moving anything. <br><a class="link" href="https://x.com/togethercompute/status/2094583517015376237?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Together AI announcement</a></p><p class="paragraph" style="text-align:left;"><b>A 2B model was trained from Qwen2-1.5B on a single RTX 5090 inside a $5,090 budget</b>. Puro-2B is an open recipe for small-model pretraining on consumer GPUs, with data, code, and weights released under Apache 2.0. The caveat is that the budget and Qwen2-1.5B comparison come from the paper’s own evaluation setup.<br><a class="link" href="https://huggingface.co/papers/2608.27370?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Puro-2B paper</a></p><p class="paragraph" style="text-align:left;"><b>Someone benchmarked Qwen3.8-Flash-Next across the whole local hardware ladder.</b> In llama.cpp, the model goes from 8.34 tok/s on CPU-only to 109.07 tok/s on a 96GB VRAM setup. The ceiling comes from an RTX 6000 PRO-class card, so this is a hardware map, not a casual laptop win.<br><a class="link" href="https://www.reddit.com/r/LocalLLaMA/comments/1w3pl64/qwen38flashnext_in_llamacpp_from_cpuonly_to_96gb/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow">Qwen3.8-Flash-Next llama.cpp thread</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That&#39;s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=agentic-video-understanding-in-gemini"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="cut-lead-review-from-hours-to-minut">Cut Lead Review From Hours To Minutes</h3><div class="image"><a class="image__link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_eee8741e-0d90-4013-8d3c-3f8f9806ea99_7395cee5&bhcl_id=5ecb84a9-4646-431a-8184-c70fed9729e8_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7843cb04-d0c8-42e8-aeda-6fae54e84ed9/Attio_banner_1.png?t=1782837603"/></a></div><p class="paragraph" style="text-align:left;">Sign up for a free trial of <a class="link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_eee8741e-0d90-4013-8d3c-3f8f9806ea99_7395cee5&bhcl_id=5ecb84a9-4646-431a-8184-c70fed9729e8_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Attio</a>, the agentic CRM.</p><p class="paragraph" style="text-align:left;">Ask Attio to build a daily workflow that surfaces the deals that need your attention today, like anything with a stage change, a recent reply, or a new signal in the last 24 hours.</p><p class="paragraph" style="text-align:left;">Review your pipeline in Claude, synced live from Attio via MCP.</p><p class="paragraph" style="text-align:left;">That&#39;s it.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_eee8741e-0d90-4013-8d3c-3f8f9806ea99_7395cee5&bhcl_id=5ecb84a9-4646-431a-8184-c70fed9729e8_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Try Attio Now</a></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=ed1b0053-380a-45c4-9271-d23e3dc53198&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Kimi K3 Runs on One CPU With 8 GB of RAM </title>
  <description>+ No OpenAI models in Cursor from Nov 12</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b2d05535-9932-4a78-8786-d7a71ecac7b9/ChatGPT_Image_Aug_30__2026__11_47_34_PM.png" length="1298062" type="image/png"/>
  <link>https://www.theunwindai.com/p/kimi-k3-runs-on-one-cpu-with-8-gb-of-ram</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/kimi-k3-runs-on-one-cpu-with-8-gb-of-ram</guid>
  <pubDate>Mon, 31 Aug 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-08-31T12:30:00Z</atom:published>
    <dc:creator>Gargi Gupta</dc:creator>
    <dc:creator>Shubham Saboo</dc:creator>
    <category><![CDATA[Daily Unwind]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;">A 2.78T-parameter model ran on one CPU with 8.24 GB of RAM.</p><p class="paragraph" style="text-align:left;">kimi-k3-in-c runs Kimi K3, all 2.78 trillion parameters, on one CPU, and the engine that does it is 176 KB of portable C99, with no BLAS, framework, or GPU.</p><p class="paragraph" style="text-align:left;">That sounds fake until you see the tradeoff. The model still needs a 1.56 TB checkpoint on disk. Also, it is painfully slow: 26.5 seconds per token at 8 GB, 19.8 at 64 GB, 5.6 once 128 GB holds the whole thing. </p><p class="paragraph" style="text-align:left;">Nobody is serving traffic from this. That is not the point. It turns the “you need a cluster for this” assumption into a concrete systems question: which bytes must be in memory, which bytes can sit on disk, and how much speed are you willing to trade for access?</p><p class="paragraph" style="text-align:left;">If you like tiny, readable systems code, this is the rabbit hole to dig.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/FareedKhan-dev/kimi-k3-in-c?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">FareedKhan-dev/kimi-k3-in-c</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Warp launched Skill Doctor, a way to improve agent skills from past sessions. </b>Skill Doctor reads old Claude Code, Codex, and Warp transcripts, scores where the agent did well or wasted time, then proposes diffs to your skill files. It’s going from “write a better prompt” to “learn from the mess your agent already made.”<br><a class="link" href="https://warp.dev/skill-doctor?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Skill Doctor: score and improve your agent skills | Warp</a></p><p class="paragraph" style="text-align:left;"><b>Tencent dropped a 770B open model that might actually be servable.</b> Hy4-preview is a 770B-parameter MoE with 49B active parameters per token and a 1M-token context window. That active-parameter count is the interesting part: it puts the model in the “huge on paper, less terrifying to run” category. Tencent released it under Apache 2.0, so teams can test it without immediately hitting a commercial-use wall.<br><a class="link" href="https://huggingface.co/tencent/Hy4-preview?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">tencent/Hy4-preview · Hugging Face</a></p><p class="paragraph" style="text-align:left;"><b>Firecrawl made web search and scraping usable before signup, no API key required</b>. You get 1,000 free credits per month, across MCP, CLI, and REST API. Specially great for prototypes, workshops, and agent demos.<br><a class="link" href="https://x.com/ericciarla/status/2093375835679977570?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Introducing Firecrawl Keyless</a></p><p class="paragraph" style="text-align:left;"><b>GLM-5.3 is now fine-tunable through Thinking Machine’s Tinker API.</b> Tinker is Thinking Machines’ training API for LoRA fine-tuning, SFT, RL, DPO, and distillation. The model is currently the strongest open-weights model on coding evals like Terminal-Bench 3.0 and DeepSWE 1.1, and great for teams with their own datasets.<br><a class="link" href="https://tinker-docs.thinkingmachines.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">tinker-docs.thinkingmachines.ai</a></p><p class="paragraph" style="text-align:left;"><b>fal is giving free MiniMax H3 Max video generations (five videos a day).</b> The model can generate a 5-second 768p clip in under 3 seconds, with free daily generations and no signup. If you work with video models, this is worth testing with your own prompts instead of trusting the cherry-picked examples.<br><a class="link" href="https://fal.ai/minimax-h3-max?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">MiniMax H3 Max</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>OpenAI is winding down model access for Cursor after SpaceX’s acquisition</b> starting November 12, 2026. OpenAI says the cancellation window comes from a change-of-control clause, and that future models including Astra will not be provided. Cursor CEO Michael Truell <a class="link" href="https://x.com/mntruell/status/2093532254006063557?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">responded</a> that OpenAI models serve about 5% of Cursor user traffic and that talks are still ongoing.<br><a class="link" href="https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">openai.com</a></p><p class="paragraph" style="text-align:left;"><b>Sometimes quiet failures happen when relevant context gets treated like authority, not prompt injection. </b>A file, snippet, or prior chat can help an agent understand a task, but it should not automatically get to decide what is true or allowed. The post shows three failures: sibling files steering code repairs, copied evidence being counted as independent sources, and conversational trust being mistaken for permission. The fix is provenance: track source, lineage, trust, and authorization before the agent acts.<br><a class="link" href="https://sunglasses.dev/blog/ai-agent-context-security-provenance-authorization?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">AI Agent Context Security Needs Provenance</a></p><p class="paragraph" style="text-align:left;"><b>Warp has run 10M Claude Code sessions inside Warp</b>, with 400K+ sessions per week and 800K monthly developers. The interesting part is Warp’s loop: users give feedback where work happens, then a separate agent turns that feedback into small skill edits. That is much better than letting every correction vanish when the session ends.<br><a class="link" href="https://claude.com/blog/how-warp-builds-self-improving-agents-on-claude?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">How Warp builds self-improving agents on Claude</a></p><p class="paragraph" style="text-align:left;"><b>TurnBench measures the part of voice agents that gets edited out of demos. </b>Sesame’s TurnBench evaluates when a voice system should speak, wait, or treat speech as an interruption. The benchmark uses 30 hours of dual-channel human conversations across 154 dialogues and 106 actors. The leaderboard shows some systems detect more turn endings but interrupt too often. While others stay careful but respond slowly. <br><a class="link" href="https://turnbench.sesame.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">TurnBench</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>OpenKB is a vectorless RAG</b> that turns raw documents into a wiki-style knowledge base and retrieves through PageIndex reasoning instead of embedding search. Not every knowledge system has to start with vector search. If your current RAG stack is brittle on source structure, this is worth trying on a small corpus.<br><a class="link" href="https://github.com/VectifyAI/OpenKB?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">OpenKB</a></p><p class="paragraph" style="text-align:left;"><b>only-cli turns websites into tiny CLIs for agents</b> so an agent can use it in hundreds of tokens instead of reading full pages. It also gets past blocks that stop naive fetchers on some sites. If page views are eating your context window, turning messy web UI into a small command surface is one of the cleaner ideas to test.<br><a class="link" href="https://github.com/only-cli/oc?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">only-cli</a> </p><p class="paragraph" style="text-align:left;"><b>Fuxi is a terminal coding agent that makes model routing part of the workflow instead of a hidden backend choice.</b> It can edit files, run shell commands, use tools, and switch across LLM providers while tracking cost. The useful thing to inspect is its routing behavior: which steps go to cheaper models, which need stronger ones, and how much the run costs as it works. <br><a class="link" href="https://github.com/fuxicodex/Fuxi?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Fuxi</a></p><p class="paragraph" style="text-align:left;"><b>Repomix is for the step right before you paste a whole repo into an agent.</b> It packs a codebase into a single agent-readable bundle with file structure, selected source content, AST-aware compression, XML-style output, and Secretlint scanning. That gives the model enough project context without dumping every raw file into the prompt. <br><a class="link" href="https://github.com/yamadashy/repomix?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Repomix</a></p><p class="paragraph" style="text-align:left;"><b>Google’s Chrome DevTools MCP gives frontend agents the browser state they have to guess from screenshots or logs.</b> It lets coding agents inspect Chrome DevTools directly: console errors, network requests, DOM state, screenshots, and performance traces. That means an agent can edit the code, reload the app, inspect the real browser, and keep debugging. <br><a class="link" href="https://github.com/ChromeDevTools/chrome-devtools-mcp?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">chrome-devtools-mcp</a></p><p class="paragraph" style="text-align:left;"><b>OpenBot is an open-source alternative to Grok Bot.</b> These are AI coworkers you can hand real work to and trust with access. Each gets a computer of its own: a real browser with its own logins, its own files, and only the tools you grant. Every action is decided before it happens and recorded after.<br><a class="link" href="https://github.com/CopilotKit/OpenBot?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">OpenBot</a></p><p class="paragraph" style="text-align:left;"><b>GenOffice is a free, open-source office suite for macOS, Windows & Linux with AI agents built in.</b> It can edit Word docs, spreadsheets, presentations, PDFs, and Markdown, which makes it more practical than tools that only handle clean text. Worth a look if your automation work involves messy office files.<br><a class="link" href="https://github.com/genspark-ai/genoffice?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">genoffice</a></p><p class="paragraph" style="text-align:left;"><b>geo-seo-claude audits how a site appears across ChatGPT, Claude, Perplexity, and Google AI Overviews.</b> It checks citation scores, AI crawler access, schema markup, brand authority, and platform-specific optimization, then produces reports. “Are AI answers citing us?” is still fuzzy for most teams, and this turns it into a concrete agent job.<br><a class="link" href="https://github.com/zubair-trabzada/geo-seo-claude?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">geo-seo-claude</a></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow"><b>Awesome LLM Apps</b></a><b> is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Claude Code’s permanent limit bump is smaller than the temporary one users have right now</b>. They are increasing their standard weekly Claude Code limits permanently by 25% for Pro, Max, Team, and seat-based Enterprise plans starting September 14. Until then, the current 50% temporary increase remains. Compared with today’s temporary level, the permanent level is about 17% lower. So if Claude Code feels unusually roomy this week, do not use that as your long-term baseline.<br><a class="link" href="https://x.com/ClaudeDevs/status/2093742321473065266?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">ClaudeDevs</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI is resetting paid Codex and ChatGPT Work usage after fixing several token-burn bugs</b> across image compaction, memory workers, goals, automations, subagents, computer history, rolling summaries, and MCP result handling. You may see usage go 10% to 50% further depending on workflow. <br><a class="link" href="https://x.com/thsottiaux/status/2093801758665715784?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Tibo</a></p><p class="paragraph" style="text-align:left;"><b>Terminal-Bench 4.0 makes cost part of the benchmark, not a side conversation. </b>After publishing a new version of the benchmark’s dataset and leaderboard, Terminal-Bench 4.0 now reports resolution rate, tokens, and cost together. Opus 5 leads at 51.8% resolution with 6.5B tokens and about $6.0K. GLM-5.3 reaches 41.8% with 8.7B tokens and about $2.7K. Fable 5 sits at 44.5%, with overlapping error bars versus GLM-5.3, but higher reported cost.<br><a class="link" href="https://www.tbench.ai/news/terminal-bench-4-0?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Terminal-Bench 4.0</a></p><p class="paragraph" style="text-align:left;"><b>Anthropic is giving scientists a cheaper Claude path, with eligibility limits. </b>Anthropic opened 10,000 free standard Claude seats for scientists for one year. Premium seats with 5x usage limits are $15 per month. The program is gated to verified principal investigators or equivalents at academic or nonprofit research institutions.<br><a class="link" href="https://www.anthropic.com/news/expanding-support-for-scientists?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow">Expanding our support for scientists</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That&#39;s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2"><b>X</b></a></span> | <span style="text-decoration:underline;"><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2"><b>LinkedIn</b></a></span><b> </b>|<b> </b><span style="text-decoration:underline;"><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2"><b>Threads</b></a></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2"><b>Awesome LLM Apps</b></a></span><b> | </b><span style="text-decoration:underline;"><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2"><b>Sponsor Us</b></a></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=kimi-k3-runs-on-one-cpu-with-8-gb-of-ram"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="stop-paying-for-10-tools-one-ai-doe">Stop Paying for 10 Tools. One AI Does It All.</h3><div class="image"><a class="image__link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_5889376d-16fe-4922-aa8c-7a8eb2f4acbe_d6ea45bd&bhcl_id=2c39ab16-4a8d-4ee3-bb53-de3b939a607d_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/95732623-bcd7-400f-8a89-857ba744f10e/SC-boost_efficiency-v5.jpg?t=1784134849"/></a></div><p class="paragraph" style="text-align:left;">Most e-commerce sellers are running their store across 6 to 10 separate tools — and spending more time managing software than growing their business. <a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_5889376d-16fe-4922-aa8c-7a8eb2f4acbe_d6ea45bd&bhcl_id=2c39ab16-4a8d-4ee3-bb53-de3b939a607d_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">StoreClaw</a> replaces your entire stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks real profit across Shopify, Amazon, and beyond.</p><p class="paragraph" style="text-align:left;">It doesn&#39;t wait for you to ask. It runs 24/7 in the background, so you wake up to a full dashboard instead of a list of things you forgot to check.</p><p class="paragraph" style="text-align:left;">Connect your store, and <a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_5889376d-16fe-4922-aa8c-7a8eb2f4acbe_d6ea45bd&bhcl_id=2c39ab16-4a8d-4ee3-bb53-de3b939a607d_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">StoreClaw</a> gets to work — no prompts, no complex setup, no six-app stack.</p><p class="paragraph" style="text-align:left;">Free to start. No credit card required.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_5889376d-16fe-4922-aa8c-7a8eb2f4acbe_d6ea45bd&bhcl_id=2c39ab16-4a8d-4ee3-bb53-de3b939a607d_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Try it for FREE today</a></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=2fe85279-a734-419f-95c1-04f06555e3f6&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>THE OFFICE Agent Harness</title>
  <description>+ $13B Hugging Face acquisition by Nvidia</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/05fddaa8-2541-4e78-b9e0-8b2c241a693b/exec-8f658dea-652b-4257-aebf-576fb544044e.png" length="956337" type="image/png"/>
  <link>https://www.theunwindai.com/p/the-office-agent-harness</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/the-office-agent-harness</guid>
  <pubDate>Fri, 28 Aug 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-08-28T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;">Someone just built <b>THE OFFICE</b> theme agent harness and made it open-source.</p><p class="paragraph" style="text-align:left;"><b>Munder Difflin</b> wraps the CLI agents you already use, like Claude Code, Codex, Gemini CLI, Kimi, Grok, OpenCode, gives them memory, wires them into a hive mind, and puts your clone in charge. Michael (ofc had to be him) is the one you talk to to get things done. </p><p class="paragraph" style="text-align:left;">It works with subscriptions and keys you already have, keeps the local version on your machine, and lets agents coordinate without everyone pushing into the same shared mess. <br><a class="link" href="https://github.com/chaitanyagiri/munder-difflin?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/chaitanyagiri/munder-difflin</a></p><p class="paragraph" style="text-align:left;">Also, do not skip By the Numbers today 🤫 </p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Cursor Cloud Agents can now start without a repo</b>. You can prompt a new web app from scratch, preview it in the browser, then save the code to a Cursor Origin repo when it is worth keeping. If you connect Vercel, Cursor can also publish the app to a live URL.<br><a class="link" href="https://cursor.com/changelog/start-from-scratch?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">cursor.com</a></p><p class="paragraph" style="text-align:left;"><b>Vercel made WebGPU easier for agents to test</b>. vgpu lets agents validate shader code in headless Node.js and CI, even when the sandbox does not have a GPU. Useful if you want agents working on visual code without relying on manual screenshot checks.<br><a class="link" href="http://vgpu.sh?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">vgpu.sh</a></p><p class="paragraph" style="text-align:left;"><b>The Claude Code team fixed a quiet cache bug that could cost money in long sessions</b>. The changelog says tool definitions were being re-rendered after OAuth token refreshes, causing a prompt-cache miss roughly once an hour. If you run long agent sessions, that is the kind of invisible leak worth upgrading for.<br><a class="link" href="https://code.claude.com/docs/en/changelog?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">code.claude.com</a></p><p class="paragraph" style="text-align:left;"><b>Google made Gemini Omni cheaper to iterate with. </b>Gemini Omni 1.1 Flash can now generate rough 360p video drafts before you pay for a higher-resolution render. These drafts are up to 60% faster and cost 1/3rd as much as 720p, so you can test a bunch of directions quickly, then upscale the one you want to keep. <br><a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">blog.google</a></p><p class="paragraph" style="text-align:left;"><b>Firecrawl added OCR to anydoc for scanned docs. </b>anydoc already turns office files and text-based PDFs into clean Markdown for agents. The new Firecrawl OCR option covers scanned pages too, with sub-5ms handling for docs that do not need OCR and 190ms median per OCR page. It is free to use with no API key.<br><a class="link" href="https://github.com/firecrawl/anydoc?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/firecrawl/anydoc</a></p><p class="paragraph" style="text-align:left;"><b>Agno, the open-source AI agent framework, released its v3.0</b>, shipping with an SDK, runtime, and control plane together. So teams can build, run, and monitor agents in the same stack. Worth a look if you are already comparing agent frameworks for production use.<br><a class="link" href="http://docs.agno.com?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">docs.agno.com</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>The best RAG stack might still start with plain search</b>. Rafael Pierre’s RAG breakdown is a good antidote to overbuilt retrieval systems: start with BM25, add query rewriting, then move to hybrid or pre-embedding only when the data proves you need it. Remember the 80/20 rule: do not build the 5% solution for a 60% problem.<br><a class="link" href="https://www.lighthousenewsletter.com/p/rag-is-simpler-than-you-think?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">lighthousenewsletter.com</a></p><p class="paragraph" style="text-align:left;"><b>Anthropic is testing a standard for agents controlling lab hardware</b>. The Model Hardware Standard is a research preview for letting agents operate programmable devices like microscopes, liquid handlers and robotic arms through shared driver primitives. They say that hardware integrations that usually take weeks or months could drop to hours or minutes.<br><a class="link" href="https://www.anthropic.com/news/model-hardware-standard-research-preview?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">anthropic.com/news</a></p><p class="paragraph" style="text-align:left;"><b>Alibaba’s latest model Qwen3.8-27B looks fine at 4-bit, but falls apart at 1-bit. </b>Quesma benchmarked several GGUF quantizations and found the 17GB <code>Q4_K_M</code> version holds up surprisingly well, while 1-bit collapses to random-chance territory on GPQA Diamond. If you are picking a local coding model for a 24GB card, this is a practical result, not a philosophical debate about quantization.<br><a class="link" href="https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">quesma.com</a></p><p class="paragraph" style="text-align:left;"><b>Someone measured Claude-ish vocabulary across 47,000+ GitHub PRs.</b> A word cluster that didn&#39;t exist in 2025 is now 45% of human-authored PRs. And the top word is &quot;load-bearing&quot;<br><a class="link" href="https://louisabraham.github.io/load-bearing/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">louisabraham.github.io/load-bearing</a></p><p class="paragraph" style="text-align:left;"><b>Anthropic says code review cannot stay line-by-line.</b> Their latest AI-native SDLC playbook argues that once agents write large chunks of code, the bottleneck moves to planning, testing, security, and deployment. So the question every team has to answer now: what replaces human line-by-line review when the diff is too big to read the old way?<br><a class="link" href="https://claude.com/blog/the-ai-native-sdlc-playbook?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">claude.com/blog</a></p><p class="paragraph" style="text-align:left;"><b>Terminal-Bench-Science is trying to measure agents on real scientific work.</b> The benchmark uses workflows contributed by working scientists, then grades concrete artifacts like analyses, simulations, proofs, code, and data products. The first release has 70 tasks, and Claude Opus 5 tops the leaderboard at only 30%, which is exactly why this is more useful than another easy eval.<br><a class="link" href="https://www.terminal-bench-science.ai/announcement?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">terminal-bench-science.ai/announcement</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Tare analyzes a Claude session and shows where the tokens actually went</b>. If your Claude Code quota disappears in ten minutes, attribution is more useful than another complaint thread. <br><a class="link" href="https://github.com/kelviq/tare?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">https://github.com/kelviq/tare</a><br><br><b>OpenSEO is an open-source alternative to Semrush and Ahrefs</b>. It exposes an MCP server so AI agents like Claude Code, OpenClaw, and Hermes can use your SEO data directly. Agent Skills are reusable workflows that guide your agent through SEO tasks using the MCP.<br><a class="link" href="https://github.com/every-app/open-seo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/every-app/open-seo</a></p><p class="paragraph" style="text-align:left;"><b>Concord is an MCP server for letting Claude Code, Codex and Cursor send messages to each other, live</b>. If your current multi-agent coordination layer is “write to a shared file and hope,” this is worth reading. <br><a class="link" href="https://github.com/Get-Concord-AI/concord-mcp?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/Get-Concord-AI/concord-mcp</a></p><p class="paragraph" style="text-align:left;"><b>Claude quickstarts now include a cookbook for running Claude Managed Agents with Vercel’s Chat SDK</b> and delivering them into Slack, WhatsApp, Discord and Teams. Useful if the thing you keep rebuilding is the chat delivery layer, not the agent itself. <br><a class="link" href="https://github.com/anthropics/claude-quickstarts?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/anthropics/claude-quickstarts</a></p><p class="paragraph" style="text-align:left;"><b>Experiential is an open-source model gateway for agent workflows</b>. It gives you one OpenAI-compatible API across hosted, BYOK and local models, plus controls for who can use which model and how much they can spend. <br><a class="link" href="https://github.com/experientiallabs/experiential?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">github.com/experientiallabs/experiential</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (134k+ </b>🌟<b> ) is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Nvidia agrees to buy Hugging Face for $12.9B.</b> No signed agreement yet, but it is big enough to make people nervous because Hugging Face is not just another AI company. It is where a lot of teams host models, datasets, Spaces, and demos. If the platform changed ownership, the question is less “is open source dead?” and more “how much of your workflow depends on one host?”<br><a class="link" href="https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">techcrunch.com</a></p><p class="paragraph" style="text-align:left;"><b>Thinking Machines is giving up to $50K in credits for open-weight safety research</b>. The grants are for projects using Tinker to study things like safer open models, hazardous-data filtering, tamper-resistant safety training and reward hacking. Good fit if you are doing actual experiments on open-weight model safety and need compute credits.<br><a class="link" href="http://thinkingmachines.ai/news/safety-research-grants/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow">thinkingmachines.ai</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-office-agent-harness"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="cut-lead-review-from-hours-to-minut">Cut Lead Review From Hours To Minutes</h3><div class="image"><a class="image__link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_8319cee5-69ab-4ef5-9672-d2cb8c717f75_7395cee5&bhcl_id=cbf3e1e8-0b9b-4552-a39e-0467acde664d_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7843cb04-d0c8-42e8-aeda-6fae54e84ed9/Attio_banner_1.png?t=1782837603"/></a></div><p class="paragraph" style="text-align:left;">Sign up for a free trial of <a class="link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_8319cee5-69ab-4ef5-9672-d2cb8c717f75_7395cee5&bhcl_id=cbf3e1e8-0b9b-4552-a39e-0467acde664d_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Attio</a>, the agentic CRM.</p><p class="paragraph" style="text-align:left;">Ask Attio to build a daily workflow that surfaces the deals that need your attention today, like anything with a stage change, a recent reply, or a new signal in the last 24 hours.</p><p class="paragraph" style="text-align:left;">Review your pipeline in Claude, synced live from Attio via MCP.</p><p class="paragraph" style="text-align:left;">That&#39;s it.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://attio.com/?utm_source=beehiiv&utm_medium=newsletter_sponsorship&utm_campaign=beehiiv-Y26&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_8319cee5-69ab-4ef5-9672-d2cb8c717f75_7395cee5&bhcl_id=cbf3e1e8-0b9b-4552-a39e-0467acde664d_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Try Attio Now</a></p></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=1dd9f06e-1d30-4219-a127-efd3dc235c92&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>OpenRouter for AI Agents </title>
  <description>+ Claude Code, Codex, Cursor, Hermes, Pi in one API</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/35fcf515-20a7-494b-958f-0c3f4034ec31/ChatGPT_Image_Aug_26__2026__10_09_24_PM.png" length="1090698" type="image/png"/>
  <link>https://www.theunwindai.com/p/openrouter-for-ai-agents</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/openrouter-for-ai-agents</guid>
  <pubDate>Thu, 27 Aug 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-08-27T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;"><b>Ox-Alpha was GLM-5.3-Flash all along</b></p><p class="paragraph" style="text-align:left;">The mystery model people were hammering on OpenRouter and OpenCode has a name now: GLM-5.3-Flash.</p><p class="paragraph" style="text-align:left;">Z.ai revealed that Ox-Alpha was its new model and released the weights the same day. It is a 320B-parameter multimodal MoE model with 18B active parameters, trained on a 30T-token multimodal corpus, and built with a hybrid sparse plus linear attention architecture to make long-context serving cheaper.</p><p class="paragraph" style="text-align:left;">The weights are on Hugging Face; the model card lists vLLM, SGLang, TokenSpeed, and KTransformers support, and Unsloth already has GGUF quantizations up. </p><p class="paragraph" style="text-align:left;">Z.ai says GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks while coming in far cheaper than premium frontier APIs. Treat the benchmark claims like you would any vendor chart, but the combination of open weights, multimodal input, long-context architecture work, and low API pricing makes this worth trying.<br><br><a class="link" href="https://huggingface.co/zai-org/GLM-5.3-Flash?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">zai-org/GLM-5.3-Flash · Hugging Face</a> <br><a class="link" href="https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">unsloth/GLM-5.3-Flash-GGUF · Hugging Face</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Qwen4 architecture. </b>Alibaba opened the weights for Qwen3.8-Flash-Next, a multimodal MoE model that’s an early preview of the architecture behind Qwen4. It supports 262K tokens natively and can stretch to 1M with YaRN. If you care about where open model architectures are going, this is more useful than another benchmark screenshot because Qwen is showing the machinery before the flagship model lands.<br><a class="link" href="https://qwen.ai/blog?id=qwen3.8-flash-next&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Qwen</a> <a class="link" href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Qwen/Qwen3.8-Flash-Next</a> | <a class="link" href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">unsloth</a></p><p class="paragraph" style="text-align:left;"><b>Google shipped Gemini 3.5 Transcribe, </b>a speech-to-text model in the Gemini API that can call other Gemini models mid-transcription to generate images or analyze files. It handles noise, jargon, fillers, live language switches, speaker attribution and word-level timestamps.<br><a class="link" href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Introducing Gemini 3.5 Transcribe</a> </p><p class="paragraph" style="text-align:left;"><b>Apple is turning the Mac Studio into a serious local AI box with M6 and M5 Ultra. </b>The new M5 Ultra supports up to 512GB of unified memory for running LLMs with hundreds of billions of parameters on device. The new M6 brings faster on-device AI to the Mac mini. If your bottleneck is “the model does not fit,” check this out.<br><a class="link" href="https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Apple introduces M6 and M5 Ultra</a></p><p class="paragraph" style="text-align:left;"><b>Firecrawl’s Developer Index is now live in Codex</b> through the OpenAI plugin marketplace. It gives Codex access to 70M+ primary sources across repos, docs, and issues, which is exactly the kind of context coding agents need when they hit an unfamiliar library. <br><a class="link" href="https://x.com/firecrawl/status/2092295030794760371?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">x.com/firecrawl</a></p><p class="paragraph" style="text-align:left;"><b>OpenComputer is basically “Firebase for agents”</b>: write an agent as a TypeScript function, deploy it, and each session gets a real Linux machine with shell, files, packages, browser, network, and MCP support. The good bit is persistence: sessions can stream, hibernate, and resume instead of starting from zero every time. <br><a class="link" href="https://opencomputer.dev/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Firebase for agents – OpenComputer</a></p><p class="paragraph" style="text-align:left;"><b>AgentSky launched an “OpenRouter for agents,” putting Claude Code, Codex, Hermes, DeepSeek Harness, Kimi Code, opencode, and more behind one API.</b> You choose the agent and the model in the same request, and AgentSky handles the cloud computer and session state. If you have been testing coding agents one by one, this makes the harness itself easier to swap.<br> <a class="link" href="https://agentsky.dev/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">AgentSky — One API, Every Agent</a></p><p class="paragraph" style="text-align:left;"><b>Monid is trying to do for tools what OpenRouter did for models</b>: one key, many providers, pay per call. Has 1,700+ tools across search, social, video, data, sales, and more, with prices shown before the agent picks one. That is useful because tool-heavy agents get expensive fast when every small workflow needs a new subscription.<br><a class="link" href="https://monid.ai/openrouter-for-agent-tools?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">The OpenRouter for agent tools </a><br><br><b>SandboxAQ open-sourced Switch for bringing agents into Slack, Teams, Discord, and other shared workrooms</b>. Agents and humans work in the same room, with the same history, rules, and context, instead of copying state across tools. It supports agents from Claude Code, Google ADK, LangChain, & OpenAI.<br><a class="link" href="https://www.sandboxaq.com/press/sandboxaq-open-sources-switch-bring-any-ai-agent-into-any-team-chat?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">SandboxAQ</a></p><p class="paragraph" style="text-align:left;"><b>Bezalel is a capability plane for agents</b>, exposed through one MCP URL. One setup gives Claude, Codex, OpenCode, Hermes, Cursor, or any MCP-speaking agent access to shared memory, email, iMessage, a cloud computer, sandboxes, hundreds of connectors, and soon virtual cards. <br><a class="link" href="https://bezalel.sh/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Bezalel — super powers for your agent</a></p><p class="paragraph" style="text-align:left;"><b>Superwhisper made its Whisper models free for all users</b>. You no longer need a Pro subscription for local, private voice-to-text, and existing Pro users got usage reset with 3,000 words for trying newer features. Nice timing, given Google shipped Gemini transcription the same day.<br><a class="link" href="https://x.com/superwhisper/status/2092660873311436832?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2092660873311436832%7Ctwgr%5E%7Ctwcon%5Es1_&ref_url=file%3A%2F%2F%2FUsers%2Fgargigupta%2F.hermes%2Fhermes-agent%2Fapps%2Fdesktop%2Frelease%2Fmac-arm64%2FHermes.app%2FContents%2FResources%2Fapp.asar.unpacked%2Fdist%2Findex.html%2F20260823_222920_1a1d05&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Superwhisper</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Switching models mid-session can wipe out your prompt-cache savings.</b> Hermes Agent creator, Teknium, says prompt caches are model-specific, so when you switch, you may repay full input-token cost for the same long context. Routing is good, but bouncing between models inside one bloated session is not free.<br><a class="link" href="https://x.com/Teknium/status/2092141955082019311?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">x.com/Teknium</a></p><p class="paragraph" style="text-align:left;"><b>Salesforce and Anthropic announced Claudeforce</b>, starting with a Salesforce in Claude plugin that has 37 prebuilt sales skills. The more interesting part is AIforce, which exposes Salesforce data, workflows, and business logic through MCP servers, APIs, and CLI tools. Most readers cannot install this tomorrow, but it shows MCP becoming enterprise plumbing. <br><a class="link" href="https://investor.salesforce.com/news/news-details/2026/Salesforce-and-Anthropic-Announce-Claudeforce-The-1-AI-Meets-the-1-AI-CRM/default.aspx?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Salesforce.com, Inc.</a></p><p class="paragraph" style="text-align:left;"><b>Your Claude Code allow list probably permits more than you meant. </b>The wildcard in a rule you wrote for one project stretches to cover the same command pointed at any folder on your machine, and Claude Code now flags those rules at startup, which is a good reason to reread the permissions you approved in a hurry.<br><a class="link" href="https://code.claude.com/docs/en/changelog?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Claude Code changelog</a></p><p class="paragraph" style="text-align:left;"><b>Do not test agent sandboxes on your real machine</b>. Trail of Bits says GPT-5.6-Cyber escaped a QEMU/KVM VM three times, eventually finding several 0-days. OpenAI’s Hugging Face incident showed agents using Artifactory as a message board, reaching the internet through SSRF, and touching real infrastructure. The boring rule wins: throwaway machines, no real keys, limited network, real logs. <a class="link" href="https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyber-capable-agents/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">VMs won&#39;t contain cyber-capable agents</a></p><p class="paragraph" style="text-align:left;"><b>Accept Markdown is a small idea that docs teams can ship today. </b>If a client sends <code>Accept: text/markdown</code>, serve a Markdown version, set <code>Vary: Accept</code>, return <code>406</code> for unsupported types, and honor q-values. Cleaner pages mean agents spend fewer tokens on nav, scripts, and layout junk. <br><a class="link" href="https://acceptmarkdown.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Serve Markdown to AI Agents with Accept Headers</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI introduced Premium seats for ChatGPT Business at $100 per month</b> when billed annually. The seat is aimed at teams that use ChatGPT, ChatGPT Work, Codex, connectors, admin controls, billing, security, and analytics as daily infrastructure. Tibo also says it removes the 5-hour limit, which is probably the line for Codex-heavy teams.<br><a class="link" href="https://x.com/OpenAI/status/2092335305366069305?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">x.com/OpenAI</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Headlong is an open-source agent microharness in under 10K lines of Bash</b>, built around a persistent thought loop. Instead of waking up only when you send a message, the agent keeps thinking, treats messages as observations, and decides when to reply. It is weird in the best way: less chatbot, more tiny always-on shell creature.<br><a class="link" href="https://github.com/laude-institute/headlong?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">laude-institute/headlong</a></p><p class="paragraph" style="text-align:left;"><b>Callstack’s agent-device lets coding agents inspect and verify running mobile apps</b> across iOS, Android, HarmonyOS, TV, web, macOS, and Linux. It uses accessibility snapshots and refs instead of screenshot guessing, then saves evidence you can replay in CI. <br><a class="link" href="https://github.com/callstack/agent-device?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">callstack/agent-device</a></p><p class="paragraph" style="text-align:left;"><b>Google’s Jot is the tiny Gemini 3.5 Transcribe demo</b> you can actually try. Hold <code>fn</code>, speak, and it writes cleaned-up text wherever your cursor is, using your own Gemini API key. <br><a class="link" href="https://github.com/google-gemini/jot-gemini-transcribe-macOS?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">google-gemini/jot-gemini-transcribe-macOS</a></p><p class="paragraph" style="text-align:left;"><b>scientific-agent-skills is a repo with 163 research skills</b> across genomics, chemistry, medicine, materials science, ML, geospatial work, and more. Even if you do not run science workflows, it is worth opening as a format reference for vertical skill packaging. <br><a class="link" href="https://github.com/K-Dense-AI/scientific-agent-skills?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">K-Dense-AI/scientific-agent-skills</a></p><p class="paragraph" style="text-align:left;"><b>AgentConnect is the open-source, multi-agent alternative to Claude Tag. </b><span style="background-color:#ffffff;">@ any agent and </span>your <span style="background-color:#ffffff;">teams and multiple AI agents work together across Slack, Telegram, Discord, Lark, GitHub, and GitLab. Use Claude Code, Codex, Grok Build, or any ACP-compatible agent in the chats and workflows your team is on.</span><br><a class="link" href="https://github.com/agentconnect-md/agentconnect?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">agentconnect-md/agentconnect</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (134k+ </b>🌟<b> ) is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Devin launched a startup program with $65K in credits and grants</b>. Approved startups get $15K in Devin credits upfront, plus up to $50K in matching grants on future usage. If coding agents are becoming a real budget line, this is Cognition trying to get early teams hooked before the bill becomes normal.<br><a class="link" href="https://devin.ai/startups?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">devin.ai/startups</a></p><p class="paragraph" style="text-align:left;"><b>Someone cloned OpenAI’s sold-out $230 Codex Micro for $35</b>. Iluvatar Labs built AgentPad13, an open-source macropad for coding agents with 13 hot-swappable keys, a rotary encoder, joystick, touch input, RGB lighting, and QMK/Vial firmware. The assembled electronics cost about $35, and even a full build comes in around $65 depending on the case, switches, and keycaps.<br><a class="link" href="https://iluvatarlabs.com/blog/2026/08/how-we-cloned-openai-codex-micro/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Iluvatar Labs</a></p><p class="paragraph" style="text-align:left;"><b>Artificial Analysis ranks Breeze TTS 2 as the top open-weight model for provider-style voices, ahead of Fish Audio S2 Pro</b>. It supports 50 languages, prompt-based voice generation, streaming, and self-hosting through Hugging Face. Fish still wins on hosted price and speed, so this is an audition-first model, not an automatic switch. Open-weight voice models are worth seriously shortlisting.<br><a class="link" href="https://x.com/ArtificialAnlys/status/2092399623839326550?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">x.com/ArtificialAnlys</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openrouter-for-ai-agents"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;" id="stop-paying-for-10-tools-one-ai-doe">Stop Paying for 10 Tools. One AI Does It All.</h3><div class="image"><a class="image__link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_e1eaa090-9664-49ad-8956-e86ce5beb0ba_d6ea45bd&bhcl_id=ef874e6d-6fdf-440e-b5ff-c70086923a02_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/95732623-bcd7-400f-8a89-857ba744f10e/SC-boost_efficiency-v5.jpg?t=1784134849"/></a></div><p class="paragraph" style="text-align:left;">Most e-commerce sellers are running their store across 6 to 10 separate tools — and spending more time managing software than growing their business. <a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_e1eaa090-9664-49ad-8956-e86ce5beb0ba_d6ea45bd&bhcl_id=ef874e6d-6fdf-440e-b5ff-c70086923a02_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">StoreClaw</a> replaces your entire stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks real profit across Shopify, Amazon, and beyond.</p><p class="paragraph" style="text-align:left;">It doesn&#39;t wait for you to ask. It runs 24/7 in the background, so you wake up to a full dashboard instead of a list of things you forgot to check.</p><p class="paragraph" style="text-align:left;">Connect your store, and <a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_e1eaa090-9664-49ad-8956-e86ce5beb0ba_d6ea45bd&bhcl_id=ef874e6d-6fdf-440e-b5ff-c70086923a02_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">StoreClaw</a> gets to work — no prompts, no complex setup, no six-app stack.</p><p class="paragraph" style="text-align:left;">Free to start. No credit card required.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.storeclaw.ai?utm_source=newsletter&utm_medium=beehive&utm_campaign=SC_boost_efficiencyv5&utm_content={{publication_alphanumeric_id}}&_bhiiv=opp_e1eaa090-9664-49ad-8956-e86ce5beb0ba_d6ea45bd&bhcl_id=ef874e6d-6fdf-440e-b5ff-c70086923a02_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Try it for FREE today</a></p></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=e83454d4-7f95-442a-aa50-670c0bc12cc4&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Microsoft&#39;s new skill tunes your agents in Claude Code, Codex, and Cursor</title>
  <description>+ Free tiers of 28 LLM providers in one API</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6f0d82fb-1a67-422b-8173-978e322c639f/upload_d7925aa71465965e.jpg" length="87669" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor</guid>
  <pubDate>Tue, 25 Aug 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-08-25T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;"><b>Your coding agent can now tune and optimize your other agents.</b> Microsoft&#39;s Agent Lightning v1.0.1 installs as a skill in Claude Code, Codex, GitHub Copilot, or any other agent harness. </p><p class="paragraph" style="text-align:left;">Give it an editable agent and a benchmark, and it reworks prompts, tools, workflows, model choice, and reasoning settings against scored results instead of vibes. </p><p class="paragraph" style="text-align:left;">Every change is measured, so the eval you keep meaning to write is the price of entry, and writing it is how you finally learn whether last Tuesday&#39;s prompt edit helped.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://github.com/microsoft/agent-lightning/releases/tag/v1.0.1?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/microsoft/agent-lightning</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Xiaomi announced the AI Cube, a desktop machine specified at 1.2 TB/s memory bandwidth.</b> Local serving hits a memory-bandwidth wall long before a FLOPs wall, so that is the number worth watching on a box like this. No first-party page resolved, so every spec is the thread&#39;s claim.<br><a class="link" href="https://www.ithome.com/0/993/546.htm?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">ithome.com</a></p><p class="paragraph" style="text-align:left;"><b>Headless Tools launched SaaS tools with no user interface at all, built to be called by agents directly.</b> Everyone else is bolting an agent API onto a product designed for a human; this one starts from the agent. <br><a class="link" href="https://hdls.tools?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">hdls.tools</a><br><br><b>Grok 4.6 is now available inside Hermes at 50% off through the Nous Research portal </b>for the next week. You can also use an existing SuperGrok or X Premium+ subscription. If you use Hermes for longer agent runs, this is a cheap week to test whether Grok earns a slot in your routing table.<br><a class="link" href="https://x.com/SpaceXAI/status/2091957125941543034?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">x.com/SpaceXAI</a></p><p class="paragraph" style="text-align:left;"><b>Dactyl builds native mobile apps from a description, runs them live in a browser simulator, and can ship to TestFlight without a Mac or Xcode</b>. Comes with live simulator: Apple sign-in, camera, sensors, and payments can work while the app is still being generated. <br><a class="link" href="https://dactyl.dev/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">dactyl.dev</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>OpenAI cut GPT-5.6 Luna pricing by 80% and Terra pricing by 20%</b> across the API, ChatGPT Work, and Codex. Luna is now $0.20 / $1.20 per million input/output tokens, and Terra is $2 / $12, so the cheap end of the GPT-5.6 family is now much cheaper for agent loops. <br><a class="link" href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">openai.com</a> </p><p class="paragraph" style="text-align:left;"><b>Claude Team and admins can now authorize MCP connectors once for the whole org through their identity provider</b>. This fixes the annoying rollout problem where every user had to do their own OAuth flow before Claude could use the same tools. For companies, MCP just got easier to deploy and audit.<br><a class="link" href="https://support.claude.com/en/articles/15537633-authorize-mcp-connectors-for-your-entire-organization?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">support.claude.com</a></p><p class="paragraph" style="text-align:left;"><b>Your agent does not have to wait for the model to finish before starting the slow work.</b> Speculative Programmatic Tool Calling predicts likely calls from partially written code, then launches search, sub-agents, or APIs early so they run alongside token generation. Worth testing if tool latency is the bottleneck in your harness.<br><a class="link" href="https://alexzhang13.github.io/blog/2026/spec-ptc/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">alexzhang13.github.io</a></p><p class="paragraph" style="text-align:left;"><b>Artificial Analysis launched benchmarks for phone-sized models across intelligence and real mobile-device inference. </b>Great for comparing small models on tool use, reasoning, speed, latency, and memory instead of relying on vendor claims or “works on my phone” demos.<br><a class="link" href="https://artificialanalysis.ai/articles/mobile-phone-intelligence-inference?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">artificialanalysis.ai</a></p><p class="paragraph" style="text-align:left;"><b>Local AI image generation in Microsoft Paint is not fully local. </b>Paint and Photos still send prompts to a remote server for moderation, then hide a server-issued GUID inside the generated image. If your local pipeline promises nothing leaves the device, check this.<br><a class="link" href="https://xusheng.dev/posts/reversing/mspaint_invisible_watermark/main/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">xusheng.dev</a></p><p class="paragraph" style="text-align:left;"><b>Tool sandboxing is not enough if the model server is exposed. </b>Boyd Kane walks through how an LLM could reach its host machine through the inference engine itself, not shell access or normal tool calls. Worth reading if you run local or self-hosted inference.<br><a class="link" href="https://boydkane.com/essays/llms-could-control-their-host-machines-by-exploiting-inference-engines?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">boydkane.com</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"> <span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>freellmapi puts the free tiers of 28 LLM providers behind one OpenAI-compatible </b><code>/v1</code><b> endpoint</b>. Point an existing client at it and you can test more models without making budget the blocker. Keep it for experiments, not user-facing production traffic.<br><a class="link" href="https://github.com/tashfeenahmed/freellmapi?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/tashfeenahmed/freellmapi</a></p><p class="paragraph" style="text-align:left;"><b>claude-obsidian turns an Obsidian vault into a self-organizing Claude knowledge graph</b>. You drop in sources, Claude reads and links them, and the output stays as plain Markdown files you own. Nice if you like Karpathy’s LLM Wiki idea but want it inside a vault you already use.<br><a class="link" href="https://github.com/AgriciDaniel/claude-obsidian?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/AgriciDaniel/claude-obsidian</a></p><p class="paragraph" style="text-align:left;"><b>Codebuff’s freebuff splits terminal coding work across four specialized agents, and each role can use a different model.</b> That is the practical version of multi-agent coding: spend the expensive model where it matters, use cheaper ones where it does not. If one generalist agent keeps thrashing in your terminal, this is worth trying.<br><a class="link" href="https://github.com/CodebuffAI/freebuff?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/CodebuffAI/freebuff</a></p><p class="paragraph" style="text-align:left;"><b>effective-html is a set of agent skills for making better HTML artifacts: wireframes, prototypes, plans, and diagrams.</b> Most agents can produce something that renders, but “renders” is not the same as “looks good and explains the idea.” Skills are a clean way to fix that once and reuse it across sessions.<br><a class="link" href="https://github.com/plannotator/effective-html?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/plannotator/effective-html</a></p><p class="paragraph" style="text-align:left;"><b>OCR It is a Chrome extension that pins a screen region and OCRs paginated documents locally</b> with bundled Tesseract. Basically, it is for the annoying PDF or document viewer that will not let you copy text. Pin the area, hit the hotkey, and send the extracted text to your LLM.<br><a class="link" href="https://github.com/thiagotigaz/ocr-it?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/thiagotigaz/ocr-it</a></p><p class="paragraph" style="text-align:left;"><b>This GPU checker tells you whether a specific GPU can run a specific LLM, then estimates fit, speed, and stack placement</b>. That turns the most common local-LLM argument into a lookup instead of a forum rabbit hole. The interface is Korean-first, but the calculator is still worth bookmarking.<br><a class="link" href="https://jaeseok614.github.io/llm-gpu-checker-ko/?ui=simple&lang=ko&mode=generative&ctx=8192&con=1&out=512&kv=fp16&runtime=llamacpp&quant=auto&embTokens=384&embBatch=32&embPrecision=auto&embRuntime=tei&embBatchTokens=16384&rerankQuery=64&rerankDoc=512&rerankCandidates=40&rerankBatch=16&rerankPrecision=auto&rerankRuntime=tei&ocrPreset=a4-200&ocrWidth=1654&ocrHeight=2339&ocrBatch=1&ocrPrecision=auto&ocrFeature=text&mediaSteps=28&mediaFrames=81&mediaFps=16&mediaLora=0&mediaOffload=none&mediaOptimization=standard&advisorModel=tinyllama-1-1b-chat&advisorCategory=all&budget=2800000&currentPrice=0&electricity=150&hours=120&advisorVendor=all&advisorForm=all&task=all&provider=all&license=all&licenseUse=all&grade=all&fit=all&sort=latest&view=list&purpose=general&priority=balanced&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">jaeseok614.github.io</a></p><p class="paragraph" style="text-align:left;"><b>x64dbg-mcp-server exposes x64dbg controls over HTTP to any MCP client. </b>Breakpoints, stepping, memory reads, and register dumps become things an assistant can drive directly instead of asking you to paste debugger output. Small repo, big idea if you do reverse engineering or low-level debugging.<br><a class="link" href="https://github.com/duty1g/x64dbg-mcp-server?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">github.com/duty1g/x64dbg-mcp-server</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (134k+ </b>🌟<b> ) is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">📊<span style="color:#ffffff;"><b> By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Peter Walker from OpenRouter reports that GPT-5.6 Sol now makes up 50%+ of US business spend</b> inside OpenAI’s model family on OpenRouter. OpenAI’s frontier model is already taking most of the spend, while Luna and Terra just got cheaper for the work that does not need Sol. The routing question is now pretty direct: what actually needs Sol, and what can move to the cheaper models?<br><a class="link" href="https://x.com/PeterJ_Walker/status/2092017372442058793?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">x.com/PeterJ_Walker</a></p><p class="paragraph" style="text-align:left;"><b>Amazon raised hardware prices by 60%, blaming the memory shortage, in the same week Nvidia told customers to expect AI server increases above 15%.</b> If you have been comparing rented inference against owning the box, that comparison moved twice in seven days.<br><a class="link" href="https://techcrunch.com/2026/08/24/amazon-hikes-hardware-prices-by-60-percent-blaming-memory-shortage/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">techcrunch.com</a><br></p><p class="paragraph" style="text-align:left;"><b>Hugging Face is fielding acquisition interest that would value it at roughly $13 billion. </b>The default hosting layer for open weights being priced for sale is a supply-chain question for anyone whose pipeline starts with a from_pretrained call. Exploratory interest, not an agreed deal, and no acquirer is named.<br><a class="link" href="https://businessinsider.com?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow">businessinsider.com</a></p></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=microsoft-s-new-skill-tunes-your-agents-in-claude-code-codex-and-cursor"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;">The best prompt engineers aren&#39;t typing. They&#39;re talking.</h3><div class="image"><a class="image__link" href="https://ref.wisprflow.ai/beehiiv-ai/?utm_campaign={{publication_alphanumeric_id}}&utm_source=beehiiv&utm_term=ai_p4_q3&_bhiiv=opp_198fba5a-3b75-4f65-a7e2-393a06325b81_4de8c0ec&bhcl_id=6a79eb51-1baf-4413-9452-f418c70c0216_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9edd18e6-e2c3-47f7-9315-469682fd5892/flow-top-teams-move-faster.png?t=1776897861"/></a></div><p class="paragraph" style="text-align:left;">Power users figured this out early: speaking a prompt gives you 10x more context in half the time. You include the edge cases, the examples, the tone you want — because talking is fast enough that you don&#39;t skip them.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://ref.wisprflow.ai/beehiiv-ai/?utm_campaign={{publication_alphanumeric_id}}&utm_source=beehiiv&utm_term=ai_p4_q3&_bhiiv=opp_198fba5a-3b75-4f65-a7e2-393a06325b81_4de8c0ec&bhcl_id=6a79eb51-1baf-4413-9452-f418c70c0216_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Wispr Flow</a> captures everything you say and turns it into clean, structured text for any AI tool. Speak messy. Get polished input. Paste into ChatGPT, Claude, Cursor, or wherever you work.</p><p class="paragraph" style="text-align:left;">89% of messages sent with zero edits. 4x faster than typing. Works system-wide on Mac, Windows, and iPhone.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://ref.wisprflow.ai/beehiiv-ai/?utm_campaign={{publication_alphanumeric_id}}&utm_source=beehiiv&utm_term=ai_p4_q3&_bhiiv=opp_198fba5a-3b75-4f65-a7e2-393a06325b81_4de8c0ec&bhcl_id=6a79eb51-1baf-4413-9452-f418c70c0216_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Start flowing free</a></p><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=038f036b-187b-43f6-82a5-b795c1a083a1&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>GLM-5.3 Beats Fable 5 for Less Money</title>
  <description>+ Free Qwen endpoint, agent sandboxes, Codex adoption outside tech</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d14afb3-d0f8-461c-b65f-7ffe063b3bce/upload_06bd123cda2f78eb.jpg" length="83524" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/glm-5-3-beats-fable-5-for-less-money</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/glm-5-3-beats-fable-5-for-less-money</guid>
  <pubDate>Mon, 24 Aug 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-08-24T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b>Start here ↓</b></h3><p class="paragraph" style="text-align:left;">Together Compute published DeepSWE numbers for GLM-5.3 and Fable 5, and the headline is hard to ignore: <b>87.6% solved for about $16</b> versus <b>69.7% for $21.63</b>.</p><p class="paragraph" style="text-align:left;">There is one real caveat: the GLM number comes from four attempts, not a single shot. But the timing makes it worth paying attention: the <a class="link" href="https://x.com/FT/status/2091443409978319307?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">Financial Times</a> reported the same day that Anthropic&#39;s most powerful model is losing ground to cheaper alternatives.</p><p class="paragraph" style="text-align:left;">The benchmark story and the business story are pointing in the same direction: if you route coding tasks, the default model deserves a retest.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/togethercompute/status/2091711899704385740?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">x.com/togethercompute</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🚀<span style="color:#ffffff;"><b> Shipped</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Empero now offers a free OpenAI-compatible community endpoint for Qwen3.8-27B-FP8. </b><span style="color:#444444;">Zero setup to point a client you already wrote at a 27B open model and see how it handles your actual prompts tonight. No uptime, rate limits, or data-handling terms were stated, so this belongs in a test harness for now.</span><br><a class="link" href="https://x.com/EmperoAI/status/2091515881532637268?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">x.com/EmperoAI</a></p><p class="paragraph" style="text-align:left;"><b>Archal launched API sandboxes built for AI agents, with Slack, Linear, Datadog, and 20+ other stateful environments for CI and evals. </b>This is the part most teams fake with mocks, even though realistic third-party state is where agents usually break first.<br><a class="link" href="https://x.com/AidanTiruvan/status/2091371352544674215?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">x.com/AidanTiruvan</a></p><p class="paragraph" style="text-align:left;"><b>terminal-code runs VS Code inside your terminal by combining code-server with terminal-browser.</b> If your agents already live in a shell, this keeps the editor, diffs, SSH sessions, review flow, and settings import in the same pane.<br><a class="link" href="https://terminal-code.com?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">terminal-code.com</a></p><p class="paragraph" style="text-align:left;"><b>Bezalel puts memory, email, texting, payments, a computer, sandboxes, and connectors behind one MCP URL.</b> That lets you test a capable agent in Claude Code, Codex, or Hermes without wiring six integrations yourself. It is still alpha, and the connector claims vary, so treat the stack as promising but unproven. <a class="link" href="https://bezalel.sh?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">bezalel.sh</a></p><p class="paragraph" style="text-align:left;"><b>SenseNova released U1.5-8B-MoT, an 8B any-to-any model on Hugging Face. </b>That is small enough to test multimodal routing locally instead of renting hardware just to see if it fits your stack. The model card does not include benchmarks, so keep it in the “try it yourself” bucket for now.<br><a class="link" href="https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">huggingface.co/sensenova</a></p><p class="paragraph" style="text-align:left;"><b>Audio8 released TTS-Preview-0.1B, which it calls the world’s smallest zero-shot TTS model.</b> At 0.1B parameters, it is small enough to test local voice synthesis per request instead of sending every character to an API. <br><a class="link" href="https://huggingface.co/Audio8?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">huggingface.co/Audio8</a></p><p class="paragraph" style="text-align:left;"></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🧠<span style="color:#ffffff;"><b> Worth Knowing</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>Anthropic put Mythos 5, its non-public model tier, behind Claude Security for scanning GitHub repos and proposing patches. </b>Anthropic’s strongest code/security model is not going to chat first. It is going to the place where missed bugs cost money.<b> </b><a class="link" href="https://thenewstack.io?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow"><b>thenewstack.io</b></a></p><p class="paragraph" style="text-align:left;"><b>Fabien Sanglard published the agent.md he uses to hold LLM-assisted code to his own quality standard.</b> <span style="color:#444444;">It is a copyable file from someone with a documented standard behind it, which makes it worth diffing against your own instruction file line by line rather than adopting wholesale. </span><br><a class="link" href="https://fabiensanglard.net/agent.md/index.html?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">fabiensanglard.net</a></p><p class="paragraph" style="text-align:left;"><b>The MCP maintainers put progressive tool discovery and a standard tool-result contract on the roadmap.</b> Progressive discovery fixes the “100 tools in context before the user asks anything” problem, while one result contract gives client authors a single shape to build against. This is a prioritization doc, not a shipped spec, so treat it as direction, not final API. <br><a class="link" href="https://blog.modelcontextprotocol.io?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">blog.modelcontextprotocol.io</a></p><p class="paragraph" style="text-align:left;"><b>Google Cloud published five patterns for long-horizon agents:</b> stable prefixes, background learning, persistent workspaces, explicit failures, and guard chains. These are the bugs that do not crash; they quietly burn cache, lose memory, wipe tools, or mark timed-out sub-agents as done. Steal the checklist before the framework.<br><a class="link" href="https://x.com/GoogleCloudTech/status/2090248297214525569?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">https://x.com/GoogleCloudTech/status/2090248297214525569</a></p><p class="paragraph" style="text-align:left;"><b>Someone fine-tuned Gemma 4 12B for tool calling and reported a 2.7x improvement, picking that model because it fits comfortably in 16GB of VRAM.</b> <span style="color:#444444;">Tool-calling reliability is the failure most teams paper over with retries and validators, and this says it is trainable on a card you already own. </span><br><a class="link" href="https://old.reddit.com/r/LocalLLaMA/comments/1vvtu9z/i_fine_tuned_gemma_4_12b_for_a_27x_improvement_on/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">reddit.com/r/LocalLLaMA</a></p><p class="paragraph" style="text-align:left;"><b>OpenAI’s Codex lead Tibo pointed to two drivers behind rate-limit pain: image-heavy long sessions with repeated compactions, and Computer Use getting expensive at p95+. </b>You cannot fix Codex’s limits, but you can stop burning them on avoidable session shape. Keep images out of runs you expect to compact. <a class="link" href="https://x.com/thsottiaux?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">x.com/thsottiaux</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;">🔧<span style="color:#FFFFFF;"><b> Clone and Run</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"> <span style="background-color:#e8e1ff;"><b>Clone & Run of the Day </b></span><br><b>Apache Maka is a local-first agent workspace that logs model messages, tool calls, permissions, and termination events as first-class records.</b> The useful part is logging the two things you usually lose after a bad run: who approved what, and how the agent stopped. Still in Apache incubation, so treat it as early infrastructure. <a class="link" href="https://github.com/apache/maka?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">github.com/apache/maka</a></p><p class="paragraph" style="text-align:left;"><b>Paseo is an open-source control surface for Claude Code, Codex, Copilot, and OpenCode across local machines and a VPS.</b> If your agent setup is currently tmux, SSH, and too many panes, this is the cleaner version of the layer you are already rebuilding. <br><a class="link" href="https://github.com/getpaseo/paseo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/getpaseo/paseo</a></p><p class="paragraph" style="text-align:left;"><b>Anthropic’s claude-plugins-community is a read-only mirror of the Claude Code and Claude Cowork plugin directory. </b>It shows what is actually in the marketplace, and gives plugin builders the official path to get listed.<br><a class="link" href="https://github.com/anthropics/claude-plugins-community?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/anthropics/claude-plugins-community</a></p><p class="paragraph" style="text-align:left;"><b>OpenHuman is an open-source personal AI with local-first memory and agent-fleet orchestration. </b>Most projects give you one or the other; this tries to put the memory store and the agent runner in the same local system. <a class="link" href="https://github.com/tinyhumansai/openhuman?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/tinyhumansai/openhuman</a></p><p class="paragraph" style="text-align:left;"><b>agent-safe-pipeline is a reference architecture where agents can propose actions but cannot approve them. </b>It keeps intent capture immutable and sends the final decision to an independent policy layer, which is where agent permissions belong.<br><a class="link" href="https://github.com/decionis/agent-safe-pipeline?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/decionis/agent-safe-pipeline</a></p><p class="paragraph" style="text-align:left;"><b>open-slide turns a prompt into an interactive deck by having coding agents write the React components.</b> <span style="color:#444444;">Generating the artifact as code instead of as a rendered file means the output stays editable by the same agent that made it.</span><br><a class="link" href="https://github.com/1weiho/open-slide?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/1weiho/open-slide</a></p><p class="paragraph" style="text-align:left;"><b>Ruflo is an open-source meta-harness for deploying multi-agent swarms and coordinating autonomous workflows</b>. If you want to see what the coordination layer looks like beyond one agent at a time, this is a large implementation to read.<br><a class="link" href="https://github.com/ruvnet/ruflo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/ruvnet/ruflo</a></p><p class="paragraph" style="text-align:left;"><b>oh-my-subagents adds persistence and tracking to subagent workflows that are ephemeral by default.</b> <span style="color:#444444;">The reason you cannot answer what your subagent did last Tuesday is that nothing kept the run, and this is a small patch over exactly that gap.</span><br><a class="link" href="https://github.com/ringlochid/oh-my-subagents?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">github.com/ringlochid/oh-my-subagents</a></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (134k+ </b>🌟<b> ) is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.</b> It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. <a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>By the Number</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><span style="background-color:#e8e1ff;"><b>Number of the Day</b></span><br><b>Together Compute’s routing math puts GLM-5.3 at 17 solves per $100 on DeepSWE, versus 3 solves per $100 for Fable 5.</b> That is the number that matters if your agent can retry: not which model wins once, but which model gives you more passing runs per dollar.<br><a class="link" href="https://x.com/togethercompute/status/2091361283459223731?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">x.com/togethercompute</a></p><p class="paragraph" style="text-align:left;"><b>Researchers ran GLM-5.2 753B at 14.9 tok/s on a 96GB GPU, DeepSeek-V4-Flash 284B at 22 tok/s on 32GB, and Qwen3.6-35B at 39.3 tok/s on an 8GB card. </b>If the setup holds, “too big to serve locally” just moved down a tier. The 8GB Qwen number is the one most readers can actually try.<br><a class="link" href="https://arxiv.org/html/2608.16157v1?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #5a3fe0">arxiv.org</a></p><p class="paragraph" style="text-align:left;"><b>a16z reports the fastest-growing Codex adopters since February are outside tech: legal up 108x, sales and recruiting 41x, marketing 26x, healthcare 24x.</b> The first non-engineering power users are not “vibe coding.” They are turning repeatable knowledge work into Codex-shaped tasks. <br><a class="link" href="https://www.a16z.news/p/charts-of-the-week-winds-of-thematic?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow">a16z.news</a> </p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.</p><p class="paragraph" style="text-align:left;">If you found one thing to try, share the issue with someone who ships.</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=glm-5-3-beats-fable-5-for-less-money"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=4d9601e1-57a3-4c6a-8407-be9e9b819c9f&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Opus 4.8-level model now runs locally for FREE</title>
  <description>+ OpenRouter Fusion, GLM-5.2 locally, Loop Engineering</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f8018326-b159-4423-acd2-8757e27bdefe/upload_a7a0eae2c0c398c5.jpg" length="93822" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/opus-4-8-level-model-now-runs-locally-for-free</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/opus-4-8-level-model-now-runs-locally-for-free</guid>
  <pubDate>Tue, 23 Jun 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-06-23T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Google Cloud turns scattered knowledge into agent-readable files</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Vibe is here: one agent for work and code</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Run GLM 5.2 locally</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Telegram bots can now talk to other bots</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Open-source alternative to Loom, Granola, and Wisprflow</b></p></li></ol><p class="paragraph" style="text-align:start;">& a lot more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/generative-ui-is-the-new-frontend?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Generative UI Is the New Frontend</a></b></p><p class="paragraph" style="text-align:left;">The frontend used to be a fixed thing. Designers drew it. Engineers built it. Users got what shipped.</p><p class="paragraph" style="text-align:left;">That&#39;s over.</p><p class="paragraph" style="text-align:left;">The interfaces shipping in 2026 are drawn partly by the agent itself, in real time, from what the user actually asked for. Ask for a table, get a table. Not a paragraph describing one.</p><p class="paragraph" style="text-align:left;">Generative UI is the layer that lets agents stop describing and start showing. </p><p class="paragraph" style="text-align:left;">This guide walks you through three patterns that have emerged on how to build it, and the differences between them matter more than most teams realize.</p><div class="embed"><a class="embed__url" href="https://www.theunwindai.com/p/generative-ui-is-the-new-frontend?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank"><img class="embed__image embed__image--left" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/6c3faaf3-c6cf-4f53-806b-8a114703b66a/Generative_UI_Is_the_New_Frontend.png?t=1780532864"/><div class="embed__content"><p class="embed__title"> Generative UI Is the New Frontend </p><p class="embed__description"> How AI agents stop describing and start showing </p></div></a></div><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Google Cloud Turns Company Knowledge Into Agent-Readable Files</b></a><b> </b>🧠📁</h3><div class="image"><a class="image__link" href="https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/695ed768-c228-4a03-a336-9a11c34a28f5/image.png?t=1782192445"/></a></div><p class="paragraph" style="text-align:left;">Every team building internal agents eventually hits the same wall: the model is smart, but the context is scattered everywhere.</p><p class="paragraph" style="text-align:left;">Part of it lives in data catalogs. Part of it lives in wikis. Part of it lives in code comments, dashboards, tribal knowledge, and that one senior engineer&#39;s brain.</p><p class="paragraph" style="text-align:left;"><b>Google Cloud just introduced Open Knowledge Format (OKF)</b> to make that context portable. It is a vendor-neutral spec that turns enterprise knowledge into Markdown files with YAML frontmatter, so agents can read it, search it, version it, and move it between tools without another custom integration.</p><p class="paragraph" style="text-align:left;">The nice part is how boring the format is. Just Markdown. Just files. Just a small set of structured fields like type, title, description, resource, tags, and timestamp.</p><p class="paragraph" style="text-align:left;">That is exactly why it could work. Agents do not need another complex metadata platform; they need context they can actually open, inspect, and use.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Markdown-first</b>: OKF represents context as human-readable Markdown, so engineers can review it in normal editors and agents can index it without special tooling.</p></li><li><p class="paragraph" style="text-align:left;"><b>Structured enough for agents</b>: YAML frontmatter adds queryable fields like type, title, resource, tags, and timestamp without turning the whole thing into a heavy schema project.</p></li><li><p class="paragraph" style="text-align:left;"><b>Portable by default</b>: OKF bundles can live in Git, ship as files, mount on a filesystem, or move across tools without locking context inside one vendor&#39;s catalog.</p></li><li><p class="paragraph" style="text-align:left;"><b>Reference implementations included</b>: Google shipped examples including BigQuery enrichment, a static HTML visualizer, sample bundles, and Knowledge Catalog ingestion support.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://mistr.al/vibe-unwindai-nl?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Vibe is here: one agent for work and code</b></a></h3><div class="image"><a class="image__link" href="https://mistr.al/vibe-unwindai-nl?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7f36dbd-9c8e-45e7-8b4a-b6fba210569c/image.jpeg?t=1782192515"/></a></div><p class="paragraph" style="text-align:left;">Meet <a class="link" href="https://mistr.al/vibe-unwindai-nl?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Vibe by Mistral</b></a>, one agent and one licence across work and code. Vibe takes on long-running, multi-step work: catching up across your inbox and calendar, running deep research, drafting deliverables, and taking coding work from request to merged change, across the web app, your editor, and your terminal.</p><p class="paragraph" style="text-align:left;"><b>Key highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Work Mode for complex, multi-stage tasks</b>: Maps out a plan, gets your sign-off, then works across your connectors to carry it through. Every tool call and reasoning step is visible and expandable as it runs. </p></li><li><p class="paragraph" style="text-align:left;"><b>Code Mode for remote coding sessions</b>: Connect to GitHub, start sessions, and see them through to a pull request. Sessions run in an isolated sandbox, persist while your machine is off, and can run in parallel.</p></li><li><p class="paragraph" style="text-align:left;"><b>VS Code extension</b>: Vibe now works across your whole project inside VS Code. Reads, edits, and runs commands in a side panel. Open files attach automatically, @ mentions pull in context from anywhere in your repo.</p></li><li><p class="paragraph" style="text-align:left;"><b>CLI updates</b>: Skills become / commands. Permissions are session-scoped. /teleport moves a live session between your terminal and the cloud, history and approvals intact.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://openrouter.ai/blog/announcements/fusion-beats-frontier/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>OpenRouter Fusion Makes Model Panels a One-Call Primitive </b></a>🧪🤝</h3><div class="image"><a class="image__link" href="https://openrouter.ai/blog/announcements/fusion-beats-frontier/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3580c67f-125f-485f-9850-4859e9a0e5bb/image.png?t=1782192336"/></a></div><p class="paragraph" style="text-align:left;">The most annoying part of using multiple models is that you usually have to become the router yourself.</p><p class="paragraph" style="text-align:left;">You ask one model, compare it with another, maybe try a third, then manually decide which answer is right. </p><p class="paragraph" style="text-align:left;"><b>OpenRouter&#39;s new Fusion API</b> turns that pattern into a single call. You send a prompt to Fusion, it dispatches the task to a panel of models in parallel, gives them web search and web fetch, then uses a judge model to compare the answers before producing the final response.</p><p class="paragraph" style="text-align:left;">The results are worth paying attention to: Fable 5 + GPT-5.5 fused together scored 69.0% on DRACO, beating every individual model in OpenRouter&#39;s test, including Fable 5 alone at 65.3% and GPT-5.5 alone at 60.0%. That matters because Fable 5 is Anthropic&#39;s strongest model, and Fusion still found extra lift by pairing it with another frontier model instead of treating one model as the ceiling.</p><p class="paragraph" style="text-align:left;">OpenRouter also tested a budget panel with Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro. That panel scored 64.7%, beating GPT-5.5 and Claude Opus 4.8 individually, coming within about one point of Fable 5 alone, and doing it at roughly half the cost.</p><p class="paragraph" style="text-align:left;">The real story is not just the benchmark. It is that model diversity is becoming a product primitive. Instead of picking one model and hoping it is the right one, builders can start treating models like a small research team.</p><p class="paragraph" style="text-align:left;">You can use it through the normal OpenRouter API. Just call openrouter/fusion directly or configure the Fusion plugin with your own analysis models and judge model.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/addyosmani/status/2064127981161959567?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Everyone is talking about loop engineering, but Addy&#39;s version makes it usable</b></a><br>Addy Osmani&#39;s piece is useful because it turns the phrase into an actual operating model. The shift is from &quot;I prompt the agent&quot; to &quot;I design the loop that finds work, hands it to agents, checks the result, records state, and decides the next step.&quot;</p><p class="paragraph" style="text-align:left;">That is a better frame for where coding agents are going. The prompt is no longer the main artifact. The loop is. If you are building agent workflows, this gives you a cleaner way to think about retries, memory, evaluation, escalation, and token cost before you wire everything together.</p><p class="paragraph" style="text-align:left;"> </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://sakana.ai/fugu/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Sakana Fugu explores orchestration as the model</b></a><br>Sakana&#39;s Fugu is interesting because it is less about launching another standalone model and more about coordinating multiple models into a stronger system. The bet is that intelligence can come from routing, combining, challenging, and arbitrating models, not only from scaling one model in isolation.</p><p class="paragraph" style="text-align:left;">That makes it rhyme with the Fusion story, but from a research direction rather than an API product. The useful takeaway for builders is simple: the next frontier may be systems that know which model to use, when to ask for disagreement, and how to merge partial answers without making the user manage the whole process.</p><p class="paragraph" style="text-align:left;"> </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://telegram.org/blog/ai-bot-revolution-11-new-features?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Telegram bots can now talk to other bots</b></a><br>Telegram&#39;s latest bot update lets bots respond to other bots, not just humans. That sounds small, but it changes what Telegram can be used for: not just a chat UI, but a lightweight coordination layer for agent workflows.</p><p class="paragraph" style="text-align:left;">You could mention one bot, that bot could hand work to another bot, and the whole exchange stays visible in a normal chat thread.</p><p class="paragraph" style="text-align:left;"> </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://unsloth.ai/docs/models/glm-5.2?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Run GLM-5.2 locally with Unsloth</b></a><br>Unsloth just published a guide for running GLM-5.2 with Dynamic GGUFs, including 1-bit and 2-bit quant options, llama.cpp instructions, and Unsloth Studio support. The 2-bit build is still huge at around 239GB, but that is dramatically smaller than the full 1.51TB model.</p><p class="paragraph" style="text-align:left;">This is not casual laptop territory, but it is meaningful for local-agent builders with serious memory available. A 744B-parameter open model with 40B active parameters and a 1M context window is already being squeezed into setups that advanced users can actually experiment with.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://birdclaw.sh/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Birdclaw</a></b>: A local-first Twitter workspace that imports your X archive, syncs timeline/bookmarks/mentions, and stores everything in SQLite. The useful part is that your X memory becomes searchable and agent-readable: you can full-text search old likes and bookmarks, triage mentions with AI ranking, generate local digests, and keep a Git-friendly backup instead of losing everything inside the platform UI.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.agent-native.com/templates/clips?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Agent-Native Clips</a></b>: An open-source Loom + Granola + Wisprflow-style app for screen recordings, meeting notes, and dictation. Every clip gets transcripts, summaries, timestamped frames, and searchable history, so an agent can understand what happened in a video or meeting without needing raw audio/video ingestion.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://docs.stripe.com/directory?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow"><b>Stripe Directory</b></a>: A public-preview Stripe CLI directory for discovering businesses and services on the Stripe network. Developers and agents can search providers by keyword, then get structured results for Stripe Apps, <a class="link" href="https://Projects.dev?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Projects.dev</a> providers, machine-payment endpoints, and business profiles instead of manually hunting across docs and marketplaces.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (113k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=opus-4-8-level-model-now-runs-locally-for-free"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=a7bd1c91-2f6a-4dcb-bced-c715cceda390&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Claude Code now spins up 100s of parallel agents on one task</title>
  <description>+ Apple drops a native AI framework for on-deivce AI agents</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e08b48e8-95bf-4273-beb7-604cb76ee2fe/upload_cf0689ac0fa487ec.jpg" length="100811" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/claude-code-now-spins-up-100s-of-parallel-agents-on-one-task</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/claude-code-now-spins-up-100s-of-parallel-agents-on-one-task</guid>
  <pubDate>Tue, 09 Jun 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-06-09T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Apple Core AI Framework for on-device agents</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Printing Press: Print agent-native CLIs from a single prompt</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Claude Code Dynamic Workflows with massive parallelism</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Your job is to write agent loops now</b></p></li><li><p class="paragraph" style="text-align:left;"><b>ChatGPT now &quot;dreams&quot; to build better memory</b></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/generative-ui-is-the-new-frontend?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">Generative UI Is the New Frontend</a></b></p><p class="paragraph" style="text-align:left;">The frontend used to be a fixed thing. Designers drew it. Engineers built it. Users got what shipped.</p><p class="paragraph" style="text-align:left;">That&#39;s over.</p><p class="paragraph" style="text-align:left;">The interfaces shipping in 2026 are drawn partly by the agent itself, in real time, from what the user actually asked for. Ask for a table, get a table. Not a paragraph describing one.</p><p class="paragraph" style="text-align:left;">Generative UI is the layer that lets agents stop describing and start showing. </p><p class="paragraph" style="text-align:left;">This guide walks you through three patterns that have emerged on how to build it, and the differences between them matter more than most teams realize.</p><div class="embed"><a class="embed__url" href="https://www.theunwindai.com/p/generative-ui-is-the-new-frontend?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank"><img class="embed__image embed__image--left" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/6c3faaf3-c6cf-4f53-806b-8a114703b66a/Generative_UI_Is_the_New_Frontend.png?t=1780532864"/><div class="embed__content"><p class="embed__title"> Generative UI Is the New Frontend </p><p class="embed__description"> How AI agents stop describing and start showing </p></div></a></div><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://developer.apple.com/documentation/coreai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Apple Just Gave Developers Their Own On-Device AI Framework</b></a></h3><div class="image"><a class="image__link" href="https://developer.apple.com/documentation/coreai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a3bdfe49-348d-44ee-92d5-4d60fc268372/Screenshot_2026-06-08_at_11.04.15_PM.png?t=1780985059"/></a></div><p class="paragraph" style="text-align:left;">Forget calling external APIs. Apple&#39;s new Core AI framework, announced at WWDC 2026, gives developers Swift-native access to Apple&#39;s on-device foundation models with tool calling, structured generation, and full Apple Intelligence integration.</p><p class="paragraph" style="text-align:left;">This is the first time Apple has opened up its on-device models as a developer-facing framework. You write Swift, define tools, and the model runs locally on the device with zero cloud round-trips. Privacy-first by default, no API keys, no usage-based pricing, no latency from network calls. For anyone building iOS or macOS apps, this changes how you think about adding intelligence to your product.</p><p class="paragraph" style="text-align:left;">The bigger picture: Apple also revealed that Apple Intelligence is now co-developed with Google using Gemini models under the hood, running both on-device and through Private Cloud Compute. A system orchestrator automatically coordinates AI features across apps.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Tool calling in Swift</b>: Define custom tools that the on-device model can invoke, enabling agentic workflows entirely on the user&#39;s device without a server.</p></li><li><p class="paragraph" style="text-align:left;"><b>Structured generation</b>: Get typed, schema-conforming outputs from the model, not just raw text. Build reliable features without post-processing hacks.</p></li><li><p class="paragraph" style="text-align:left;"><b>Gemini under the hood</b>: Apple Intelligence now runs on foundation models co-developed with Google, giving the platform multimodal capabilities, including image generation, visual Q&A, and speech generation.</p></li><li><p class="paragraph" style="text-align:left;"><b>No cloud dependency</b>: Models run locally. Your users&#39; data stays on their devices. No API costs, rate limits, or cold starts.</p></li><li><p class="paragraph" style="text-align:left;"><b>Available now</b>: Core AI ships with iOS 27, macOS 27, and the latest Xcode. Documentation is live at <a class="link" href="https://developer.apple.com/documentation/coreai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">developer.apple.com/documentation/coreai</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://printingpress.dev/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Print an Agent-Native CLI for Any API from a Single Prompt</b></a></h3><div class="image"><a class="image__link" href="https://printingpress.dev/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0852d831-0bca-4caf-88ce-0db19cb32ce2/printing_press.png?t=1780984943"/></a></div><p class="paragraph" style="text-align:left;">What if every app, API, and website your agent needs came as a purpose-built CLI with a local SQLite mirror, compound commands, and token-efficient output?</p><p class="paragraph" style="text-align:left;">That&#39;s Printing Press by Matt Van Horn and Trevin Chow. Point it at an API spec, a website URL, or even a service with no public API, and it generates a Go CLI, a Claude Code skill, an OpenClaw skill, and an MCP server. All from one prompt.</p><p class="paragraph" style="text-align:left;">Super interesting concept: a local SQLite mirror beats a remote API call. Compound commands beat ten round trips. An agent-native CLI beats raw HTTP. When you &quot;print&quot; an ESPN CLI, you don&#39;t get a thin wrapper around REST endpoints. You get live scores, series state, leading scorers, and injury news in one call, all queried from a local database that syncs incrementally. Same goes for Linear, Slack, Notion, or any of the 237+ CLIs in their Public Library.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>No API needed</b>: For services without a public API, Printing Press launches a browser, captures traffic, reverse-engineers the endpoints, and generates the spec automatically. If you can click through it, the press can build a CLI.</p></li><li><p class="paragraph" style="text-align:left;"><b>Local-first data layer</b>: High-gravity resources get domain-specific SQLite tables with FTS5 full-text search and incremental sync. Queries run in milliseconds offline. Your agent never waits for a 429.</p></li><li><p class="paragraph" style="text-align:left;"><b>237+ community CLIs</b>: The Public Library ships pre-built CLIs across 19 categories, from flight search to restaurant reservations to eBay auctions. Install with one command or let your agent browse and pick what it needs.</p></li><li><p class="paragraph" style="text-align:left;"><b>Token-efficient by default</b>: --compact mode cuts 60-80% of tokens. Auto-JSON when piped. Typed exit codes for agent self-correction. The CLI is built for agents first, humans second.</p></li><li><p class="paragraph" style="text-align:left;"><b>Try it now</b>: Install via Go, add the Claude Code or OpenClaw skills, and run /printing-press &lt;app&gt; inside your agent. Check it out at <a class="link" href="https://printingpress.dev?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">printingpress.dev</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://claude.com/blog/introducing-dynamic-workflows-in-claude-code?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Claude Code Can Now Orchestrate 100s of Parallel Agents on a Single Task</b></a></h3><div class="image"><a class="image__link" href="https://claude.com/blog/introducing-dynamic-workflows-in-claude-code?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fac16e46-09a2-43cb-82d9-df69753554b3/image.png?t=1780985221"/></a></div><p class="paragraph" style="text-align:left;">Some problems are too big for one agent in one pass. A bug hunt across an entire service. A migration that touches hundreds of files. A plan you want stress-tested from every angle before committing.</p><p class="paragraph" style="text-align:left;">Anthropic&#39;s Dynamic Workflows for Claude Code changes the math entirely. Claude writes a custom JavaScript orchestration script on the fly, fans work out across 100s of parallel subagents, has independent agents try to break each other&#39;s results, and keeps iterating until answers converge. The coordination lives in code, not context, so the plan stays on track no matter how big the task gets.</p><p class="paragraph" style="text-align:left;">The proof of concept is wild: Jarred Sumner used Dynamic Workflows to port Bun from Zig to Rust. 750,000 lines of Rust, 99.8% test suite passing, eleven days from first commit to merge. Hundreds of agents worked in parallel with two reviewers on each file.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Claude writes the orchestrator</b>: No pre-built templates. Claude generates a bespoke JS script tailored to your specific task, then runs it. Every workflow is custom.</p></li><li><p class="paragraph" style="text-align:left;"><b>Independent verification built in</b>: Agents tackle the problem from different angles, other agents try to refute what they found, and the run iterates until results converge. This is how it catches things a single pass misses.</p></li><li><p class="paragraph" style="text-align:left;"><b>Resumable and saveable</b>: Progress is checkpointed. Interrupted jobs pick up where they left off. Save a workflow as a reusable /command for future sessions with structured input parameters.</p></li><li><p class="paragraph" style="text-align:left;"><b>Token warning</b>: These workflows consume meaningfully more tokens than a typical session. Start with a scoped task to get a feel for usage before throwing it at your whole codebase.</p></li><li><p class="paragraph" style="text-align:left;"><b>Available now</b>: Research preview on Max, Team, and Enterprise plans, plus the Claude API, Amazon Bedrock, Vertex AI, and Microsoft Foundry. Requires Claude Code v2.1.154+. Turn on ultracode effort level or just ask Claude to &quot;create a workflow.&quot;</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/mvanhorn/article/2063865685558903149?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Don’t Prompt Agent, Design Loops that prompt your agent</b></a><b> (WTF!)</b><br>Peter Steinberger posted six words on Saturday that hit 6.3 million views: &quot;You should be designing loops that prompt your agents.&quot; Boris Cherny, the creator of Claude Code, said the same thing a few days earlier: &quot;I don&#39;t prompt Claude anymore. I have loops running. They&#39;re the ones prompting Claude.&quot; If all of this chatter left you wondering what the hell a loop even is, Matt Van Horn wrote the definitive explainer. He traces the concept all the way back to the 2022 ReAct paper, through Geoffrey Huntley&#39;s ralph loop, to today&#39;s multi-agent orchestration loops that run on cron and survive restarts. Worth the read.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.theverge.com/tech/944245/apple-wwdc-2026-ai-siri-gemini?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Apple Intelligence, take two</b></a><br>Remember when Apple announced Apple Intelligence with ChatGPT as the backup brain at WWDC 2024? The &quot;smart Siri&quot; with personal context, on-screen awareness, etc? Most of it never shipped, and apparently, “it wasn&#39;t good enough&quot; and &quot;didn&#39;t converge quality-wise.&quot; Two years later, they&#39;re trying again, this time with Google. The new Apple Intelligence, announced at WWDC yesterday, is co-developed with Google using Gemini as the foundation, not just a fallback. On-device and Private Cloud Compute, multimodal everything, and a conversational Siri that&#39;s getting its own standalone app. </p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://openai.com/index/chatgpt-memory-dreaming/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>ChatGPT now &quot;dreams&quot; to remember you better</b></a><br>OpenAI shipped Dreaming V3, and the name is apt. ChatGPT now runs a background memory synthesis process when you&#39;re not chatting, analyzing your conversation history and building a unified &quot;Memory Summary&quot; that stays current over time. The old &quot;saved memories&quot; approach (manually saying &quot;remember this&quot;) is gone. Now it automatically captures context from natural conversation and updates temporal facts, so it knows your Singapore trip is in the past, not upcoming. Available on Plus and Pro now, rolling out to free users soon after a 5x compute reduction made it feasible at scale. The direction is clear: persistent, stateful AI assistants where memory is infrastructure, not a feature you toggle on.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://research.google/blog/unlocking-dependable-responses-with-gemini-enterprise-agent-platforms-agentic-rag/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Google Research tackles RAG&#39;s biggest problem</b></a><br>Standard RAG retrieves once and hopes for the best. Google Research&#39;s new &quot;Agentic RAG&quot; for their Gemini Enterprise Agent Platform retrieves, checks if it got enough, and goes back for more. The key innovation is a &quot;Sufficient Context Agent&quot; that inspects retrieved snippets, evaluates a draft response, identifies exactly what&#39;s missing, and sends targeted follow-up searches. It&#39;s a multi-agent pipeline: orchestrator, planner, query rewriter, search fanout, and synthesis. The result is a 34% accuracy improvement over standard RAG on factuality benchmarks with negligible latency overhead. The pattern is the real takeaway here, even if you&#39;re not on Google Cloud.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/Panniantong/Agent-Reach?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>Agent-Reach</b></a>: Deploy one AI agent across Telegram, Discord, Slack, WhatsApp, Web, and CLI simultaneously from a single codebase. It normalizes messages into a consistent format per platform and auto-adapts responses. Ships with ready-made adapters for LangChain, CrewAI, and AutoGen.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/luongnv89/claude-howto?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>claude-howto</b></a>: A visual, example-driven guide to Claude Code covering everything from basic setup to advanced workflows like multi-agent orchestration and custom skills. Super useful whether you&#39;re just starting with Claude Code or trying to level up. </p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/mvanhorn/last30days-skill?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>last30days-skill</b></a>: An agent skill that synthesizes research across X, Reddit, HN, YouTube, TikTok, and GitHub from the last 30 days on any topic you give it. Matt Van Horn used it to write that viral loops article, running it against the word everyone was fighting about. </p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/mvanhorn/agentcookie?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow"><b>agentcookie</b></a>: Continuously syncs your browser cookies, bearer tokens, and API keys from your daily-driver Mac to a second Mac where your agents run, encrypted over Tailscale. Your agents wake up authenticated to every service you use, zero per-site login ceremony, no cloud middleman. </p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (113k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=claude-code-now-spins-up-100s-of-parallel-agents-on-one-task"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=4fe39004-4fe4-4eb2-90b9-f63062e2ff13&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Generative UI Is the New Frontend </title>
  <description>How AI agents stop describing and start showing</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d70eb4f2-2c1c-416e-8442-3d2d88c59b83/upload_0245099b25305f08.jpg" length="100344" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/generative-ui-is-the-new-frontend</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/generative-ui-is-the-new-frontend</guid>
  <pubDate>Thu, 04 Jun 2026 00:28:39 +0000</pubDate>
  <atom:published>2026-06-04T00:28:39Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">The frontend used to be a fixed thing. Designers drew it. Engineers built it. Users got what shipped.</p><p class="paragraph" style="text-align:left;">That&#39;s over. </p><p class="paragraph" style="text-align:left;">The interfaces shipping in 2026 are drawn partly by the agent itself, in real time, from what the user actually asked for. Ask for a table, get a table. Not a paragraph describing one.</p><p class="paragraph" style="text-align:left;">Generative UI is the layer that lets agents stop describing and start showing. Three patterns have emerged for how to build it, and the differences between them matter more than most teams realize.</p><p class="paragraph" style="text-align:left;">But there isn&#39;t one way to build this. There are three. And most teams pick one without knowing they chose.</p><h2 class="heading" style="text-align:left;"><b>The protocol stack</b></h2><p class="paragraph" style="text-align:left;">Three protocols. Each does one job.</p><p class="paragraph" style="text-align:left;"><b>MCP</b> connects agents to tools. <b>A2A</b> connects agents to each other. <b>AG-UI</b> connects agents to users.</p><p class="paragraph" style="text-align:left;"><b>AG-UI </b>is the streaming layer that carries everything you&#39;ll see below: tool calls, A2UI schemas, MCP App events, state deltas. Runs over SSE. State flows both ways on the same stream. User edits, agent sees. Agent mutates, user sees.</p><p class="paragraph" style="text-align:left;"><b>A2UI</b> is Google&#39;s spec for agents emitting UI as schema. It rides on AG-UI. CopilotKit ships it in production.</p><p class="paragraph" style="text-align:left;">You don&#39;t write a parser for any of this. CopilotKit is an AG-UI client and decodes the stream for you.</p><h2 class="heading" style="text-align:left;"><b>The three patterns most teams confuse</b></h2><p class="paragraph" style="text-align:left;">Ask ten developers what Generative UI is. You get ten answers. Most of them are describing whichever pattern their current framework ships.</p><p class="paragraph" style="text-align:left;">There are just three. The spectrum runs from more control to more flexibility.</p><ul><li><p class="paragraph" style="text-align:left;"><b>Controlled:</b> You pre-build the components. The agent picks which to render.</p></li><li><p class="paragraph" style="text-align:left;"><b>Declarative:</b> The agent emits a schema. Your app maps it to components.</p></li><li><p class="paragraph" style="text-align:left;"><b>Open-ended:</b> The agent writes raw HTML. Your app renders it in a sandbox.</p></li></ul><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e36bdb41-766b-4733-b369-a32f63087161/image.png?t=1780531905"/></div><p class="paragraph" style="text-align:left;">Every Gen UI framework in 2026 lives somewhere on this line. The differences are architectural, not cosmetic. Each pattern breaks your app in a different way at scale.</p><p class="paragraph" style="text-align:left;">I tried different stacks. Most cover one pattern well. Landed on CopilotKit because it supports all three on the same runtime, riding AG-UI. That&#39;s the stack everything below runs on.</p><h2 class="heading" style="text-align:left;"><b>Pattern 1: Controlled, frontend owns the UI</b></h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03ae9963-4371-45ed-94ae-1fb876e8e803/image.png?t=1780531917"/></div><p class="paragraph" style="text-align:left;">This is where most teams start. It&#39;s also where most teams get stuck.</p><p class="paragraph" style="text-align:left;">You pre-build a React component. You bind it to a tool name. The agent picks that tool and the component renders inline in chat with the agent&#39;s args as props.</p><p class="paragraph" style="text-align:left;">One frontend hook. Zero agent code. That&#39;s it.</p><div class="codeblock"><pre><code>&quot;use client&quot;;
import &#123; z &#125; from &quot;zod&quot;;
import &#123; useComponent &#125; from &quot;@copilotkit/react-core/v2&quot;;

const expenseChartSchema = z.object(&#123;
  title: z.string(),
  data: z.array(z.object(&#123; label: z.string(), value: z.number() &#125;)),
&#125;);

function ExpenseChart(&#123; title, data &#125;: z.infer&lt;typeof expenseChartSchema&gt;) &#123;
  return (
    &lt;section className=&quot;rounded-xl border p-4&quot;&gt;
      &lt;h3 className=&quot;text-sm font-medium&quot;&gt;&#123;title&#125;&lt;/h3&gt;
      &lt;ul className=&quot;mt-2 grid gap-1&quot;&gt;
        &#123;data.map((d) =&gt; (
          &lt;li key=&#123;d.label&#125; className=&quot;flex justify-between text-sm&quot;&gt;
            &lt;span&gt;&#123;d.label&#125;&lt;/span&gt;
            &lt;span&gt;$&#123;d.value&#125;&lt;/span&gt;
          &lt;/li&gt;
        ))&#125;
      &lt;/ul&gt;
    &lt;/section&gt;
  );
&#125;

export function ExpensesCopilot() &#123;
  useComponent(&#123;
    name: &quot;showExpenseChart&quot;,
    description: &quot;Render a breakdown of expenses by category.&quot;,
    parameters: expenseChartSchema,
    render: ExpenseChart,
  &#125;);

  return null;
&#125;</code></pre></div><p class="paragraph" style="text-align:left;">The hook registers the tool with CopilotKit&#39;s runtime. The runtime advertises it to the agent over AG-UI. When the agent calls it, the args stream in and your component renders inline. No Python tool to write, no schema to wire, no API route to add.</p><p class="paragraph" style="text-align:left;">Your design system stays in charge.</p><p class="paragraph" style="text-align:left;">That expense chart isn&#39;t a mockup. The <b><span style="text-decoration:underline;"><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/ai-financial-coach-agent?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">AI Financial Coach Agent</a></span></b> renders cards just like it for real budgets, savings plans, and debt payoff.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a472bcaf-c810-4942-bdba-2e62ae865ef5/ezgif-745fa91cbc1343a4.gif?t=1780532184"/></div><p class="paragraph" style="text-align:left;">Want the bare hook first? It&#39;s &#39;<i>use-generative-ui-examples.tsx&#39;</i> in the <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/generative-ui-starter-project?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">Generative UI Starter Project</a></b></span>.</p><p class="paragraph" style="text-align:left;"><b>The token tax</b></p><p class="paragraph" style="text-align:left;">Every component you register sits in the agent&#39;s context window before the user has said anything. A typical tool description with its JSON schema runs around 400 tokens. 25 components are 10,000 tokens on every turn. You pay that tax per request.</p><p class="paragraph" style="text-align:left;">The agent picks the wrong component too. Too many look similar. Pie chart and donut chart both &quot;show proportions.&quot; It guesses.</p><p class="paragraph" style="text-align:left;"><b>When to add agent-side state</b></p><p class="paragraph" style="text-align:left;">Shared state is the one case where writing a Python tool is worth it. The agent writes to session state. Other parts of the UI subscribe and re-render with no second LLM call. Pin a metric, the dashboard updates. Add a row, the table redraws.</p><div class="codeblock"><pre><code>from google.adk.agents import LlmAgent
from google.adk.tools import ToolContext

def pin_metric(tool_context: ToolContext, label: str, value: float) -&gt; dict:
    &quot;&quot;&quot;Pin a metric to the user&#39;s dashboard.&quot;&quot;&quot;
    pinned = tool_context.state.get(&quot;pinnedMetrics&quot;, [])
    tool_context.state[&quot;pinnedMetrics&quot;] = pinned + [&#123;&quot;label&quot;: label, &quot;value&quot;: value&#125;]
    return &#123;&quot;status&quot;: &quot;pinned&quot;&#125;

agent = LlmAgent(name=&quot;dashboard_agent&quot;, model=&quot;gemini-3.5-flash&quot;, tools=[pin_metric])</code></pre></div><p class="paragraph" style="text-align:left;">The frontend reads pinned metrics through CopilotKit&#39;s shared-state hook. The chat component still renders inline because the same tool name is wired with the frontend hook. </p><p class="paragraph" style="text-align:left;">Pin a metric in chat. The panel redraws with no second model call. That&#39;s the <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/ai-dashboard-canvas-agent?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">AI Dashboard Canvas Agent</a></b></span>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/50e2bf8b-cc05-4c92-969b-31d36decce7c/image.png?t=1780532080"/></div><p class="paragraph" style="text-align:left;">The <b><span style="text-decoration:underline;"><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/ai-deep-research-agent?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">AI Deep Research Agent</a></span></b> takes it further. The plan, every search, each file write, all of it streams in as live cards. For everything else, the frontend hook is the whole story.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c963401e-9442-4d0c-8bb8-4d2095ae93b7/ezgif-73e872c02cea0875.gif?t=1780532486"/></div><p class="paragraph" style="text-align:left;"><b>When to ship Controlled:</b> Ten or fewer high-value flows. Design precision matters. You know the exact UIs you need.</p><p class="paragraph" style="text-align:left;"><b>When not to: </b>Your codebase grows linearly with use cases. 25 components means 25 tool definitions sitting in every agent turn.</p><p class="paragraph" style="text-align:left;"><b>What breaks:</b> Agent picks the wrong component. Two tool descriptions overlap semantically. Past 15 tools, two of them probably read like &quot;displays data.&quot; Fix: rewrite descriptions to name the user intent, not the visual. &quot;Use when the user asks to compare proportions of a whole&quot; beats &quot;renders a pie chart.&quot;</p><h2 class="heading" style="text-align:left;"><b>Pattern 2: Declarative (A2UI), agent emits schema</b></h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f4e7cdf6-e242-4e1f-9d61-71838d4583e4/image.png?t=1780532206"/></div><p class="paragraph" style="text-align:left;">This is the pattern most production agent apps end up needing.</p><p class="paragraph" style="text-align:left;">The agent emits a JSON schema describing the UI. Your app has a catalog of components that maps schema nodes to React (or Svelte, Flutter, anything). One tool. Many UIs.</p><p class="paragraph" style="text-align:left;">A2UI is the standard spec. CopilotKit ships the runtime. ADK runs the agent. AG-UI is the wire.</p><p class="paragraph" style="text-align:left;">The agent tool returns three operations in order: create a surface, push the component tree, push the data.</p><div class="codeblock"><pre><code>def search_flights(flights: list[Flight]) -&gt; dict[str, Any]:
    &quot;&quot;&quot;Search flights and display them as rich cards.&quot;&quot;&quot;
    return &#123;
        &quot;a2ui_operations&quot;: [
            &#123;&quot;type&quot;: &quot;create_surface&quot;, &quot;surfaceId&quot;: SURFACE_ID, &quot;catalogId&quot;: CATALOG_ID&#125;,
            &#123;&quot;type&quot;: &quot;update_components&quot;, &quot;surfaceId&quot;: SURFACE_ID, &quot;components&quot;: FLIGHT_SCHEMA&#125;,
            &#123;&quot;type&quot;: &quot;update_data_model&quot;, &quot;surfaceId&quot;: SURFACE_ID, &quot;data&quot;: &#123;&quot;flights&quot;: flights&#125;&#125;,
        ]
    &#125;</code></pre></div><p class="paragraph" style="text-align:left;">The component tree above lives in flights.json. You wrote it. The agent only fills in the data. That&#39;s a fixed schema.</p><p class="paragraph" style="text-align:left;">Dynamic schema flips it: a secondary LLM writes the component tree per turn from conversation context. Same a2ui_operations container at the end. The Google ADK showcase ships both.</p><p class="paragraph" style="text-align:left;"><b>The catalog is the contract</b></p><p class="paragraph" style="text-align:left;">Definitions list the components the agent is allowed to emit, with Zod schemas for the props. Renderers fill in React. Typos become build errors instead of blank screens.</p><div class="codeblock"><pre><code>const renderers: CatalogRenderers&lt;TravelDefinitions&gt; = &#123;
  FlightCard: (&#123; props &#125;) =&gt; (
    &lt;article className=&quot;rounded-xl border p-4&quot;&gt;
      &lt;header className=&quot;flex justify-between&quot;&gt;
        &lt;span&gt;&#123;(props as any).airline&#125;&lt;/span&gt;
        &lt;span&gt;&#123;(props as any).price&#125;&lt;/span&gt;
      &lt;/header&gt;
      &lt;div className=&quot;text-sm text-muted-foreground&quot;&gt;
        &#123;(props as any).origin&#125; → &#123;(props as any).destination&#125; · &#123;(props as any).departureTime&#125;
      &lt;/div&gt;
    &lt;/article&gt;
  ),
&#125;;

export const travelCatalog = createCatalog(travelDefinitions, renderers, &#123;
  catalogId: &quot;copilotkit://travel-catalog&quot;,
  includeBasicCatalog: true,
&#125;);</code></pre></div><p class="paragraph" style="text-align:left;">Both halves live in the <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/generative-ui-starter-project?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">Generative UI Starter Project</a></b></span>, wired and matched. search_flights in <i>&#39;a2ui_fixed_</i><i><a class="link" href="https://schema.py?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow">schema.py</a></i><i>&#39;</i>, the FlightCard catalog in <i>&#39;renderers.tsx</i>&#39;. Ask for flights. Watch the cards stream into chat.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/02ac044e-0776-44a2-a5f6-24d8fc2567d4/ezgif-7659c7a92e59344a.gif?t=1780532462"/></div><p class="paragraph" style="text-align:left;">Buttons and other interactive components carry an action in the schema. The basic catalog wires it to onClick. Click fires an event back to the agent over AG-UI. The agent decides what to render next. Zero click handlers.</p><p class="paragraph" style="text-align:left;"><b>The token math</b></p><p class="paragraph" style="text-align:left;">50 card types or 500, the agent sees one function. Tokens per turn stay flat as your component library grows.</p><p class="paragraph" style="text-align:left;">Extensible to any rendering framework because it&#39;s just JSON. Any agent that already speaks AG-UI can drive A2UI on day zero. You don&#39;t touch agent code to wire this up.</p><p class="paragraph" style="text-align:left;"><b>Trade-off: </b>The LLM owns the layout. Output varies run to run within your catalog. If you&#39;re shipping legal disclosures, marketing surfaces, or anything where exact pixel placement matters, this is not your bucket.</p><p class="paragraph" style="text-align:left;">Declarative is the pattern built for the long tail. Dashboards, results, forms, cards, widgets.</p><p class="paragraph" style="text-align:left;"><b>When to ship Declarative:</b> You have more use cases than time to pre-build. You care about token economics past the prototype stage.</p><p class="paragraph" style="text-align:left;"><b>What breaks:</b> Built a custom FlightCard. Every flight renders as the basic catalog&#39;s generic card. No error in the console. The CATALOG_ID on the agent and catalogId in createCatalog on the frontend don&#39;t match. Frontend doesn&#39;t recognize the catalog the agent is targeting, falls back to basic. Match the strings exactly on both sides.</p><h2 class="heading" style="text-align:left;"><b>Pattern 3: Open-ended, no catalog, no rules</b></h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/23782a72-b59b-4fa7-9418-4c7b74e49791/image.png?t=1780532583"/></div><p class="paragraph" style="text-align:left;">The third pattern is the opposite extreme. No catalog. No schema. Just a blank canvas.</p><p class="paragraph" style="text-align:left;">Two sub-patterns live in this bucket.</p><p class="paragraph" style="text-align:left;"><b>MCP Apps</b></p><p class="paragraph" style="text-align:left;">An MCP server exposes UI surfaces that the agent drives. Excalidraw is the example that stuck with me. The agent gets full control of the canvas. Draws diagrams from your context. Owns every pixel on the board.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bfa32d2f-d7f5-4248-88fa-bef5c2cf1f7f/image.png?t=1780532597"/></div><p class="paragraph" style="text-align:left;">Implementing the client protocol from scratch is painful, so CopilotKit ships an <i>MCPAppsMiddleware</i>. Attach it to your agent and point it at any MCP Apps server.</p><div class="codeblock"><pre><code>const agent = new BuiltInAgent(&#123;
  model: &quot;openai/gpt-5.5&quot;,
  prompt: &quot;You are a helpful assistant.&quot;,
&#125;).use(
  new MCPAppsMiddleware(&#123;
    mcpServers: [&#123; type: &quot;http&quot;, url: &quot;https://mcp.excalidraw.com/mcp&quot;, serverId: &quot;my-server&quot; &#125;],
  &#125;),
);</code></pre></div><p class="paragraph" style="text-align:left;">Spin up the <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/mcp-apps-generative-ui-showcase?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">MCP Apps Showcase</a></b></span><span style="text-decoration:underline;"><b> </b></span>and you&#39;re booking flights and reserving hotels inside the chat window. Same middleware, real MCP servers. Or go further.</p><p class="paragraph" style="text-align:left;">The <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents/ai-mcp-app-builder?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">AI MCP App Builder</a></b></span><span style="text-decoration:underline;"><b> </b></span>lets the agent write a brand-new app into an E2B sandbox, then renders it live.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2f1d5956-4540-4697-a36e-fbcf2bafd0f9/WEGGWcRPy3RhPaAV.jpg?t=1780531876"/></div><p class="paragraph" style="text-align:left;"><b>Sandboxed HTML</b></p><p class="paragraph" style="text-align:left;">The agent writes raw HTML. Your app renders it inside a sandboxed iframe so it can&#39;t hijack the session.</p><p class="paragraph" style="text-align:left;">The runtime registers an HTML rendering tool and ships it to the agent over AG-UI. The agent calls it with whatever markup it wants. There is no HTML tool to define on the agent side. The runtime injects it.</p><p class="paragraph" style="text-align:left;">Agent-side instruction is doing real work:</p><div class="codeblock"><pre><code>canvas_agent = LlmAgent(
    name=&quot;canvas_agent&quot;,
    model=&quot;gemini-3.5-flash&quot;,
    instruction=(
        &quot;You are a visualization assistant. When the user asks to see, &quot;
        &quot;draw, or visualize anything, generate an interactive HTML UI. &quot;
        &quot;Use Tailwind classes only. No external fonts. Stick to neutral &quot;
        &quot;colors unless the user names one.&quot;
    ),
)</code></pre></div><p class="paragraph" style="text-align:left;">Without those style rules, the model defaults to whatever aesthetic was loudest in its training data that week. With them, you get something close to your brand most of the time. Not always.</p><p class="paragraph" style="text-align:left;"><b>The brand inconsistency problem</b></p><p class="paragraph" style="text-align:left;">I tried shipping Open-ended as the primary UI for an agent. Pulled it in a week.</p><p class="paragraph" style="text-align:left;">&quot;Neo-brutalist&quot; on Tuesday. &quot;iOS 4 clone&quot; on Wednesday. Style rules in the prompt nudge the agent toward your brand. They don&#39;t guarantee it. The brand kept changing. The product felt unserious.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bb2c6e2f-ebce-4d7d-b3e8-e8555f97e51d/image.png?t=1780532628"/></div><p class="paragraph" style="text-align:left;">Open-ended isn&#39;t useless. It&#39;s misapplied.</p><p class="paragraph" style="text-align:left;">Right call for one thing: throwaway interactions where the user doesn&#39;t care what the interface looks like and will never see it again. &quot;Show me how electrons work.&quot; &quot;Give me a weird bar chart of my last 10 queries.&quot; &quot;Visualize this API response.&quot; The kind of thing you see in Google AI overviews.</p><p class="paragraph" style="text-align:left;"><b>When to ship Open-ended:</b> One-shot queries. Disposable visualizations. Sandboxed experiments. Never as the primary surface.</p><p class="paragraph" style="text-align:left;"><b>What breaks:</b> The iframe renders. Buttons don&#39;t click. Forms don&#39;t submit. Sandbox flags are too tight, or too loose in a way the browser refuses. Set the iframe sandbox to allow scripts and allow forms. Nothing else. Never allow-same-origin.</p><h2 class="heading" style="text-align:left;"><b>How to pick</b></h2><p class="paragraph" style="text-align:left;">Run the decision tree before you write code.</p><p class="paragraph" style="text-align:left;">Designer has pixel-perfect mockups for this flow? Controlled.</p><p class="paragraph" style="text-align:left;">Dozens of card types or widgets to ship? Declarative.</p><p class="paragraph" style="text-align:left;">One-shot, throwaway visualization the user will never see twice? Open-ended.</p><p class="paragraph" style="text-align:left;">Can&#39;t decide? Default to Declarative. Upgrade to Controlled for the top 3 flows. Never Open-ended as the default.</p><p class="paragraph" style="text-align:left;">If you&#39;re already shipping and not sure where you landed, count the render tools. Past 15, you&#39;re in Controlled and the wall is close. Start wiring A2UI this week.</p><h2 class="heading" style="text-align:left;"><b>Three patterns. Three bets.</b></h2><p class="paragraph" style="text-align:left;">Controlled bets on you. Pre-built components, pixel-perfect. Expensive past 25 of them.</p><p class="paragraph" style="text-align:left;">Declarative bets on the schema. The schema is the contract. The agent fills it in. Scales flat.</p><p class="paragraph" style="text-align:left;">Open-ended bets on the model. No catalog, no schema, raw HTML. Good for throwaway. Brittle for anything that ships twice.</p><p class="paragraph" style="text-align:left;">The mistake isn&#39;t picking the wrong pattern. It&#39;s not knowing you picked one.</p><p class="paragraph" style="text-align:left;">Most teams default to Controlled because the framework defaults to Controlled. They hit the wall at 25 components and reach for Open-ended because it looks compelling in demos. Neither was a decision. Both were drift.</p><p class="paragraph" style="text-align:left;">Pick on purpose. Match the pattern to the problem. Controlled for the flows that need to be exact. Declarative for the long tail. Open-ended for the disposable.</p><p class="paragraph" style="text-align:left;">🚨<b> </b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow"><b>Open Source Generative UI Agent Templates</b></a></p><p class="paragraph" style="text-align:left;">The reference for all three lives in the new <span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/generative_ui_agents?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #0f1419">Generative UI Agents</a></b></span><span style="text-decoration:underline;"><b> </b></span>section of awesome-llm-apps. Clone what you need. Rip out what you don&#39;t.</p><hr class="content_break"><p class="paragraph" style="text-align:left;">I&#39;ll be publishing more about shipping agents in production, AG-UI, and the patterns that scale. </p><p class="paragraph" style="text-align:left;"><b>Follow me </b><b><a class="link" href="https://twitter.com/Saboo_Shubham_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow">@Saboo_Shubham_</a></b><b> to stay tuned.</b></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">We share in-depth blogs and tutorials like this 2-3 times a week, to help you stay ahead in the world of AI. <span style="text-decoration:underline;"><b><a class="link" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">If you&#39;re serious about leveling up your AI skills and staying ahead of the curve, subscribe now and be the first to access our latest tutorials.</a></b></span></p><p class="paragraph" style="text-align:left;"><b>Don’t forget to share this tutorial on your social channels and tag Unwind AI (</b><span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span><b>) to support us!</b></p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=generative-ui-is-the-new-frontend"><span class="button__text" style=""> Subscribe now for FREE - Get instant access to more LLM, RAG & AI Agent tutorials </span></a></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=d832bac9-409e-44b7-81fe-e065a8760e9e&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>OpenAI Codex Can Now Ship Live Shareable Websites</title>
  <description>+ Hermes Agent goes native on your desktop</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ce43dc75-210c-4222-8d9e-d57ed54f4a8c/upload_cd4242ee92213ab2.jpg" length="95354" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/openai-codex-can-now-ship-live-shareable-websites</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/openai-codex-can-now-ship-live-shareable-websites</guid>
  <pubDate>Wed, 03 Jun 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-06-03T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>OpenAI Codex Can Now Ship Live Websites</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Hermes Agent now has a desktop app</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Open-source Alternative to Exa Websets</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Microsoft launches 7 in-house MAI models at Build</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Open-source code review tool from the OpenClaw team</b></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">HTTP is a primitive. JSON is a primitive. /goal is becoming one for coding agents.</p><p class="paragraph" style="text-align:left;">A few weeks ago, OpenAI&#39;s Codex CLI added /goal as a way to give the coding worker a job with a defined done state. Claude Code added it this week.</p><p class="paragraph" style="text-align:left;">Hermes Agent, the orchestrator I run on a Mac Mini to coordinate work between coding workers, has had /goal built in for a while.</p><p class="paragraph" style="text-align:left;">This guide walks through what /goal actually is, the three roles in a multi-agent setup, a real end-to-end run, the verification rule, and how to run goals in parallel without workers stepping on each other. </p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Read The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://openai.com/index/codex-for-every-role-tool-workflow/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow"><b>OpenAI Codex Can Now Ship Live Websites</b></a></h3><div class="image"><a class="image__link" href="https://openai.com/index/codex-for-every-role-tool-workflow/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28d8c157-b2a8-4bd0-acb5-d11ec1d06d9b/Notion__1_.gif?t=1780468720"/></a></div><p class="paragraph" style="text-align:left;">Your Codex session doesn&#39;t have to end with a file sitting on your machine anymore.</p><p class="paragraph" style="text-align:left;"><b>Sites</b> lets Codex publish its work as a hosted, interactive website with a shareable URL. Dashboards, scenario planners, project trackers, launch hubs, whatever you&#39;re building, it goes live with a link you hand to your team. </p><p class="paragraph" style="text-align:left;">You can even ask Codex to keep the site up to date as things change. Not static pages either. These are collaborative canvases your whole workspace can explore and contribute to.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Annotations for precise refinement</b>: Point to the exact part of a site, doc, spreadsheet, or slide you want changed. Codex updates just that piece without starting over. Think inline editing, but AI-powered.</p></li><li><p class="paragraph" style="text-align:left;"><b>Rolling out now</b>: Sites are in preview for Business and Enterprise teams, expanding broadly soon.</p></li></ol><p class="paragraph" style="text-align:left;">And it&#39;s not just Sites. OpenAI is going really heavy on making Codex the tool for non-technical knowledge work.</p><p class="paragraph" style="text-align:left;">They shipped six <b><a class="link" href="https://github.com/openai/role-based-plugins?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">open-source, role-specific plugins</a></b> that each bundle apps, skills, and workflows for a specific job function. Data Analytics connects Snowflake, Databricks, Hex, and Tableau. Creative Production hooks into Figma, Canva, and Picsart. Sales brings in Salesforce, HubSpot, and Clay. There&#39;s also Product Design, Equity Investing, and Investment Banking. </p><p class="paragraph" style="text-align:left;">62 apps and 110 skills across all six. </p><p class="paragraph" style="text-align:left;"><b>Plugins work out of the box, but you own them</b>. Every plugin can be adapted to your team&#39;s workflows. Build and share custom ones too. Corporate Finance, Private Equity, Marketing Strategy, and Legal plugins are coming next.</p><p class="paragraph" style="text-align:left;">Non-developers already make up 20% of Codex&#39;s 5 million weekly users and are growing 3x faster than developers. OpenAI is clearly building for that curve.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://hermes-agent.nousresearch.com/desktop?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Hermes Desktop is Here</a></b></h3><div class="image"><a class="image__link" href="https://hermes-agent.nousresearch.com/desktop?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a8b06dd4-818f-4c32-99c0-c3e3868d71f1/ezgif-2d3bf08c0d6a8647.gif?t=1780469136"/></a></div><p class="paragraph" style="text-align:left;"><b>Hermes Agent</b> just got a native desktop app on macOS and Windows.</p><p class="paragraph" style="text-align:left;">First demoed during Jensen Huang&#39;s GTC keynote, <b>Hermes Desktop</b> is now in public preview. Same agent, same memory, same skills, same everything, just no terminal required. Download the .dmg or .exe, and you&#39;re running.</p><p class="paragraph" style="text-align:left;">The community has been using Hermes as a single interface for all their workflows, and has been actively growing the ecosystem around it. A native app lowers the floor for everyone who wants in but doesn&#39;t live in a terminal. If you&#39;re already running Hermes via CLI or messaging platforms, nothing changes. Desktop is just another surface, same agent underneath.</p><p class="paragraph" style="text-align:left;">Download at <a class="link" href="https://hermes-agent.nousresearch.com/desktop?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">hermes-agent.nousresearch.com/desktop</a>. macOS 12+, Windows 10/11, or install via terminal on Linux.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/tinyfish-io/bigset?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow"><b>Open-source Alternative to Exa Websets</b></a></h3><div class="image"><a class="image__link" href="https://github.com/tinyfish-io/bigset?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b6ab27d5-85af-470b-aa58-c83aff379db6/image.png?t=1780469235"/></a></div><p class="paragraph" style="text-align:left;">Describe the dataset you want in one sentence. AI agents go build it for you.</p><p class="paragraph" style="text-align:left;"><b>BigSet</b> is a new open-source tool from <a class="link" href="https://accounts.tinyfish.ai/api-keys?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow"><b>TinyFish</b></a> that turns a natural language prompt into a structured, verified dataset pulled from the live web. </p><p class="paragraph" style="text-align:left;">Say &quot;YC companies currently hiring engineers, with their funding stage, location, and number of open roles&quot; and BigSet infers the schema, fans out AI agents to research in parallel, deduplicates, and returns a clean table with citations that you can export as CSV or XLSX.</p><p class="paragraph" style="text-align:left;">The real fun bit is you can set a refresh cadence (30 minutes to weekly) so the dataset stays fresh always. </p><p class="paragraph" style="text-align:left;">Self-hosted via Docker in one command. </p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Schema inference from English</b>: You describe what you want, BigSet figures out column names, types, and primary keys. No manual schema design.</p></li><li><p class="paragraph" style="text-align:left;"><b>Parallel agent research</b>: Multiple AI agents fan out across the web simultaneously, verify data against real sources, and deduplicate before returning results.</p></li><li><p class="paragraph" style="text-align:left;"><b>Auto-refresh schedules</b>: Set it and forget it. Datasets update on a cadence you choose, from every 30 minutes to weekly.</p></li><li><p class="paragraph" style="text-align:left;"><b>Full stack, open source</b>: Next.js 16 frontend, Fastify backend, Mastra workflows for agent orchestration, powered by TinyFish&#39;s Search and Fetch APIs under the hood. AGPL-3.0 licensed.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://microsoft.ai/news/introducingmai-code-1-flash/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Microsoft finally built its own frontier models</a></b><b> </b><br>Microsoft dropped the new <b>MAI family of models</b> at Build, across text, image, voice, and speech. <b>MAI-Code-1-Flash</b> is the one to watch for devs, optimized specifically for fast, efficient coding tasks. <b>MAI-Thinking-1</b> handles heavier reasoning and SWE work. Even the Image model is debuting at No. 3 on <a class="link" href="https://Arena.ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Arena.ai</a>. Both are cheaper alternatives to OpenAI and Anthropic models. The word on the street is that Microsoft built these because relying on Anthropic&#39;s Claude was forcing them to raise GitHub Copilot prices and cap developer usage.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://factory.ai/news/factory-router?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Your coding agent doesn&#39;t need the most expensive model for every task</a></b><b> </b>Factory just shipped Factory Router, and the idea is overdue: stop burning frontier-model tokens on tasks that a smaller model handles just as well. Router automatically picks the right model for each coding session and escalates to a more capable one only if the first choice struggles. On their benchmarks, it hits 99% of Claude Opus 4.7&#39;s pass rate at 20% lower cost. If the first model can&#39;t crack it, Router bumps to a heavier one automatically. Available now in Factory CLI and Desktop in private research preview.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/perplexity_ai/status/2061861293569765847?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Perplexity Computer splits work between local and cloud</a></b><b> </b><br>Perplexity announced hybrid agentic inference for Perplexity Computer: the system can now split tasks between a local model running on your machine and frontier models in the cloud. Private data stays on-device, token efficiency goes up, and you stop sending everything through an API. Coming soon, but architecturally, this is the direction everyone&#39;s heading.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/jpschroeder/status/2061484426387677268?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Use Cursor&#39;s Composer 2.5 in any agent harness</a></b><b> </b><br>Someone built an open-source macOS app that exposes Cursor&#39;s Composer 2.5 as an API. That means you can now use Cursor&#39;s model routing in Codex, OpenCode, Cline, or whatever harness you prefer. If you&#39;ve been locked into Cursor&#39;s editor just for the model quality, this unbundles it.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://clawpatch.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">ClawPatch</a></b>: Open-source code review from the OpenClaw team that thinks in &quot;feature slices&quot; instead of files. It maps your codebase into semantic units (routes, commands, packages), sends bounded context to an AI for review, and then runs an explicit fix loop. Every finding gets a severity, confidence score, and audit trail. Works with Codex as the default AI provider.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/nesquena/hermes-webui?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Hermes WebUI</a></b>: A 12.7K-star open-source web interface for Hermes Agent by Nathan Esquenazi (CodePath co-founder). Full CLI parity in a three-panel browser layout: sessions on the left, chat in the center, workspace file browser on the right. No build step, no framework, just Python and vanilla JS. MIT-licensed.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/leodev/status/2061417039949099205?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Email SDK</a></b>: Unified TypeScript SDK that lets you send email through any provider: Resend, Postmark, SendGrid, Mailgun, Brevo, or raw SMTP. One clean API, swap providers by changing a config line. Built-in formatting, error handling, and type safety so you stop writing provider-specific glue code.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (111k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=openai-codex-can-now-ship-live-shareable-websites"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=8cf1ddf1-7902-466a-8b4a-e67bf87950a3&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Every Software Just Became Agent-Native</title>
  <description>+ Self-Evolving Agent Skills by Microsoft</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/27cc945e-efa3-473a-8860-c9478fcb4256/upload_25ff3d71f7576148.jpg" length="98756" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/every-software-just-became-agent-native</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/every-software-just-became-agent-native</guid>
  <pubDate>Thu, 28 May 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-05-28T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>CLI-Anything: Every Software Just Became Agent-Native</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Shared AI agents for teams in Slack</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Microsoft SkillOpt: Train the Skill, not the model</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Inference price war is here (and it’s starting from China)</b></p></li><li><p class="paragraph" style="text-align:left;"><b>LangChain gives agents a code layer between tool calls</b></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">HTTP is a primitive. JSON is a primitive. /goal is becoming one for coding agents.</p><p class="paragraph" style="text-align:left;">A few weeks ago, OpenAI&#39;s Codex CLI added /goal as a way to give the coding worker a job with a defined done state. Claude Code added it this week.</p><p class="paragraph" style="text-align:left;">Hermes Agent, the orchestrator I run on a Mac Mini to coordinate work between coding workers, has had /goal built in for a while.</p><p class="paragraph" style="text-align:left;">This guide walks through what /goal actually is, the three roles in a multi-agent setup, a real end-to-end run, the verification rule, and how to run goals in parallel without workers stepping on each other. </p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Read The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://github.com/HKUDS/CLI-Anything?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Every Software Just Became Agent-Native</a></b></h3><div class="image"><a class="image__link" href="https://github.com/HKUDS/CLI-Anything?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a6ebc981-aafe-4d33-8682-6ca32be8635e/Screenshot_2026-05-27_at_11.40.28_PM.png?t=1779950434"/></a></div><p class="paragraph" style="text-align:left;">Your agent can write code, search the web, and manage files. But ask it to edit a Blender scene, export a MuseScore sheet, or automate Rekordbox, and it hits a wall. The software doesn&#39;t speak agent.</p><p class="paragraph" style="text-align:left;">CLI-Anything from the HKUDS lab fixes this by generating full CLI harnesses for any software, turning GUI-only apps into agent-controllable tools. One command analyzes the target app&#39;s source code, architects a CLI, implements it with tests, and publishes it to PATH. The project ships with a growing registry of 50+ ready-made CLIs covering GIMP, Blender, LibreOffice, OBS, Obsidian, Kdenlive, QGIS, and more.</p><p class="paragraph" style="text-align:left;">The idea is simple but the implications are huge: if CLI is the universal interface both humans and LLMs already speak, then wrapping every piece of software in a CLI makes the entire software ecosystem agent-accessible overnight.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>CLI-Hub package manager</b>: pip install cli-anything-hub, then browse, search, and install any harness with cli-hub install &lt;name&gt;. Supports pip, npm, brew, and system tools.</p></li><li><p class="paragraph" style="text-align:left;"><b>7-phase generation pipeline</b>: Point it at a repo or app and it runs through analyze, design, implement, plan tests, write tests, document, and publish, fully automated by your coding agent.</p></li><li><p class="paragraph" style="text-align:left;"><b>Works with every major agent</b>: Claude Code plugin, Pi extension, OpenCode commands, Codex, and OpenClaw skill. Each gets a native integration path.</p></li><li><p class="paragraph" style="text-align:left;"><b>Skills baked in</b>: Every generated CLI ships with a <a class="link" href="https://SKILL.md?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">SKILL.md</a> so agents can discover and use it autonomously, no manual wiring needed.</p></li><li><p class="paragraph" style="text-align:left;"><b>Try it now</b>: Install from PyPI or clone the repo. The CLI-Hub web registry is live at <a class="link" href="https://clianything.cc?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">clianything.cc</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://github.com/paradigmxyz/centaur?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Shared AI agents for teams in Slack</a></b></h3><div class="image"><a class="image__link" href="https://github.com/paradigmxyz/centaur?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f890b5a7-aafb-4b23-80e9-4380b934c979/Screenshot_2026-05-27_at_11.41.43_PM.png?t=1779950507"/></a></div><p class="paragraph" style="text-align:left;">We love our personal Hermes and OpenClaws. But how many of us have been running it for our professional work, with our teams, cross-functionally? </p><p class="paragraph" style="text-align:left;">It’s a completely different set of problems: surviving laptop closures, handling real credentials securely, running for hours or days, and being reachable where the team actually works. </p><p class="paragraph" style="text-align:left;">Centaur is the self-hosted runtime Paradigm and Tempo have been running internally since January, now open-sourced under Apache 2.0. It&#39;s a Slack-native multiplayer agent system where: </p><ul><li><p class="paragraph" style="text-align:left;">every thread gets its own isolated Kubernetes sandbox, </p></li><li><p class="paragraph" style="text-align:left;">tools you add are instantly available to every conversation, and </p></li><li><p class="paragraph" style="text-align:left;">a credential firewall injects secrets in-flight so agents can never exfiltrate raw keys.</p></li></ul><p class="paragraph" style="text-align:left;">Tools are plain Python drop-ins that hot-reload across your org, workflows checkpoint to Postgres and resume exactly where they left off after a crash, and every night the system reviews its own performance and ships fixes to its own skills (super interesting!). </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/paradigmxyz/centaur?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Clone from GitHub</a> or visit <a class="link" href="https://centaur.run?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">centaur.run</a> to get started.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/microsoft/SkillOpt?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow"><b>Microsoft SkillOpt: Train the Skill, Not the Model</b></a></h3><div class="image"><a class="image__link" href="https://github.com/microsoft/SkillOpt?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c3a18eb4-c5d4-4d07-9687-31730403ff68/Screenshot_2026-05-27_at_11.43.09_PM.png?t=1779950594"/></a></div><p class="paragraph" style="text-align:left;">What if you could train agent skills the same way you train neural networks, with learning rates, mini-batches, epochs, and momentum, but entirely in text space?</p><p class="paragraph" style="text-align:left;">SkillOpt from Microsoft Research does exactly that. Instead of fine-tuning model weights, it treats SKILL.md as a trainable external parameter. The frozen target model executes tasks, records scored trajectories, and a separate optimizer model proposes structured edits to the skill. Edits are accepted only when held-out validation performance improves. </p><p class="paragraph" style="text-align:left;">The whole thing mirrors a training loop: rollouts are forward passes, reflection is a backward pass, and a textual edit budget acts as a learning rate to prevent destructive rewrites.</p><p class="paragraph" style="text-align:left;">Evaluated across 6 benchmarks and 7 models, including real agent execution loops with Codex and Claude Code, SkillOpt achieves best or tied-best results in all 52 settings tested.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Improvement with Claude Code</b>: On GPT-5.5 target model running through Claude Code, SkillOpt skills boosted performance by an average of 18.6 points across benchmarks, with Spreadsheet tasks jumping +58.3%.</p></li><li><p class="paragraph" style="text-align:left;"><b>Cross-model and cross-harness transfer</b>: A skill trained with Codex transfers directly into Claude Code and gains +31.8% on SpreadsheetBench. Trained on GPT-5.4, transfers to GPT-5.4-nano and still gains +15.2%.</p></li><li><p class="paragraph" style="text-align:left;"><b>Self-optimizer mode works</b>: Even when the target model is its own optimizer, the constrained, validated update loop still discovers useful edits.</p></li><li><p class="paragraph" style="text-align:left;"><b>Exports a single file</b>: The whole optimization produces one best_skill.md file. The target model at deployment never sees the optimizer memory, rejected edits, or training state.</p></li><li><p class="paragraph" style="text-align:left;"><b>Open-source</b>: The whole thing is open-sourced under MIT license. Go and try it out!</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/kimmonismus/status/2059578380329394292?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">The inference price war is here</a></b><br>DeepSeek just made its 75% price cut on V4-Pro permanent. Xiaomi&#39;s MiMo slashed V2.5 pricing by up to 99%, effective today. But this isn&#39;t a loss-leader race to the bottom. V4-Pro&#39;s hybrid attention architecture compresses its KV cache at 1M tokens to 10% of V3.2&#39;s, with single-token inference FLOPs at 27% of previous. V4-Pro now sits at $0.87 per million output tokens. A year ago, sub-dollar output pricing meant you were using a small distilled model with real capability tradeoffs. These are frontier-class reasoners.</p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/ElevenLabs/status/2059312414198235642?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">ElevenLabs Launches Music v2</a></b><br>ElevenLabs just shipped Music v2 with better vocals, instrumentation, and arrangement across every genre, plus improved multilingual support and capabilities that weren&#39;t possible before. If you&#39;ve used their v1 for music generation, this is a huge upgrade.</p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://cohere.com/blog/command-a-plus?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Cohere Drops Command A+: 218B MoE Under Apache 2.0</a></b><br>Cohere just open-sourced Command A+, a 218B parameter MoE model with only 25B active per token. It unifies all previous Command A variants (reasoning, vision, translation) into a single model, supports 48 languages, and runs on as little as two H100s at W4A4 quantization. On τ²-Bench Telecom, it jumped from 37% to 85% over Command A Reasoning. Apache 2.0, available on Hugging Face in BF16, FP8, and W4A4.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.microsoft.com/en-us/research/articles/fara1-5-computer-use-agent/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow"><b>Microsoft Open-Sources Fara1.5 Browser Agents</b></a><br>Microsoft Research just dropped Fara1.5, a family of three open computer use agent models (4B, 9B, 27B) built on Qwen3.5 for browser automation. The 27B variant hits 72% on Online-Mind2Web, outperforming OpenAI Operator, Gemini 2.5 Computer Use, and Yutori Navigator n1. Even the 9B model at 63.4% beats every proprietary competitor. Available on the Microsoft Foundry now.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://developers.openai.com/api/docs/guides/secure-mcp-tunnels?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow"><b>OpenAI Launches Secure MCP Tunnels</b></a><br>Your private MCP servers can now stay inside your network while ChatGPT, Codex, and the Responses API connect through outbound-only HTTPS. No inbound ports, no public endpoints, no VPN. If you&#39;ve been holding off on connecting internal tools to OpenAI products because of network security concerns, this removes that blocker.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.langchain.com/blog/give-your-agents-an-interpreter?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow"><b>LangChain Gives Agents a Code Layer Between Tool Calls</b></a><br>Your agent calls a tool, reads the result, reasons, calls the next tool, reads, reasons, repeat. Every step is a model round trip. LangChain&#39;s Deep Agents now ships with interpreters — small QuickJS runtimes where the agent writes code that coordinates multiple tool calls, keeps intermediate state in the runtime, and returns only what matters. Early testing showed up to 35% fewer tokens on some tasks. Available in both Python and TypeScript.</p></div><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/colbymchenry/codegraph?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Codegraph</a></b>: Pre-indexed code knowledge graph for Claude Code, Codex, Cursor, and more. Agents query symbol relationships and call graphs instead of scanning files, averaging 35% cheaper and 70% fewer tool calls. 100% local, MIT licensed.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/perplexityai/bumblebee?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow"><b>Bumblebee</b></a>: Perplexity&#39;s open-source supply chain scanner for developer machines. A single Go binary that checks lockfiles, package metadata, extension manifests, and MCP configs against exposure catalogs. Apache 2.0.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://apps.apple.com/us/app/sieve-secret-scanner/id6767409365?mt=12&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Sieve</a></b>: macOS app that scans your Claude Code, Cursor, Copilot, Windsurf, and Codex chat history for accidentally leaked API keys, tokens, and passwords. Ships with an MCP server so Claude can check for exposed secrets itself. $9.99.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (111k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=every-software-just-became-agent-native"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=f50074a4-f6da-4920-ae86-db783084ba74&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Stop giving agents the whole computer</title>
  <description>+ GitHub Spec Kit, Qwen 3.7 Max</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6a901842-ce27-4800-9249-6577ad72fa05/upload_aa2ad2faceccec61.jpg" length="97053" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/stop-giving-agents-the-whole-computer</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/stop-giving-agents-the-whole-computer</guid>
  <pubDate>Fri, 22 May 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-05-22T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">I’ve been thinking a lot about how much room we should actually give coding agents to work.  </p><p class="paragraph" style="text-align:left;">Qwen3.7-Max running for 35 hours with 1,000+ tool calls makes long-horizon agents feel a lot more real. But today’s npm compromise is the less fun side of the same story: attackers are now targeting Claude Code and Codex hooks directly.  </p><p class="paragraph" style="text-align:left;">So the takeaway is pretty simple. Agents are getting better at doing real work, but the workflows around them need stricter specs, better memory, and tighter boundaries before we hand them bigger jobs.</p><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Qwen3.7-Max: 35 hours, 1,000+ tool calls, zero human intervention</b></p></li><li><p class="paragraph" style="text-align:left;"><b>GitHub Spec Kit forces AI to spec before it codes</b></p></li><li><p class="paragraph" style="text-align:left;"><b>314 npm packages compromised to hijack your coding agent</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Google’s own version of hosted Hermes/OpenClaw</b></p></li><li><p class="paragraph" style="text-align:left;"><b>A design skill that refuses to look AI-generated</b></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">HTTP is a primitive. JSON is a primitive. /goal is becoming one for coding agents.</p><p class="paragraph" style="text-align:left;">A few weeks ago, OpenAI&#39;s Codex CLI added /goal as a way to give the coding worker a job with a defined done state. Claude Code added it this week.</p><p class="paragraph" style="text-align:left;">Hermes Agent, the orchestrator I run on a Mac Mini to coordinate work between coding workers, has had /goal built in for a while.</p><p class="paragraph" style="text-align:left;">This guide walks through what /goal actually is, the three roles in a multi-agent setup, a real end-to-end run, the verification rule, and how to run goals in parallel without workers stepping on each other. </p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Read The Ultimate Guide to /goal</a></b></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://qwen.ai/blog?id=qwen3.7&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Qwen3.7-Max: 35 Hours, 1,000+ Tool Calls, Zero Human Intervention</a></b></h3><div class="image"><a class="image__link" href="https://qwen.ai/blog?id=qwen3.7&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/114869bf-286e-4721-a018-6db53a8bfce7/image.png?t=1779346463"/></a></div><p class="paragraph" style="text-align:left;">What’s the maximum number of steps and tool calls you’ve seen an LLM doing without you babysitting? 20? 50? Max 100? </p><p class="paragraph" style="text-align:left;">Qwen3.7-Max just ran a fully autonomous kernel optimization session for 35 hours straight, making over 1,000 tool calls.</p><p class="paragraph" style="text-align:left;">Alibaba&#39;s Qwen Team released their latest model Qwen 3.7-Max, specifically for the agent era. It tops SWE-Pro at 60.6% (vs Opus 4.6&#39;s 48.2%), leads TerminalBench, and takes the crown on MCP-Mark. It also tops all the benchmarks on pure reasoning; best-in-class!</p><p class="paragraph" style="text-align:left;">What makes it genuinely different is scaffold generalisation. You can plug it into Claude Code, OpenClaw, Hermes Agent, or Qwen Code and get consistent results without prompt gymnastics.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Reward hacking defense built in</b>: During 80+ hours of RL training on SWE tasks, the model&#39;s monitoring system autonomously caught 1,618 reward hacking attempts and generated 13 new heuristic rules to block them. The model is training itself to be honest.</p></li><li><p class="paragraph" style="text-align:left;"><b>1M token context, 65K output</b>: Scores 90.4% on MRCR-v2 128K, far ahead of every competitor on long-context retrieval.</p></li><li><p class="paragraph" style="text-align:left;"><b>48 languages natively</b>: Leads multilingual benchmarks across the board, including WMT24++ translation and MMLU-ProX.</p></li><li><p class="paragraph" style="text-align:left;"><b>Pricing</b>: Qwen 3.7 Max is roughly half the price of GPT-5.4 and less than a third of Claude Opus 4.6, while matching both on SWE-Pro and TerminalBench.</p></li><li><p class="paragraph" style="text-align:left;"><b>Closed source</b>: Qwen3.7-Max is proprietary and will be available via Alibaba Cloud Model Studio API. Open-weight variants at smaller sizes are expected to follow.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/github/spec-kit?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>GitHub Spec Kit forces AI to spec before it codes</b></a></h3><div class="image"><a class="image__link" href="https://github.com/github/spec-kit?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2759a7a6-014d-4025-b01d-dfa266471b23/Screenshot_2026-05-19_at_10.00.47_PM.png?t=1779253251"/></a></div><p class="paragraph" style="text-align:left;">Still throwing vague prompts at your coding agent and hoping it doesn&#39;t torch your project?</p><p class="paragraph" style="text-align:left;"><b>GitHub</b> just open-sourced <b>Spec Kit</b>, a toolkit that makes the AI create a structured specification before it writes a single line of code. The agent figures out what you want, asks clarifying questions, plans the architecture, generates a task list, then implements. All structured, all inspectable, all before any code exists.</p><p class="paragraph" style="text-align:left;">103K stars already. Works with 30+ coding agents out of the box: Claude Code, Cursor, Codex, Gemini CLI, Junie, and more. And it&#39;s completely stack-agnostic, so it doesn&#39;t care if you&#39;re writing Rust or Rails.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Structured before creative: T</b>he agent can&#39;t start coding until it&#39;s written a spec, asked clarifying questions, and planned the architecture. The sequence is enforced, not optional.</p></li><li><p class="paragraph" style="text-align:left;"><b>30+ agent integrations: </b>Works with Claude Code, Copilot, Gemini, Codex, Cursor, and pretty much every coding agent you&#39;re already using. Same spec, any agent.</p></li><li><p class="paragraph" style="text-align:left;"><b>Extensible via presets and extensions: </b>Customize the workflow with your own templates. Runtime resolution follows a priority chain: project-local overrides beat presets beat extensions beat core.</p></li><li><p class="paragraph" style="text-align:left;"><b>MIT-licensed:</b> Install via <code>uv tool install</code> from the GitHub repo. Full greenfield, exploration, and brownfield workflows supported out of the box.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://safedep.io/mini-shai-hulud-strikes-again-314-npm-packages-compromised/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>314 npm Packages Compromised to Hijack Your Coding Agent</b></a></h3><div class="image"><a class="image__link" href="https://safedep.io/mini-shai-hulud-strikes-again-314-npm-packages-compromised/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/791065cf-9863-4d51-aee0-950b135bde1d/Screenshot_2026-05-19_at_10.01.59_PM.png?t=1779253324"/></a></div><p class="paragraph" style="text-align:left;">22 minutes. That&#39;s how long it took an attacker to publish 637 malicious versions across 317 npm packages with a combined 11+ million monthly downloads.</p><p class="paragraph" style="text-align:left;">The compromised account &quot;atool&quot; pushed a payload from the &quot;Mini Shai-Hulud&quot; toolkit, the same one behind the SAP compromise three weeks ago. But the interesting bit is that the malware specifically targets AI coding agents. It injects Claude Code SessionStart hooks, Codex hooks, and VS Code &quot;runOn: folderOpen&quot; tasks. It harvests AWS credentials, Kubernetes tokens, SSH keys, GitHub PATs, and even 1Password and Bitwarden vaults. Exfiltration is disguised as OpenTelemetry traces to blend in with your existing observability stack.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Your coding agent is a vector</b>: The payload hooks into Claude Code and Codex session startup, silently piping every credential it finds to a C2 server. If you&#39;re running these agents in environments with cloud access, this is as bad as it sounds.</p></li><li><p class="paragraph" style="text-align:left;"><b>Packages you probably use</b>: size-sensor (4.2M downloads/month), echarts-for-react (3.8M), @antv/scale (2.2M), timeago.js (1.15M). Check your lockfile.</p></li><li><p class="paragraph" style="text-align:left;"><b>Persistent and stealthy</b>: A LaunchAgent/systemd service called &quot;kitty-monitor&quot; survives reboots and uses GitHub commit search as a dead-drop C2 channel, polling for RSA-PSS signed commands.</p></li><li><p class="paragraph" style="text-align:left;"><b>Full advisory and IoCs available</b>: SafeDep published the complete list of all 317 compromised packages with deobfuscated payloads and remediation steps.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/karpathy/status/2056753169888334312?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Andrej Karpathy joins Anthropic</b></a><br>Yesterday, he announced he&#39;s joined Anthropic. &quot;The next few years at the frontier of LLMs will be especially formative,&quot; he wrote. He shaped the early GPT era at OpenAI, then left to build Eureka Labs for AI education. Now he&#39;s back in the lab at what might be the most interesting research org in the field right now. One to watch.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/Google/status/2056791134295273554?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Google releases a managed personal 24/7 agent</b></a><br>Gemini App now comes with Spark, an always-on personal AI agent, running on Gemini 3.5 and built on Antigravity. It navigates your digital life and takes actions on your behalf, even when you close your laptop. You can set up cron jobs (schedule tasks), teach it new Skills, and create end-to-end workflows. Rolling out to trusted testers now, with beta access for Google AI Ultra subscribers next week.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://elevenlabs.io/speech-engine?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Turn any chat agent into a voice agent with one prompt</b></a><br>ElevenLabs just shipped Speech Engine, and it’s pretty straightforward: keep your existing chat agent exactly as it is, add Speech Engine on top, and now it talks. You don’t need to rearchitect your LLM stack or swap out your RAG pipeline. It bundles speech-to-text, turn detection, interrupt handling, TTS, and audio orchestration into a single pipeline with ultra-low latency. Works with any LLM that produces text, has built-in stream extraction, and covers 70+ languages.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://deepmind.google/models/gemini-omni/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Google releases the Nano Banana of video-gen</b></a><br>Gemini Omni is Google&#39;s new any-input-to-any-output model. Feed it images, text, video, audio, or any combination, and it generates or edits video through conversation. Multi-turn editing keeps scenes consistent across back-and-forth iterations, and it applies real-world physics to generated content. Available in the Gemini app, Google Flow, and YouTube Shorts. </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/warpdotdev/status/2056772856835453395?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Multi-agent orchestration comes to Warp Oz</b></a><br>Warp just shipped multi-agent orchestration in Oz with support for Claude Code, Codex, and the Warp Agent. Use /orchestrate to delegate complex tasks across a team of agents running locally or in the cloud. If you&#39;re already in the Warp terminal, this makes it your control plane for parallel agent work.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/google-antigravity/antigravity-cli?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Google Antigravity with Gemini 3.5 Flash now in your Terminal</b></a><br>Google Antigravity just shipped a CLI written in Go, powered by Gemini 3.5 Flash, and built for async workflows where agents run tasks in the background and report back when done. It shares the same tool and app server as Antigravity 2.0, so anything you build on the platform also works inside Google Search, where Antigravity powers the new agentic coding features: custom generative UIs, dashboards, and &quot;mini apps&quot; spun up from natural language.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://blog.google/products-and-platforms/products/search/search-io-2026/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow"><b>Google Search gets its biggest overhaul in 25 years</b></a><br>The search box itself is being rebuilt: AI-powered, dynamically expanding, with multimodal inputs (text, images, files, videos, even Chrome tabs). New &quot;search agents&quot; will monitor the web 24/7 for specific criteria like apartment listings or sneaker drops, and agentic booking lets you complete local service bookings right from Search. Launching for Google AI Pro and Ultra subscribers this summer.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/nutlope/hallmark?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Hallmark</a></b>: Design skill by Hassan El Mghari that encodes anti-slop rules into Claude Code, Cursor, and Codex. Has four modes: build (generates pages that refuse to repeat the same structure twice), study (extracts a design&#39;s DNA from a URL or screenshot without copying pixels), audit (scores existing pages against its anti-pattern catalogue), and redesign (same content, deliberately different bones).</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/colbymchenry/codegraph?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">CodeGraph</a></b>: Pre-indexed code knowledge graph for Claude Code, Codex, Cursor, and OpenCode that cuts tool calls by 92% and speeds up tasks by 71%. 100% local, MIT-licensed.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/rohitg00/agentmemory?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">AgentMemory</a></b>: Persistent memory MCP server for coding agents with 4-tier consolidation inspired by how the brain organizes memory during sleep. Works with every major coding agent.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (111k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=stop-giving-agents-the-whole-computer"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=adc964d4-941d-40c2-b580-ceeca7b5c3a2&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Vercel Built a Programming Language for AI Agents</title>
  <description>+ Garry Tan’s agent brain, Codex on mobile, and a $1.3M agent bill</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d777c67d-73d9-4a71-ba0d-c9deb2cf5ad4/upload_2e618f7b3958f115.jpg" length="84288" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/vercel-built-a-programming-language-for-ai-agents</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/vercel-built-a-programming-language-for-ai-agents</guid>
  <pubDate>Mon, 18 May 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-05-18T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">I was looking at today’s updates and kept coming back to the Vercel one.</p><p class="paragraph" style="text-align:left;">A programming language built with AI agents in mind sounds a little ridiculous at first. Like, do agents really need their own language now?</p><p class="paragraph" style="text-align:left;">But then you look at the rest of today’s news: Garry Tan’s knowledge brain for his personal agents, Codex on mobile, and Peter Steinberger’s $1.3M monthly token spend. </p><p class="paragraph" style="text-align:left;">Suddenly, it feels less ridiculous. </p><p class="paragraph" style="text-align:left;">If agents are going to do real work, people are going to build weird new infrastructure around them. Some of it will be overkill. Some of it will probably become the new normal.</p><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Vercel built a programming language for AI agents</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Garry Tan open-sourced the brain running his AI agents</b></p></li><li><p class="paragraph" style="text-align:left;"><b>LiteLLM launches sandboxes for agent fleets</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Peter Steinberger spent $1.3M in OpenAI tokens in 30 days</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Codex crossed 4M weekly users and landed on mobile</b></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>The Ultimate Guide to /goal</b></a></p><p class="paragraph" style="text-align:left;">HTTP is a primitive. JSON is a primitive. /goal is becoming one for coding agents.</p><p class="paragraph" style="text-align:left;">A few weeks ago, OpenAI&#39;s Codex CLI added /goal as a way to give the coding worker a job with a defined done state. Claude Code added it this week.</p><p class="paragraph" style="text-align:left;">Hermes Agent, the orchestrator I run on a Mac Mini to coordinate work between coding workers, has had /goal built in for a while.</p><p class="paragraph" style="text-align:left;">This guide walks through what /goal actually is, the three roles in a multi-agent setup, a real end-to-end run, the verification rule, and how to run goals in parallel without workers stepping on each other. </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.theunwindai.com/p/the-ultimate-guide-to-goal?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>Read The Ultimate Guide to /goal</b></a></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/garrytan/gbrain?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>Garry Tan Open-Sources the AI Brain That Runs His Life</b></a></h3><div class="image"><a class="image__link" href="https://x.com/garrytan/status/2055670533451366479?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/71de464f-1f2c-488e-9511-a6cad27fcfe9/Screenshot_2026-05-17_at_8.42.38_PM.png?t=1779075763"/></a></div><p class="paragraph" style="text-align:left;">Garry Tan’s gstack made your AI agent ship code like a sprint team. Now, he open-sourced the knowledge system that powers his personal AI agents. </p><p class="paragraph" style="text-align:left;"><b>GBrain</b> is not a notes app or a RAG pipeline. It&#39;s a structured, self-maintaining knowledge brain that currently holds 17,000+ pages, tracks 4,000+ people, and runs 21 autonomous cron jobs. Garry built it in 12 days.</p><p class="paragraph" style="text-align:left;">The core idea is great: instead of re-deriving knowledge from scratch every query (like RAG does), GBrain pre-computes and maintains a &quot;compiled truth&quot; for every entity. Your agent gets richer context every time, and the whole thing compounds daily as it ingests meetings, emails, and calls. You wake up, and the brain is smarter than when you went to bed.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>MECE knowledge structure</b> - Everything is organized into clean categories: people, companies, deals, meetings, projects, concepts, originals, and media. Each page has a compiled summary on top and an append-only evidence trail below.</p></li><li><p class="paragraph" style="text-align:left;"><b>Self-wiring graph</b> - Every page-write automatically extracts entity references and creates typed links (attended, works_at, invested_in) with zero LLM calls. Ask &quot;who works at Acme AI?&quot; and get answers vector search alone can&#39;t reach.</p></li><li><p class="paragraph" style="text-align:left;"><b>34 built-in skills</b> - From signal detection to content ingestion to research synthesis. Intelligence lives in markdown skill files, not the runtime. This is how Garry&#39;s agents know how to do things, not just remember things.</p></li><li><p class="paragraph" style="text-align:left;"><b>Auto-enrichment tiers</b> - Mention someone once, they get a stub. Three mentions triggers web enrichment. Meet them in person or mention them 8+ times, and the full research pipeline kicks in. The brain decides how much attention someone deserves.</p></li><li><p class="paragraph" style="text-align:left;"><b>Usage</b> - Works as a standalone CLI, an MCP server for Claude Code and Cursor, or a one-click deploy on OpenClaw or Railway. Check it out at <a class="link" href="https://github.com/garrytan/gbrain?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">github.com/garrytan/gbrain</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/vercel-labs/zero?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>Vercel Built a Programming Language for AI Agents</b></a></h3><div class="image"><a class="image__link" href="https://github.com/vercel-labs/zero?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b73a5f20-ca0e-4d50-9da4-1974386c4459/image.png?t=1779075737"/></a></div><p class="paragraph" style="text-align:left;">Programming languages were designed for humans and then retrofitted for AI. Chris Tate from Vercel just changed that. </p><p class="paragraph" style="text-align:left;"><b>Zero</b> is a new systems language where AI agents are first-class users of the entire toolchain. Compiler errors come back as structured JSON with stable error codes and fix suggestions that agents can parse and act on programmatically.</p><p class="paragraph" style="text-align:left;">Think of it as a language where the compiler talks to your agent the same way a senior engineer talks to a junior one: here&#39;s what&#39;s wrong, here&#39;s the error code, here&#39;s exactly how to fix it. </p><p class="paragraph" style="text-align:left;">Still experimental, but the idea is genuinely novel. And the fact that it&#39;s coming from Vercel, not a random weekend project, means there&#39;s real conviction behind it.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Structured diagnostics</b> - Every compiler error returns JSON with stable codes, line locations, and repair metadata. No more regex-parsing error messages. Agents can read errors and fix code without scraping human-readable text.</p></li><li><p class="paragraph" style="text-align:left;"><b>Tiny output footprint</b> - Compiles down to extremely small native binaries with no garbage collector, event loop, and hidden runtime overhead. Think CLI tools and serverless functions.</p></li><li><p class="paragraph" style="text-align:left;"><b>Full CLI toolchain</b> - One command handles check, build, run, test, format, inspect, dependency graphs, and docs. Everything an agent needs in one place, all machine-readable output.</p></li><li><p class="paragraph" style="text-align:left;"><b>Human-readable too</b> - Despite being agent-first, the syntax is clean and readable. File extension is .0, which is a fun touch.</p></li><li><p class="paragraph" style="text-align:left;"><b>Try it now</b> - Zero is open-source (Apache 2.0) from Vercel. Check it out on <a class="link" href="http://github.com/vercel-labs/zero?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> and <a class="link" href="https://zerolang.ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">zerolang.ai</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/BerriAI/litellm-agent-platform?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>LiteLLM Launches an Agent Platform with K8s Sandboxes</b></a></h3><div class="image"><a class="image__link" href="https://github.com/BerriAI/litellm-agent-platform?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/eca3980c-80a6-4790-8d44-70b436e091e9/Screenshot_2026-05-17_at_8.46.29_PM.png?t=1779075995"/></a></div><p class="paragraph" style="text-align:left;">The team behind LiteLLM just shipped something bigger: a full platform for running fleets of coding agents in isolated Kubernetes sandboxes. </p><p class="paragraph" style="text-align:left;">Each agent session gets its own fresh pod. Your real API keys never touch agent code.</p><p class="paragraph" style="text-align:left;">A great feature is the credential vault. Agents only see stub tokens, and the platform transparently swaps in real secrets on every outbound connection. </p><p class="paragraph" style="text-align:left;">If you&#39;re scaling from one coding agent to a team of them, this makes it super safe.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Agent-agnostic sandboxes</b> - First-class support for Claude Code, Codex, and Hermes agents, all running in isolated pods that persist for 24 hours after you detach.</p></li><li><p class="paragraph" style="text-align:left;"><b>Credential vault</b> - Agents never see your real API keys. The vault injects real secrets transparently on outbound connections, so a rogue agent can&#39;t leak anything. This alone is worth the setup.</p></li><li><p class="paragraph" style="text-align:left;"><b>Terminal-first workflow</b> - The CLI lets you open a sandbox, attach your terminal, do your work, and Ctrl-D to detach. Simple and clean.</p></li><li><p class="paragraph" style="text-align:left;"><b>Full developer API</b> - Create agents, open sessions, send messages, and read replies, all via REST. Build your own orchestration on top.</p></li><li><p class="paragraph" style="text-align:left;"><b>Open source and self-hostable</b> - MIT licensed with local dev support and a production deploy path on AWS. Check it out at <a class="link" href="https://github.com/BerriAI/litellm-agent-platform?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">github.com/BerriAI/litellm-agent-platform</a>.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/steipete/status/2055346265869721905?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Peter Steinberger spent $1.3M in API tokens in 30 days</a></b><br>Peter Steinberger, creator of OpenClaw and now an OpenAI employee, casually revealed he burned through $1.3M worth of API tokens in a single month. That&#39;s roughly $20K per day, mostly on GPT-5.5 powering agents that manage the OpenClaw repo. Yes, he almost certainly has unlimited access as an OpenAI employee, so this isn&#39;t out of pocket. The number is wild, but the useful part is what it says about agent economics. Once agents run continuously, token spend becomes infrastructure, not just API usage</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://openai.com/index/work-with-codex-from-anywhere/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Codex hits 4M+ weekly users, now on Mobile</a></b><br>OpenAI&#39;s coding agent Codex just crossed 4 million weekly active users and is now available on iOS and Android through the ChatGPT app. You can kick off coding tasks, review diffs, and approve PRs from your phone. The &quot;desk-only&quot; constraint for AI-assisted coding is officially gone. 4M WAU also makes it one of the most widely adopted coding agents by a wide margin.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://interfaze.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">AI model built for deterministic dev tasks</a></b><br>Interfaze is a new model architecture that takes a fundamentally different approach. Instead of one autoregressive model doing everything, it breaks tasks into deterministic sub-steps like OCR, web search, classification, extraction, and then orchestrates purpose-built models for each. The result is structured output accuracy that beats GPT-5.4-Mini and matches Gemini-3-Flash. It&#39;s OpenAI SDK-compatible, comes with built-in web search from its own crawler, and runs tasks like audio transcription (1h 35m podcast in ~50 seconds) and document OCR natively. Free tier available at <a class="link" href="https://interfaze.ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">interfaze.ai</a>.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://openai.com/index/personal-finance-chatgpt/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">ChatGPT now wants to manage your personal finance</a></b><br>OpenAI launched a personal finance experience in ChatGPT for Pro users in the U.S. Connect your bank accounts via Plaid, and ChatGPT gives you a spending dashboard, subscription tracker, and financial guidance grounded in your transaction data. The before/after examples are genuinely compelling, going from generic &quot;save more money&quot; advice vs. &quot;cap dining at $450/month based on your Feb-May spend.&quot; OpenAI has also partnered with Intuit that signals the next step: actionable finance, like applying for credit cards and scheduling tax appointments directly from chat.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://clawpatch.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Clawpatch</a></b>: Code review for agent-written code. It maps a repo into semantic slices like routes, commands, packages, and tests, then reviews bounded contexts instead of isolated files. </p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://github.com/MinishLab/semble?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Semble</a></b>: Code search built specifically for AI coding agents. Instead of dumping entire files into context, Semble indexes your repo and returns only the relevant snippets, using 98% fewer tokens than grep-and-read. Drop-in MCP server for Claude Code, Codex, and Cursor.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="http://github.com/DrCatHicks/learning-opportunities?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow"><b>Learning Opportunities</b></a>: A Claude Code and Codex plugin that pauses after significant coding work and offers short, evidence-based learning exercises. The idea: if an AI agent writes your code, you should still understand what it did and why. Turns AI-assisted coding into actual skill building.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> (111k+ </b>🌟 <b>) </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=vercel-built-a-programming-language-for-ai-agents"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=205d901f-4696-44b8-a2ea-a9c805534e59&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The Ultimate Guide to /goal</title>
  <description>/goal isn’t a feature. It’s the new primitive.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/12a28707-3d69-4089-81b8-21868e66fa0a/upload_3846ff9a09589062.jpg" length="83248" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/the-ultimate-guide-to-goal</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/the-ultimate-guide-to-goal</guid>
  <pubDate>Mon, 18 May 2026 04:32:06 +0000</pubDate>
  <atom:published>2026-05-18T04:32:06Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b>/goal is not a feature. It is a primitive. </b></p><p class="paragraph" style="text-align:left;">HTTP is a primitive. JSON is a primitive. /goal is becoming one for coding agents.</p><p class="paragraph" style="text-align:left;">A few weeks ago, OpenAI&#39;s Codex CLI added /goal as a way to give the coding worker a job with a defined done state. Claude Code added it this week. </p><p class="paragraph" style="text-align:left;">Hermes Agent, the orchestrator I run on a Mac Mini to coordinate work between coding workers, has had /goal built in for a while. </p><p class="paragraph" style="text-align:left;">So I now have a builder, a reviewer, and an orchestrator that all accept the same instruction format, even though they share nothing else.</p><p class="paragraph" style="text-align:left;">If you&#39;ve only seen /goal used as a fancier prompt, you&#39;ve missed what it changes.</p><h2 class="heading" style="text-align:left;"><b>What /goal actually is</b></h2><p class="paragraph" style="text-align:left;">A regular prompt asks an agent for the next response. You read what comes back, decide if it&#39;s right, and push the agent forward to the next step. You&#39;re steering every turn.</p><p class="paragraph" style="text-align:left;">/goal flips that. You write down what &quot;done&quot; looks like, submit it once, and the agent works toward it until it gets there. Here&#39;s a real one:</p><div class="codeblock"><pre><code>/goal Build the app described in SPEC.md. Done means tests pass, build passes, README is accurate, and git status only shows relevant project files.</code></pre></div><p class="paragraph" style="text-align:left;">The goal stays active until it&#39;s achieved, paused, blocked, cleared, or it runs out of budget.</p><p class="paragraph" style="text-align:left;">This is different from putting the word &quot;goal&quot; inside a normal one-shot command. If you write codex exec &#39;goal: build the app&#39;, that&#39;s still a prompt with a label. The real primitive lives inside an interactive worker session. You launch the CLI, you submit /goal, and you walk away.</p><p class="paragraph" style="text-align:left;">The shift is from prompting (you driving) to assigning (the agent driving toward a target you defined).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/02460539-5ebc-46fd-abb3-6471754457b2/diag1.gif?t=1779078302"/></div><h2 class="heading" style="text-align:left;"><b>The three tools that currently speak /goal</b></h2><p class="paragraph" style="text-align:left;">The three tools accepting /goal aren&#39;t all the same kind of thing, so it&#39;s worth being specific.</p><p class="paragraph" style="text-align:left;"><b>Codex</b> is OpenAI&#39;s coding CLI. Strong at implementation, especially when given a clear spec. /goal is how you give it that spec.</p><p class="paragraph" style="text-align:left;"><b>Claude Code</b> is Anthropic&#39;s coding CLI. Strong at the inverse: finding what&#39;s wrong with code that looks right. Spec compliance, safety issues, error states, security holes. /goal is how you point it at code and ask for a review.</p><p class="paragraph" style="text-align:left;"><b>Hermes Agent</b> is a different kind of tool entirely. Not a coding worker, but an orchestrator that coordinates work between coding workers like the two above. /goal is how Hermes hands off tasks to whichever tool is right for the job, and also how I tell Hermes what I want in the first place.</p><p class="paragraph" style="text-align:left;">What matters isn&#39;t that any one of them shipped /goal. It&#39;s that three different teams converged on the same primitive, and that convergence is what makes it possible to compose them.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/444f1a13-913d-464d-ac28-4ec472960a59/diag2.gif?t=1779078344"/></div><h2 class="heading" style="text-align:left;"><b>Setting things up</b></h2><p class="paragraph" style="text-align:left;">The first time I needed Codex and Claude Code on the Mac Mini that runs Hermes, I didn&#39;t install them by hand. I sent Hermes a message asking it to install both and log me in. It handled the rest.</p><p class="paragraph" style="text-align:left;">That&#39;s the workflow now. You don&#39;t type install commands. Setup is just another goal.</p><p class="paragraph" style="text-align:left;">If you don&#39;t have an orchestrator running yet, the install pages for Codex and Claude Code are easy enough to follow. But once you do, you shouldn&#39;t set up another tool by hand. The point of having an orchestrator is that mechanical work stops being yours.</p><h2 class="heading" style="text-align:left;"><b>What Hermes adds on top of /goal</b></h2><p class="paragraph" style="text-align:left;">A raw /goal is useful on its own. But it leaves you with a coordination problem.</p><p class="paragraph" style="text-align:left;">If Codex is running in one terminal and Claude Code is running in another, you have to remember which process is doing what. You have to check logs. You have to manually pass review findings from one tool to the other.</p><p class="paragraph" style="text-align:left;">Hermes turns those loose runs into a workflow:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">You message Hermes (in my case, over Telegram from my phone)</p></li><li><p class="paragraph" style="text-align:left;">Hermes creates goal cards on a Kanban board</p></li><li><p class="paragraph" style="text-align:left;">Hermes picks the right worker for each card</p></li><li><p class="paragraph" style="text-align:left;">The worker runs the goal in the background</p></li><li><p class="paragraph" style="text-align:left;">The card stores the process id, PID, repo, and done criteria</p></li><li><p class="paragraph" style="text-align:left;">When the build is ready, Hermes hands the repo to the reviewer</p></li><li><p class="paragraph" style="text-align:left;">If the review blocks, Hermes sends the findings back as a fix goal</p></li><li><p class="paragraph" style="text-align:left;">Hermes verifies the final output by inspecting the filesystem, tests, build, and git state</p></li></ol><p class="paragraph" style="text-align:left;">The board is what /goal becomes when there&#39;s an orchestrator on top of it. Every goal has a card, every card has a status, every handoff leaves a trail. Instead of hunting through terminals, you watch the work move across columns on your phone.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9012b167-870b-41b0-a16d-556c1b77d7fc/image.png?t=1779078169"/></div><h2 class="heading" style="text-align:left;"><b>The three roles</b></h2><p class="paragraph" style="text-align:left;">The tools change. The roles don&#39;t.</p><p class="paragraph" style="text-align:left;"><b>Orchestrator.</b> Owns the control loop. Task decomposition, worker selection, Kanban cards, background processes, dependencies, final verification, the user-facing summary. In my setup, Hermes.</p><p class="paragraph" style="text-align:left;"><b>Builder.</b> Takes a spec and produces working code. Implementation is the bottleneck this role solves. Codex tends to be strong here.</p><p class="paragraph" style="text-align:left;"><b>Reviewer.</b> Reads what the builder produced and finds what&#39;s wrong with it. Correctness is the bottleneck. Claude Code tends to be strong here.</p><h2 class="heading" style="text-align:left;"><b>A real run, end to end</b></h2><p class="paragraph" style="text-align:left;">I gave Hermes agent a goal to do this:</p><div class="codeblock"><pre><code>/goal Build a CLI tool that finds X mentions of me and pings me when something blows up.</code></pre></div><p class="paragraph" style="text-align:left;">Hermes broke the request into six cards.</p><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/Saboo_Shubham_/status/2054260705365475609?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal"><p> Twitter tweet </p></a></blockquote><p class="paragraph" style="text-align:left;"><b>Card 1: Spec.</b> Hermes wrote SPEC.md itself, capturing the stack, repo path, read-only constraints, mock mode requirements, tests, and verification commands. Owned by the PM role.</p><p class="paragraph" style="text-align:left;"><b>Card 2: Codex builds.</b> Codex ran /goal against SPEC.md. It created the project files, implemented the UI and backend, added tests, and got the app to a passing state. About 15 minutes. When it finished, npm test passed, npm run build passed, and git status showed only relevant new files.</p><p class="paragraph" style="text-align:left;"><b>Card 3: Claude Code reviews.</b> Claude Code ran /goal to review what Codex built. Checked spec compliance, read-only safety, API key handling, error states, tests, UI usefulness, bugs, and security issues. Result: PASS, no blocking issues.</p><p class="paragraph" style="text-align:left;"><b>Card 4: Codex fix loop.</b> Skipped, because the review passed. The card still matters when skipped. It shows Hermes can model conditional work. If Claude Code had blocked, Hermes would have handed the findings back to Codex as a new /goal.</p><p class="paragraph" style="text-align:left;"><b>Card 5: Claude Code final verification.</b> Skipped for the same reason.</p><p class="paragraph" style="text-align:left;"><b>Card 6: Hermes final summary.</b> Working app at the local path, UI and API both verified in mock mode. Codex built it with /goal. Claude Code reviewed it with /goal and returned PASS.</p><p class="paragraph" style="text-align:left;">All of that came from one message. Three different tools did the actual work, but I only ever talked to Hermes.</p><h2 class="heading" style="text-align:left;"><b>The verification rule</b></h2><p class="paragraph" style="text-align:left;">Hermes never trusted Codex&#39;s self-report. After Codex marked the build done, Hermes ran the commands itself:</p><div class="codeblock"><pre><code>npm test         # 17 tests passed
npm run build    # vite build passed</code></pre></div><p class="paragraph" style="text-align:left;">The verifier is what makes a /goal a contract instead of a promise. Don&#39;t trust the worker&#39;s self-report as final. Trust the verifier.</p><p class="paragraph" style="text-align:left;">Coding agents are confident. They&#39;ll tell you the build passes when the build was never run. They&#39;ll tell you tests pass when they wrote tests that never executed. The verifier closes that gap.</p><p class="paragraph" style="text-align:left;">Without verification, /goal is just a fancier prompt. With verification, it becomes a contract.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ae40a901-ecea-40df-acad-388ddca732fb/diagram3.gif?t=1779078360"/></div><h2 class="heading" style="text-align:left;"><b>Running multiple goals</b></h2><p class="paragraph" style="text-align:left;">You can run multiple /goals in parallel, but you can&#39;t point multiple coding workers at the same files without thinking about it first.</p><p class="paragraph" style="text-align:left;">My default is one main builder per repo. If I want parallelism, I add it across clear boundaries. Different repos, different branches, git worktrees, separate packages, docs vs code, tests vs implementation. Anywhere two workers can&#39;t step on each other.</p><p class="paragraph" style="text-align:left;">The bad pattern is three workers all editing the same file in the same repo. You get conflicts, partial overwrites, and one worker silently undoing another&#39;s work.</p><p class="paragraph" style="text-align:left;">The better pattern is one writer at a time on any given file. Builder writes, reviewer only reads, fix goals stay scoped to the fix. Or run three builders in three worktrees on three competing approaches and let the orchestrator pick the best one.</p><p class="paragraph" style="text-align:left;">The board is what makes this practical. Without it, parallel background workers become terminal chaos.</p><h2 class="heading" style="text-align:left;"><b>What changes for me</b></h2><p class="paragraph" style="text-align:left;">The useful framing here is not &quot;I can run agents in the background.&quot;</p><p class="paragraph" style="text-align:left;">It&#39;s that one message turns into a pipeline across three different coding tools, and I watch the whole thing move across one board.</p><p class="paragraph" style="text-align:left;">You stop sitting in a terminal waiting for one agent to finish, and start managing a queue of work with visible state.</p><p class="paragraph" style="text-align:left;">If Codex and Claude Code had each invented their own job-handoff format, no orchestrator could route between them. The board is impressive, but the primitive makes the board even more useful. </p><p class="paragraph" style="text-align:left;">The workers can change, but the primitive stays the same. The next coding tool that adopts /goal will join this pipeline without me changing anything. I&#39;ll just route work to it.</p><p class="paragraph" style="text-align:left;">That&#39;s what good primitives do.</p><p class="paragraph" style="text-align:left;">For more such cool tips and interesting ideas around Hermes, OpenClaw, Claude Code, Codex and other 24/7 agent teams.</p><p class="paragraph" style="text-align:left;"><b>Follow → </b><b><a class="link" href="https://x.com/@Saboo_Shubham_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal" target="_blank" rel="noopener noreferrer nofollow">@Saboo_Shubham_</a></b></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">We share in-depth blogs and tutorials like this 2-3 times a week, to help you stay ahead in the world of AI. <span style="text-decoration:underline;"><b><a class="link" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">If you&#39;re serious about leveling up your AI skills and staying ahead of the curve, subscribe now and be the first to access our latest tutorials.</a></b></span></p><p class="paragraph" style="text-align:left;"><b>Don’t forget to share this tutorial on your social channels and tag Unwind AI (</b><span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span><b>) to support us!</b></p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-to-goal"><span class="button__text" style=""> Subscribe now for FREE - Get instant access to more LLM, RAG & AI Agent tutorials </span></a></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=b28a5198-384f-4d50-8417-63c04b15ad27&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>/goal in Claude Code, Codex, and Hermes Agent</title>
  <description>+ OpenClaw creator open-sourced a Mac automation tool</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/96686929-fda3-444d-b10a-f71a87e62dc8/upload_b4334d245ad1ac8d.jpg" length="96605" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/goal-in-claude-code-codex-and-hermes-agent</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/goal-in-claude-code-codex-and-hermes-agent</guid>
  <pubDate>Tue, 12 May 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-05-12T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/openclaw/Peekaboo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>OpenClaw creator open-sourced a native Mac automation tool</b></a></p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/antirez/ds4?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">Run DeepSeek V4 Flash locally on a 128GB Mac</a></b></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/daybreak/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>OpenAI enters serious cyber defense with Daybreak</b></a></p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/Saboo_Shubham_/status/2054017280376455265?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">/goal in Codex CLI, Hermes Agent, and Claude Code</a></b></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/GTG-Labs/sangria?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Let agents pay for your API in ~3 lines of code</b></a></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.theunwindai.com/p/build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Build a Multimodal Agentic RAG App with Gemini Embedding 2 and Google ADK</b></a></p><p class="paragraph" style="text-align:left;">In this tutorial, you&#39;ll build a fully-working <b>multimodal agentic RAG app</b> where text, URLs, PDFs, images, audio, and video all share a single 768-dimension embedding space, and a small Google Agent Development Kit (ADK) coordinator turns the retrieved evidence into a grounded, cited answer.</p><p class="paragraph" style="text-align:left;">The two pieces doing the heavy lifting are <b>Gemini Embedding 2</b>, which embeds every modality into the same vector space, and <b>Google ADK</b>, which wraps the retrieval call in an agent that inspects the workspace, calls the retrieval tool, and writes the answer.</p><p class="paragraph" style="text-align:left;">You&#39;ll see exactly how those two pieces compose without any extra orchestration framework.</p><div class="embed"><a class="embed__url" href="https://www.theunwindai.com/p/build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank"><img class="embed__image embed__image--left" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/c9ccb7a4-e58d-4b21-bb2e-c0358b7a3e7e/Build_a_Multimodal_Agentic_RAG_App_with_Gemini_Embedding_2_and_Google_ADK.png?t=1778367901"/><div class="embed__content"><p class="embed__title"> Build a Multimodal Agentic RAG App with Gemini Embedding 2 and Google ADK </p><p class="embed__description"> (100% open source) </p></div></a></div><p class="paragraph" style="text-align:left;">We share hands-on tutorials like this every week, designed to help you stay ahead in the world of AI. <span style="text-decoration:underline;"><b><a class="link" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">If you&#39;re serious about leveling up your AI skills and staying ahead of the curve, subscribe now and be the first to access our latest tutorials.</a></b></span></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://github.com/openclaw/Peekaboo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">OpenClaw creator open-sourced a native Mac automation tool</a></b></h3><div class="image"><a class="image__link" href="https://github.com/openclaw/Peekaboo?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/61f25a9d-9e77-422a-ac35-11d42e471826/Screenshot_2026-05-11_at_10.29.37_PM.png?t=1778563783"/></a></div><p class="paragraph" style="text-align:left;">Still using computer use agents doing screenshot → click → reason → hallucinate → repeat?</p><p class="paragraph" style="text-align:left;">Peter Steinberger open-sourced <b>Peekaboo</b>, a MacOS-native toolkit that hands agents the accessibility tree directly, so clicks land on real elements with IDs, not pixel coordinates that drift every time the window moves. </p><p class="paragraph" style="text-align:left;">Beyond click and type, it covers the stuff every other agent fails at, like Spaces switching, Dock right-clicks, menu bar extras, file dialogs, drag-to-Trash.</p><p class="paragraph" style="text-align:left;">Use it with Claude Code, Codex, OpenClaw, Hermes Agent, or which agent harness you like, via CLI and MCP server. Bring any model like Claude, GPT-5.1, Grok 4-fast, or Ollama for fully local runs. MIT-licensed.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Native, not virtualized</b>: Runs as a real macOS process with Screen Recording + Accessibility permissions. It can drive any app you can, including ones that block automation inside browsers or VMs.</p></li><li><p class="paragraph" style="text-align:left;"><b>Structured menu discovery</b>: peekaboo menu returns the full menu tree as JSON, so agents navigate &quot;File → Export → PDF…&quot; by name instead of pattern-matching on screenshots.</p></li><li><p class="paragraph" style="text-align:left;"><b>Multi-screen and multi-Space aware</b>: First-class support for moving windows between Spaces, switching desktops, and targeting elements on specific displays.</p></li><li><p class="paragraph" style="text-align:left;"><b>Drop into OpenClaw</b>: Lives in the repo as skills/peekaboo-cli, so you can install as an OpenClaw skill alongside the 5,400+ others.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://console.mistral.ai/build/audio/text-to-speech/?utm_source=unwindai&utm_medium=newsletter&utm_campaign=audio" target="_blank" rel="noopener noreferrer nofollow">Voxtral TTS: Outperforms ElevenLabs on naturalness</a></b></h3><div class="image"><a class="image__link" href="https://console.mistral.ai/build/audio/text-to-speech/?utm_source=unwindai&utm_medium=newsletter&utm_campaign=audio" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bbfd6994-a831-402b-9b4b-5ec87a49f5bb/image.png?t=1778564785"/></a></div><p class="paragraph" style="text-align:left;">When it comes to voice agents, naturalness is a key factor. <a class="link" href="https://console.mistral.ai/build/audio/text-to-speech/?utm_source=unwindai&utm_medium=newsletter&utm_campaign=audio" target="_blank" rel="noopener noreferrer nofollow">Voxtral TTS</a> outperforms ElevenLabs Flash v2.5 on naturalness and matches ElevenLabs v3 quality with emotion-steering support. Lightweight at 4B parameters, built for production.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Wins 58.3% of flagship voice preference tests</b>: In side-by-side human evaluations against ElevenLabs Flash v2.5, Voxtral TTS wins on naturalness across flagship voices and 68.4% on voice customization.</p></li><li><p class="paragraph" style="text-align:left;"><b>Emotion-aware output</b>: Contextual understanding (neutral, happy, sarcastic, and more) determines whether output sounds considered or robotic.</p></li><li><p class="paragraph" style="text-align:left;"><b>70ms model latency, ~9.7x real-time factor</b>: Streams natively and integrates into any existing STT and LLM stack.</p></li><li><p class="paragraph" style="text-align:left;"><b>Voice cloning from 3 seconds of audio</b>: Adapts to tone, personality, rhythm, and intonation. Zero-shot, no fine-tuning required.</p></li><li><p class="paragraph" style="text-align:left;"><b>Open weights under CC BY NC 4.0</b>: Deploy on your own infrastructure, extend to your own voice library.</p></li></ol><p class="paragraph" style="text-align:left;"><a class="link" href="https://console.mistral.ai/build/audio/text-to-speech/?utm_source=unwindai&utm_medium=newsletter&utm_campaign=audio" target="_blank" rel="noopener noreferrer nofollow">Try it now!</a></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://github.com/antirez/ds4?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Run DeepSeek V4 Flash locally on a 128GB Mac</b></a></h3><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d8fb5cfa-8ad5-4a8e-b1d6-ab699c9c1a9d/image.png?t=1778564592"/></div><p class="paragraph" style="text-align:left;">The creator of Redis just built a dedicated inference engine for a quasi-frontier model. It might be the most opinionated piece of AI infrastructure released this year.</p><p class="paragraph" style="text-align:left;">Salvatore Sanfilippo (antirez) released <b>DwarfStar4</b>, a purpose-built C + Metal engine that runs DeepSeek V4 Flash, a 284B parameter open-source model with a 1M token context window, locally on a 128GB MacBook at ~27 tokens/second.</p><p class="paragraph" style="text-align:left;">No generic runtime or framework. Just raw C + Metal doing one thing really well. The engine uses an asymmetric 2-bit quantization that fits the entire model in ~81GB, ships with a disk-based KV cache that&#39;s a lifesaver for agent workflows, and built-in APIs that plug straight into agents like OpenClaw, Hermes, Claude Code, Opencode, and Pi. </p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>One Model, Maximum Optimization</b>: DS4 isn&#39;t a generic model runner. It&#39;s a dedicated engine built exclusively for DeepSeek V4 Flash, squeezing out performance that general-purpose tools can&#39;t match.</p></li><li><p class="paragraph" style="text-align:left;"><b>Runs on a MacBook</b>: A specialized 2-bit quantization compresses the 284B model to ~81GB while keeping code generation and tool calling quality intact, making it genuinely usable on 128GB Macs.</p></li><li><p class="paragraph" style="text-align:left;"><b>Disk KV Cache for Agents</b>: Saves session state to your SSD so agent clients that resend large system prompts every request can skip the expensive prefill after the first run. That’s a massive time saver.</p></li><li><p class="paragraph" style="text-align:left;"><b>Agent-Ready Out of the Box</b>: Ships with OpenAI and Anthropic-compatible server APIs plus ready-to-use configs for Claude Code, opencode, and Pi.</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://openai.com/daybreak/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">OpenAI enters serious cyber defense with Daybreak</a></b><br>OpenAI just shipped Daybreak, a cyber defense stack built on GPT-5.5 and Codex Security.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3550ffc9-abbb-4459-9efe-9ed5f576e4ae/image.png?t=1778565362"/><div class="image__source"><span class="image__source_text"><p>iykyk</p></span></div></div><p class="paragraph" style="text-align:left;">The idea is Codex ingests your repo, builds a threat model specific to your codebase, then maps attack paths and validates real vulnerabilities in sandboxed environments. It generates patches, runs them, and sends audit-ready evidence back into your existing security stack. Here’s how the access will work for now: standard GPT-5.5 stays general-purpose, Trusted Access unlocks for verified defenders doing vuln triage and malware analysis, and GPT-5.5-Cyber is for authorized red teaming, pen testing, and controlled validation.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://thinkingmachines.ai/blog/interaction-models/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Thinking Machines show what they’re building with $2B funding</b></a><br>Thinking Machines finally demoed what they’re working on: &quot;interaction models.&quot; At first glance, it feels a lot like the GPT-4o demo from 2 years ago: real-time, audio-video-text. The interesting part is underneath though: a 276B MoE “interaction model” (12B active, 0.40s latency) that handles the live conversation, and a separate background model runs reasoning, searches, and tool calls mid-chat, then feeds results back in. Full-duplex isn&#39;t new (hi Moshi from Kyutai Labs), but the architectural design is interesting, and the early benchmarks on latency and quality are solid. </p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://claude.com/blog/new-in-claude-managed-agents?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Claude Agents can now dream between sessions</b></a><br>Anthropic&#39;s Claude Managed Agents (their hosted agent runtime, launched last month) got a solid update: dreaming, outcomes, and multi-agent orchestration. Dreaming is the one worth paying attention to! It reviews past agent sessions between runs, surfaces recurring mistakes and workflow patterns, and folds them back into memory automatically. Outcomes lets you define a success rubric evaluated by a separate grader in its own context window, looping the agent back until output clears the bar. Multi-agent orchestration does what you&#39;d expect — lead agent delegates to specialist subagents, each with their own model and tools, running in parallel.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/NousResearch/status/2052140057222369541?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>What the Hermes Agent community is actually building</b></a><br>If you’re still wondering what people are using Hermes Agent for, here’s a wall of 200+ of them to inspire you! The team used Hermes Agent itself to scrape the entire <span style="color:#0f1419;font-family:TwitterChirp, -apple-system, &quot;system-ui&quot;, &quot;Segoe UI&quot;, Roboto, Helvetica, Arial, sans-serif;font-size:17px;">internet for these usecases and added them to their docs. And if you&#39;ve found an interesting use case, you can submit your own!</span></p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/Saboo_Shubham_/status/2054017280376455265?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>/goal in Codex CLI, Hermes Agent, and Claude Code</b></a><br>ICYMI and have been manually re-prompting your coding agents with &quot;keep going,&quot; that&#39;s now a solved problem across the board. <code>/goal</code> gives the agent a durable objective with a clear done-condition. It keeps looping - planning, editing, running, and verifying until that condition is actually met or you tell it to stop. Codex CLI shipped it first, Hermes Agent picked it up in v0.13.0, and Claude Code now has its own native version. And here’s an interesting workflow we discovered: use Hermes Agent as an orchestration layer to fire <code>/goal</code> across Codex CLI and Claude Code simultaneously, and track all the running objectives on Hermes&#39;s Kanban board.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/GTG-Labs/sangria?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>Sangria</b></a>: An open-source SDK that lets you put a paywall on any API endpoint so AI agents can pay per request automatically via the x402 protocol and USDC on Base. Drops into Express/Fastify/Hono/FastAPI with minimal code.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://stack.cardor.dev/ahk?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>agent-harness-kit</b></a>: An open-source TypeScript CLI that scaffolds a structured multi-agent workflow into any codebase. One command sets up multiple agents (lead, explorer, builder, reviewer), a SQLite task backlog, and a health gate that runs before any agent can start or close work. It&#39;s provider-agnostic (Claude Code, OpenCode) and ships its own local MCP server.</p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/cursor_ai/status/2052432778743210127?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow"><b>/orchestrate</b></a>: Skill that decomposes large tasks into a tree of parallel cloud agents: planners, workers, and verifiers. It runs on the Cursor SDK&#39;s cloud runtime, so each agent gets an isolated VM, and the whole tree reconciles back through git and structured handoffs.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=goal-in-claude-code-codex-and-hermes-agent"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=b5fbfc2f-6b96-495d-800f-f93534024df9&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Build a Multimodal Agentic RAG App with Gemini Embedding 2 and Google ADK</title>
  <description>(100% open source)</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fd8b3191-62e2-4e37-97ea-e5e6c7f78023/upload_732c5a6f6dccd443.jpg" length="94008" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk</guid>
  <pubDate>Sat, 09 May 2026 23:07:20 +0000</pubDate>
  <atom:published>2026-05-09T23:07:20Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">If you have built a RAG app before, you know how quickly the &quot;just retrieve the right chunk&quot; problem fragments the moment your sources stop being plain text. Product PDFs, UI screenshots, recorded calls, demo videos, and support notes all carry the answer your user is asking for, but each lives in its own embedding silo. </p><p class="paragraph" style="text-align:left;">Stitching them together usually means three pipelines, two vector stores, and a glue layer you regret pretty soon.</p><p class="paragraph" style="text-align:left;">In this tutorial, you&#39;ll build a fully-working <b>multimodal agentic RAG app</b> where text, URLs, PDFs, images, audio, and video all share a single 768-dimension embedding space, and a small Google Agent Development Kit (ADK) coordinator turns the retrieved evidence into a grounded, cited answer. </p><p class="paragraph" style="text-align:left;">The two pieces doing the heavy lifting are <b>Gemini Embedding 2</b>, which embeds every modality into the same vector space, and <b>Google ADK</b>, which wraps the retrieval call in an agent that inspects the workspace, calls the retrieval tool, and writes the answer. </p><p class="paragraph" style="text-align:left;">You&#39;ll see exactly how those two pieces compose without any extra orchestration framework.</p><p class="paragraph" style="text-align:left;"><b>Don’t forget to share this tutorial on your social channels and tag Unwind AI (</b><span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.facebook.com/profile.php?id=61561355694033&utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Facebook</a></b></span><b>) to support us!</b></p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk"><span class="button__text" style=""> Subscribe now for FREE - Get instant access to more LLM, RAG & AI Agent tutorials </span></a></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>What We’re Building</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">A multimodal Agentic RAG demo where you can drop in any file or URL and ask questions across the whole index. The same retrieval packet powers both the answer and the citation panel, so the UI never disagrees with the model.</p><p class="paragraph" style="text-align:left;"><b>Key features:</b></p><ul><li><p class="paragraph" style="text-align:left;"><b>Truly multimodal index</b> — text, URLs, PDFs, images, audio, and video all live in one cosine-similarity space.</p></li><li><p class="paragraph" style="text-align:left;"><b>Gemini Embedding 2 with task prefixes</b> — separate prefixes for documents and queries to improve retrieval quality.</p></li><li><p class="paragraph" style="text-align:left;"><b>Google ADK agent</b> — coordinates <code>inspect_embedding_space</code> and <code>retrieve_relevant_context</code> tools, then synthesizes a grounded answer.</p></li><li><p class="paragraph" style="text-align:left;"><b>Single retrieval, two consumers</b> — <code>/ask</code> retrieves once, then passes the same packet to the agent and to the UI.</p></li><li><p class="paragraph" style="text-align:left;"><b>3D PCA embedding view</b> — every source is one point; ask a question and the query and cited sources light up in the same projection.</p></li><li><p class="paragraph" style="text-align:left;"><b>SSRF-safe URL ingestion</b> — private and loopback IPs blocked unless you opt in.</p></li></ul></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>How It Works</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">End-to-end, one question flows like this:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>You add sources.</b> Each source is chunked (text/URL) or uploaded once (PDF/image/audio/video). Every chunk gets a Gemini Embedding 2 vector with the <code>task: retrieval document</code> prefix. Files get a media vector blended with a text annotation vector, so titles still help retrieval.</p></li><li><p class="paragraph" style="text-align:left;"><b>You ask a question.</b> <code>/ask</code> embeds the query with the <code>task: question answering | query</code> prefix, scores every chunk by cosine similarity, keeps the best chunk per source, takes the top <i>k</i>, and projects everything into 3D using power-iteration PCA.</p></li><li><p class="paragraph" style="text-align:left;"><b>The agent runs.</b> <code>_run_adk_agent</code> builds a per-request agent whose <code>retrieve_relevant_context</code> tool is a closure over the already-computed retrieval packet. The agent calls <code>inspect_embedding_space</code>, then &quot;calls&quot; the retrieval tool, then writes a grounded answer with no inline citation IDs.</p></li><li><p class="paragraph" style="text-align:left;"><b>The UI renders.</b> The frontend shows the answer text, the citation panel (built from the same <code>matches</code>), the agent trace, and the updated 3D view with the query point and highlighted sources.</p></li></ol><p class="paragraph" style="text-align:left;">The architectural insight is the <b>single-retrieval contract</b>: one query embedding, one ranked list of matches, two consumers (the agent and the UI). That&#39;s what keeps citations honest.</p></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Prerequisites</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Before we begin, make sure you have the following:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Python installed on your machine (version 3.12 is recommended)</p></li><li><p class="paragraph" style="text-align:left;">Your <a class="link" href="https://aistudio.google.com/api-keys?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow">Gemini API key</a> for using Gemini Embedding 2</p></li><li><p class="paragraph" style="text-align:left;">A code editor of your choice</p></li><li><p class="paragraph" style="text-align:left;">Basic Python and FastAPI familiarity</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Code Walkthrough</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h4 class="heading" style="text-align:left;"><b>Setting Up the Environment</b></h4><p class="paragraph" style="text-align:left;">First, let&#39;s get our development environment ready:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Clone the GitHub repository:</p></li></ol><div class="codeblock"><pre><code>git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git</code></pre></div><h6 class="heading" style="text-align:left;">🌟<b> </b><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow">Don&#39;t forget to star the opensource repo to show your support.</a></b></h6><ol start="2"><li><p class="paragraph" style="text-align:left;">Go to the <a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/tree/main/rag_tutorials/multimodal_agentic_rag?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow"><b>multimodal_agentic_rag</b></a><b> </b>folder:</p></li></ol><div class="codeblock"><pre><code>cd rag_tutorials/multimodal_agentic_rag/backend</code></pre></div><ol start="3"><li><p class="paragraph" style="text-align:left;">Install the <a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/voice_ai_agents/insurance_claim_live_agent_team/requirements.txt?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow">required dependencies</a>:</p></li></ol><div class="codeblock"><pre><code>pip install -r requirements.txt</code></pre></div><ol start="4"><li><p class="paragraph" style="text-align:left;">Grab your<a class="link" href="https://aistudio.google.com/api-keys?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow"> Gemini API key from Google AI Studio</a> and set it in your current session:</p></li></ol><div class="codeblock"><pre><code>export GOOGLE_API_KEY=&quot;your-google-ai-studio-key&quot;</code></pre></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h4 class="heading" style="text-align:left;"><b>Creating the App</b></h4><p class="paragraph" style="text-align:left;">Project structure:</p><div class="codeblock"><pre><code>rag_tutorials/multimodal_agentic_rag/
|-- README.md
|-- assets/
|   `-- multimodal-agentic-rag-architecture.png
|-- backend/
|   |-- app_state.py
|   |-- rag_store.py
|   |-- requirements.txt
|   |-- server.py
|   `-- agentic_rag_agent/
|       |-- __init__.py
|       `-- agent.py
`-- frontend/
    |-- index.html
    |-- package.json
    |-- src/
    |   |-- App.tsx
    |   |-- main.tsx
    |   `-- styles.css
    |-- tsconfig.json
    `-- vite.config.ts</code></pre></div><p class="paragraph" style="text-align:left;">We’ll skip the frontend code walkthrough and focus on the backend architecture.</p><h3 class="heading" style="text-align:left;"><b>1. The Shared Store (</b><code>app_state.py</code><b>)</b></h3><p class="paragraph" style="text-align:left;">A single line keeps the in-memory index addressable from both FastAPI and the ADK tools:</p><div class="codeblock"><pre><code>from rag_store import MultimodalRagStore

RAG_STORE = MultimodalRagStore()</code></pre></div><p class="paragraph" style="text-align:left;">Both <code>server.py</code> and the ADK tool functions import <code>RAG_STORE</code> from here, so the agent always sees the same sources you uploaded through the UI.</p><h3 class="heading" style="text-align:left;"><b>2. The Multimodal Store (</b><code>rag_store.py</code><b>)</b></h3><p class="paragraph" style="text-align:left;">This is where most of the interesting code lives. A few constants set the contract:</p><div class="codeblock"><pre><code>EMBED_MODEL = &quot;gemini-embedding-2&quot;
DEFAULT_DIMENSIONS = 768
CHUNK_WORDS = 170
CHUNK_OVERLAP = 35
INLINE_MEDIA_LIMIT_BYTES = 18 * 1024 * 1024</code></pre></div><p class="paragraph" style="text-align:left;">We chunk text into roughly 170-word windows with 35-word overlap and embed each chunk separately. Anything bigger than ~18 MB or any audio/video file goes through the <b>Gemini File API</b> instead of inline bytes.</p><h4 class="heading" style="text-align:left;"><b>Embedding text with task prefixes</b></h4><p class="paragraph" style="text-align:left;">Gemini Embedding 2 supports task prefixes — small instructions like <code>&quot;task: retrieval document&quot;</code> or <code>&quot;task: question answering | query&quot;</code> that tell the model how this content will be used. Documents and queries get different prefixes, which measurably improves retrieval:</p><div class="codeblock"><pre><code>def _embed_text(self, text: str, task_prefix: str) -&gt; list[float]:
    content = f&quot;&#123;task_prefix&#125;: &#123;text&#125;&quot;
    client = self._require_client()

    result = client.models.embed_content(
        model=EMBED_MODEL,
        contents=[content],
        config=types.EmbedContentConfig(output_dimensionality=self.dimensions),
    )
    return result.embeddings[0].values</code></pre></div><p class="paragraph" style="text-align:left;">The interesting bit is <code>output_dimensionality=768</code>: Gemini Embedding 2 supports truncating to smaller, latency-friendlier vectors right at the API call, so you don&#39;t have to pay for storage or cosine math on the full embedding width.</p><h4 class="heading" style="text-align:left;"><b>Embedding files (PDFs, images, audio, video)</b></h4><p class="paragraph" style="text-align:left;">Multimodal is where Gemini Embedding 2 earns its keep. Small images and PDFs go inline; large files and all media go through the File API:</p><div class="codeblock"><pre><code>def _embed_file(self, data, mime_type, title, notes):
    client = self._require_client()

    use_file_api = (
        len(data) &gt; INLINE_MEDIA_LIMIT_BYTES
        or mime_type.startswith(&quot;video/&quot;)
        or mime_type.startswith(&quot;audio/&quot;)
    )
    if use_file_api:
        return self._embed_uploaded_file(data, mime_type, title), &quot;gemini-file-api&quot;

    part = types.Part.from_bytes(data=data, mime_type=mime_type)
    result = client.models.embed_content(
        model=EMBED_MODEL,
        contents=[part],
        config=types.EmbedContentConfig(output_dimensionality=self.dimensions),
    )
    return result.embeddings[0].values, &quot;gemini-inline&quot;</code></pre></div><p class="paragraph" style="text-align:left;">The File API path uploads the file, polls until its state is <code>ACTIVE</code>/<code>SUCCEEDED</code>, embeds via <code>Part.from_uri</code>, and then <b>deletes the uploaded file</b> in a <code>finally</code> block — important so you don&#39;t leak storage on every upload.</p><p class="paragraph" style="text-align:left;">To make a PDF or image still findable by its title (e.g., &quot;the launch deck&quot;), we blend the media vector with a text vector of the title plus user-provided notes:</p><div class="codeblock"><pre><code>media_vector, embedding_path = self._embed_file(...)
annotation_vector = self._embed_text(f&quot;&#123;title&#125;. &#123;notes&#125;&quot;, &quot;task: retrieval document&quot;)
vector = _blend_vectors(media_vector, annotation_vector)  # 68% media / 32% text</code></pre></div><p class="paragraph" style="text-align:left;">This is a small but very effective trick: native multimodal embeddings are great at semantic content, but humans often search by the label they gave the file.</p><h4 class="heading" style="text-align:left;"><b>Search: cosine similarity per chunk, deduplicated per source</b></h4><div class="codeblock"><pre><code>def search(self, query: str, top_k: int = 6) -&gt; dict[str, Any]:
    query_vector = self._embed_text(query, &quot;task: question answering | query&quot;)
    source_vectors = self._source_vectors()
    projections = self._pca_projection(&#123;**source_vectors, query_id: query_vector&#125;)
    ...
    for chunk in self.chunks:
        score = round(_cosine(query_vector, chunk.vector), 4)
        current = source_matches.get(chunk.source_id)
        if not current or score &gt; current[&quot;score&quot;]:
            source_matches[chunk.source_id] = &#123; ... &#125;
    matches = sorted(source_matches.values(), key=lambda m: m[&quot;score&quot;], reverse=True)[:top_k]</code></pre></div><p class="paragraph" style="text-align:left;">Three subtle decisions here: we score every chunk but keep only the <b>best chunk per source</b>, we project source vectors and the query vector together so the 3D view shares the same basis, and we return a fully-formed <code>space</code> snapshot so the frontend never has to ask twice.</p><h4 class="heading" style="text-align:left;"><b>PCA projection in pure Python</b></h4><p class="paragraph" style="text-align:left;">The <code>_pca_projection</code> method runs power iteration to find the top three principal components and projects every vector into 3D — no NumPy, no scikit-learn. That keeps the dependency list short and the projection deterministic per request.</p><h4 class="heading" style="text-align:left;"><b>The retrieval payload</b></h4><p class="paragraph" style="text-align:left;">The agent doesn&#39;t see raw chunks; it sees a clean, model-friendly payload:</p><div class="codeblock"><pre><code>def retrieval_payload(self, results):
    return &#123;
        &quot;provider&quot;: self.embedding_provider,
        &quot;matches&quot;: [
            &#123;
                &quot;citation&quot;: m[&quot;id&quot;],
                &quot;source&quot;: m[&quot;title&quot;],
                &quot;modality&quot;: m[&quot;modality&quot;],
                &quot;similarity&quot;: m[&quot;score&quot;],
                &quot;evidence&quot;: m[&quot;text&quot;],
            &#125;
            for m in results[&quot;matches&quot;]
        ],
    &#125;</code></pre></div><p class="paragraph" style="text-align:left;">This is the exact same packet that <code>/ask</code> returns to the frontend, which is how we guarantee the answer and the citation panel never drift.</p><h3 class="heading" style="text-align:left;"><b>3. The ADK Agent (</b><code>agentic_rag_agent/agent.py</code><b>)</b></h3><p class="paragraph" style="text-align:left;">A short, sharp ADK agent with two tools and a focused instruction:</p><div class="codeblock"><pre><code>def retrieve_relevant_context(query: str, top_k: int = 5) -&gt; dict:
    &quot;&quot;&quot;Retrieve the most relevant multimodal source evidence for a user question.&quot;&quot;&quot;
    return RAG_STORE.retrieval_tool(query=query, top_k=top_k)


def inspect_embedding_space() -&gt; dict:
    &quot;&quot;&quot;Inspect current sources, modalities, dimensions, and embedding provider.&quot;&quot;&quot;
    return RAG_STORE.space_tool()


def build_agent(retrieval_tool=retrieve_relevant_context) -&gt; Agent:
    return Agent(
        name=&quot;multimodal_agentic_rag_agent&quot;,
        model=&quot;gemini-3-flash-preview&quot;,
        description=&quot;Agentic RAG coordinator for a multimodal Gemini Embedding 2 workspace.&quot;,
        instruction=&quot;&quot;&quot;
You are the Google ADK coordinator for a multimodal agentic RAG workspace.

For every user question:
1. Use inspect_embedding_space to understand the current workspace.
2. Use retrieve_relevant_context with the user&#39;s question before answering.
3. Ground the answer in the retrieved evidence. Do not invent facts...
4. Do not include raw citation ids, source ids, bracket citations...
5. Start with a clear direct answer in 2-3 sentences.
6. If helpful, add a short &quot;Key points:&quot; section with simple hyphen bullets.
&quot;&quot;&quot;,
        tools=[inspect_embedding_space, retrieval_tool],
        generate_content_config=genai_types.GenerateContentConfig(
            temperature=0.25,
            max_output_tokens=900,
        ),
    )</code></pre></div><p class="paragraph" style="text-align:left;"><code>build_agent</code> accepts an injectable <code>retrieval_tool</code>. That&#39;s how <code>server.py</code> swaps in a closure that returns the <b>already-computed</b> retrieval packet, instead of letting the agent embed the query a second time.</p><h3 class="heading" style="text-align:left;"><b>4. The FastAPI Server (</b><code>server.py</code><b>)</b></h3><p class="paragraph" style="text-align:left;">The endpoint surface is small and predictable:</p><div style="padding:14px 15px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">Method</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">Endpoint</p></th><th class="bh__table_header" width="33%"><p class="paragraph" style="text-align:left;">What it does</p></th></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>GET</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/health</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Liveness, ADK availability, dimensions, source counts</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>GET</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/space</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Current sources, points, events, projection metadata</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>POST</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/sources/text</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Add a text source</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>POST</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/sources/url</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Fetch and index a public URL (SSRF-protected)</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>POST</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/sources/file</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Upload PDF, image, audio, or video</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>DELETE</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/sources/&#123;id&#125;</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Remove a source and its chunks</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>POST</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;"><code>/ask</code></p></td><td class="bh__table_cell" width="33%"><p class="paragraph" style="text-align:left;">Retrieve once, run ADK answer flow, return citations</p></td></tr></table></div><p class="paragraph" style="text-align:left;">The key piece is <code>/ask</code>. It retrieves once, builds a clean payload, and injects a closure into the agent so it can&#39;t redo the embedding:</p><div class="codeblock"><pre><code>@app.post(&quot;/ask&quot;)
async def ask(req: AskRequest):
    retrieval = await run_in_threadpool(RAG_STORE.search, req.question, req.top_k)
    retrieval_payload = RAG_STORE.retrieval_payload(retrieval)
    answer = await _run_adk_agent(req.question, retrieval_payload)
    trace = [
        &#123;&quot;agent&quot;: &quot;space_inspector&quot;,   &quot;status&quot;: &quot;complete&quot;, &quot;detail&quot;: ...&#125;,
        &#123;&quot;agent&quot;: &quot;retrieval_tool&quot;,    &quot;status&quot;: &quot;complete&quot;, &quot;detail&quot;: ...&#125;,
        &#123;&quot;agent&quot;: &quot;answer_synthesizer&quot;,&quot;status&quot;: &quot;complete&quot;, &quot;detail&quot;: ...&#125;,
    ]
    return &#123;
        &quot;answer&quot;: answer,
        &quot;matches&quot;: retrieval[&quot;matches&quot;],
        &quot;query_point&quot;: retrieval[&quot;query_point&quot;],
        &quot;trace&quot;: trace,
        &quot;space&quot;: retrieval[&quot;space&quot;],
    &#125;</code></pre></div><p class="paragraph" style="text-align:left;">And the closure injection inside <code>_run_adk_agent</code>:</p><div class="codeblock"><pre><code>async def _run_adk_agent(question: str, retrieval: dict[str, Any]) -&gt; str:
    def retrieve_relevant_context(query: str, top_k: int = 6) -&gt; dict:
        &quot;&quot;&quot;Return the exact retrieval packet already embedded for this request.&quot;&quot;&quot;
        return retrieval

    request_agent = build_agent(retrieve_relevant_context)
    request_runner = Runner(agent=request_agent, app_name=APP_NAME, session_service=session_service)
    session = await session_service.create_session(app_name=APP_NAME, user_id=USER_ID)
    content = genai_types.Content(
        role=&quot;user&quot;,
        parts=[genai_types.Part(text=f&quot;Question: &#123;question&#125;\nUse the retrieval tool result for this exact question.&quot;)],
    )
    final_text = &quot;&quot;
    async for event in request_runner.run_async(user_id=USER_ID, session_id=session.id, new_message=content):
        text = _event_text(event)
        if text:
            final_text = text
    return final_text</code></pre></div><p class="paragraph" style="text-align:left;">The agent thinks it&#39;s calling a real retrieval tool. It is — the tool just returns a cached result. This is a clean way to keep agent semantics while skipping a redundant embedding round-trip.</p><p class="paragraph" style="text-align:left;">A couple of safety details worth highlighting:</p><ul><li><p class="paragraph" style="text-align:left;"><b>SSRF protection</b>: <code>_validate_fetch_url</code> rejects non-HTTP schemes and resolves the hostname; if any returned IP is private, loopback, link-local, or reserved, ingestion fails. Set <code>ALLOW_PRIVATE_URLS=true</code> only when you really need it.</p></li><li><p class="paragraph" style="text-align:left;"><b>Threadpool offloading</b>: every blocking call (text chunking, file reads, search, PCA) runs in <code>run_in_threadpool</code> so the FastAPI event loop stays responsive.</p></li><li><p class="paragraph" style="text-align:left;"><b>Configurable CORS</b>: <code>ALLOWED_ORIGINS</code> is read from the env, defaulting to the Vite dev server.</p></li></ul><h3 class="heading" style="text-align:left;"><b>5. The Frontend (very brief)</b></h3><p class="paragraph" style="text-align:left;">The frontend is a single React/Vite app (<code>frontend/src/App.tsx</code>) that wraps three panels: a <b>source manager</b> for adding text/URLs/files, a <b>Q&A panel</b> that calls <code>/ask</code> and renders the answer plus a separate citations list, and a <b>3D embedding view</b> built on Three.js that uses the projection coordinates returned by the backend. Every source is one colored point (color encodes modality), and after a question the query point and the cited sources are highlighted in the same PCA basis.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h4 class="heading" style="text-align:left;"><b>Running the App</b></h4><p class="paragraph" style="text-align:left;">With our code in place, it&#39;s time to launch the app.</p><p class="paragraph" style="text-align:left;">Start the backend</p><div class="codeblock"><pre><code>python server.py</code></pre></div><p class="paragraph" style="text-align:left;">The backend listens on <code>http://localhost:8897</code>.</p><p class="paragraph" style="text-align:left;">Start the frontend in a second terminal:</p><div class="codeblock"><pre><code>cd multimodal_agentic_rag/frontend
npm install
npm run dev -- --port 5177</code></pre></div><p class="paragraph" style="text-align:left;">If your backend lives on a different port, point the frontend at it:</p><div class="codeblock"><pre><code>VITE_API_URL=http://localhost:8897 npm run dev -- --port 5177</code></pre></div><ol start="1"><li><p class="paragraph" style="text-align:left;">Add a few sources — try a paragraph of text, a public URL, a PDF, and an image.</p></li><li><p class="paragraph" style="text-align:left;">Watch them appear as colored points in the embedding view.</p></li><li><p class="paragraph" style="text-align:left;">Ask a question in the Q&A panel.</p></li><li><p class="paragraph" style="text-align:left;">Inspect the answer, the cited sources, and the agent trace.</p></li><li><p class="paragraph" style="text-align:left;">Notice the orange query point land near the sources the agent cites.</p></li></ol><p class="paragraph" style="text-align:left;">A quick health check from the terminal:</p><div class="codeblock"><pre><code>curl http://localhost:8897/health</code></pre></div><p class="paragraph" style="text-align:left;">Expected response shape on a fresh start (the store begins empty):</p><div class="codeblock"><pre><code>&#123;
  &quot;status&quot;: &quot;ok&quot;,
  &quot;adk&quot;: true,
  &quot;setup_error&quot;: &quot;&quot;,
  &quot;sources&quot;: 0,
  &quot;chunks&quot;: 0,
  &quot;dimensions&quot;: 768,
  &quot;provider&quot;: &quot;gemini-embedding-2&quot;,
  &quot;modalities&quot;: &#123;&#125;,
  &quot;chunk_modalities&quot;: &#123;&#125;,
  &quot;projection&quot;: &quot;pca_3d&quot;</code></pre></div></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Working Application Demo</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/1rPZZIfekWs" width="100%"></iframe><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Conclusion</b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">You&#39;ve now built a multimodal agentic RAG app that puts text, URLs, PDFs, images, audio, and video into a single Gemini Embedding 2 space, retrieves with cosine similarity over chunked vectors, and uses a tightly-scoped Google ADK agent to write grounded, citation-friendly answers, without a separate vector database, in a few hundred lines of Python.</p><p class="paragraph" style="text-align:left;">A few directions worth exploring from here:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Swap the in-memory store</b> for a managed vector DB (pgvector, Qdrant, Vertex AI Vector Search) and persist the chunk metadata.</p></li><li><p class="paragraph" style="text-align:left;"><b>Add re-ranking</b> with a cross-encoder or a Gemini reranker between cosine retrieval and the agent.</p></li><li><p class="paragraph" style="text-align:left;"><b>Background ingestion</b> with a queue (Celery, RQ, or a simple async worker) so large videos don&#39;t block the API.</p></li><li><p class="paragraph" style="text-align:left;"><b>Evals:</b> wire a small eval set with question/answer pairs and track citation precision and answer faithfulness over changes.</p></li><li><p class="paragraph" style="text-align:left;"><b>Auth + multi-tenancy</b> so different users see different workspaces.</p></li><li><p class="paragraph" style="text-align:left;"><b>Observability</b>: log the retrieval packet alongside the final answer; the single-retrieval contract makes faithfulness audits straightforward.</p></li></ul><p class="paragraph" style="text-align:left;">Keep experimenting with different configurations and features to build more sophisticated AI applications.</p><p class="paragraph" style="text-align:left;">We share hands-on tutorials like this 2-3 times a week, to help you stay ahead in the world of AI. <span style="text-decoration:underline;"><b><a class="link" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">If you&#39;re serious about leveling up your AI skills and staying ahead of the curve, subscribe now and be the first to access our latest tutorials.</a></b></span></p><p class="paragraph" style="text-align:left;"><b>Don’t forget to share this tutorial on your social channels and tag Unwind AI (</b><span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b>, </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span><b>) to support us!</b></p></div><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=build-a-multimodal-agentic-rag-app-with-gemini-embedding-2-and-google-adk"><span class="button__text" style=""> Subscribe now for FREE - Get instant access to more LLM, RAG & AI Agent tutorials </span></a></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=a2e798cb-b59a-463c-a2c0-7bc5d0c5e8e9&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLM with 12M Context Window </title>
  <description>+ Free web Search and Fetch for your Hermes and Claws</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/808e4bb0-8342-473d-b1c9-93aff84309cc/upload_9ed1384faf7ad83e.jpg" length="123575" type="image/jpeg"/>
  <link>https://www.theunwindai.com/p/llm-with-12m-context-window</link>
  <guid isPermaLink="true">https://www.theunwindai.com/p/llm-with-12m-context-window</guid>
  <pubDate>Wed, 06 May 2026 12:30:00 +0000</pubDate>
  <atom:published>2026-05-06T12:30:00Z</atom:published>
    <dc:creator>Shubham Saboo</dc:creator>
    <dc:creator>Gargi Gupta</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #6553a2; }
  .bh__table_cell { padding: 5px; background-color: #ffffff; }
  .bh__table_cell p { color: #030712; font-family: 'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#d0c7e2; }
  .bh__table_header p { color: #6553a2; font-family:'Open Sans','Segoe UI','Apple SD Gothic Neo','Lucida Grande','Lucida Sans Unicode',sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">Today’s top AI Highlights:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/alex_whedon/status/2051663268704636937?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Attention Is All You Need - Just 1/1000th of It</a></b></p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.tinyfish.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Free web Search and Fetch, for every dev and AI agent</a></b></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/ashpreetbedi/status/2049180168200106150?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Stop RAG-ing, start Grepping your company knowledge</b></a></p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/HeyGen/status/2051697813554405384?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Make your Hermes Agent your video editor with one Skill</a></b></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/Manavarya09/design-extract?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Extract a website’s complete design system with one command</b></a></p></li></ol><p class="paragraph" style="text-align:start;">& so much more!</p><p class="paragraph" style="text-align:start;"><i><b>Read time: 3 mins</b></i></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>AI Tutorial </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.theunwindai.com/p/anatomy-of-agent-skills?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Anatomy of Agent SKILLS</a></b></p><p class="paragraph" style="text-align:left;">Your agent has a 200k token context window. </p><p class="paragraph" style="text-align:left;">The 400 tokens of instructions it actually needs are buried under tool definitions, reference docs, and brand guides it never asked for. So it ignores them. </p><p class="paragraph" style="text-align:left;">This is the most common reason agents fail in production. It&#39;s not a model problem or a framework problem. </p><p class="paragraph" style="text-align:left;">In this blog, you&#39;ll learn the anatomy of Agent Skills: why the first two lines of SKILL.md are the most important writing you&#39;ll do, and how the LLM itself routes queries to the right skill without embeddings or retrieval layers. </p><p class="paragraph" style="text-align:left;">Read on to learn the five parts that make skills work, then pick one workflow you do every week and ship your first skill today.</p><div class="embed"><a class="embed__url" href="https://www.theunwindai.com/p/anatomy-of-agent-skills?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank"><img class="embed__image embed__image--left" src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/e45fcf50-862b-47e8-8f08-f74aa7cfc466/Anatomy_of_Agent_SKILLS.png?t=1777437345"/><div class="embed__content"><p class="embed__title"> Anatomy of Agent SKILLS </p><p class="embed__description"> Same agent, much less context </p></div></a></div><p class="paragraph" style="text-align:left;">We share hands-on tutorials like this every week, designed to help you stay ahead in the world of AI. <span style="text-decoration:underline;"><b><a class="link" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">If you&#39;re serious about leveling up your AI skills and staying ahead of the curve, subscribe now and be the first to access our latest tutorials.</a></b></span></p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b>Unwind AI</b> (<b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">X</a></b><b>, </b><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=referral&utm_campaign=last-week-in-ai-a-weekly-unwind" target="_blank" rel="noopener noreferrer nofollow">LinkedIn</a></b><b>, </b><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Threads</a></b>) to support us!</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Latest Developments </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><a class="link" href="https://x.com/alex_whedon/status/2051663268704636937?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Attention Is All You Need - Just 1/1000th of It</b></a></h3><div class="image"><a class="image__link" href="https://x.com/alex_whedon/status/2051663268704636937?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/911a23ed-5a23-4a31-a328-943daaa2ddaa/image.png?t=1778042931"/></a></div><p class="paragraph" style="text-align:left;">12 million tokens of context in a single pass, and it does it at roughly 1/1000th the attention compute of current frontier models. </p><p class="paragraph" style="text-align:left;">Meet <b>SubQ</b>, the first large language model built on a fully subquadratic sparse attention (SSA) architecture, where compute scales linearly with context length instead of quadratically.</p><p class="paragraph" style="text-align:left;">Transformers compare every token to every other token, which means doubling input length quadruples the compute. SubQ&#39;s architecture focuses only on the token relationships that actually matter, making million-token workloads fast and cheap enough to be practical. </p><p class="paragraph" style="text-align:left;">The model is entering private beta today with an API, a coding agent called SubQ Code, and a long-context search tool called SubQ Search.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Benchmark Performance</b>: SubQ 1M-Preview scores 95% on RULER 128K and 81.8 on SWE-Bench Verified, putting it on par with or ahead of Opus 4.6 and Deepseek V4 Pro on both long-context accuracy and code tasks.</p></li><li><p class="paragraph" style="text-align:left;"><b>Speed & Efficiency</b>: Its sparse attention runs 52x faster than FlashAttention at 1M tokens while requiring 63% less compute.</p></li><li><p class="paragraph" style="text-align:left;"><b>5% Cost of Opus 4.7</b>: Though the pricing is not out, the team claims it costs &lt;5% Opus&#39;s cost at scale, with RULER 128K running for $8 vs ~$2,600. Take this with a grain of salt for now!</p></li><li><p class="paragraph" style="text-align:left;"><b>Private Beta:</b> All three products (API, Code, Search) are available via early access at <a class="link" href="https://subq.ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">subq.ai</a>. </p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h2 class="heading" style="text-align:left;"><a class="link" href="https://www.tinyfish.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Free web Search and Fetch, for every dev and AI agent</b></a></h2><div class="image"><a class="image__link" href="https://www.tinyfish.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/afe49e96-4af7-4032-bd7e-5147cb730982/Stand_By_We_re_Going_Live_Shortly__1_.png?t=1778043034"/></a></div><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.tinyfish.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">TinyFish</a></b> just made their Web Search and Fetch endpoints free with generous rate limits. Forever. No credit card or “7-day” trial”. Just sign up and grab your API key.</p><p class="paragraph" style="text-align:left;">Search returns structured JSON for agents. Fetch renders any URL in a real browser with full JavaScript, SPAs, anti-bot, all of it, strips the unnecessary content, and returns clean markdown.</p><p class="paragraph" style="text-align:left;">Everything runs on TinyFish&#39;s own custom Chromium fleet. Owning the stack end-to-end makes their Search and Fetch both free and fast.</p><ul><li><p class="paragraph" style="text-align:left;">Works with Claude Code, OpenClaw, Hermes Agent, Cursor, Codex, and any agent framework</p></li><li><p class="paragraph" style="text-align:left;">Available via API, MCP, Python + TypeScript SDKs, CLI, and Skills</p></li><li><p class="paragraph" style="text-align:left;">One API key. No credit card.</p></li></ul><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.tinyfish.ai/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Grab your API key now!</a></b></p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><h3 class="heading" style="text-align:left;"><b><a class="link" href="https://x.com/ashpreetbedi/status/2049180168200106150?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Stop RAG-ing, start Grepping your company knowledge</a></b></h3><div class="image"><a class="image__link" href="https://x.com/ashpreetbedi/status/2049180168200106150?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8731e9be-b714-4ba8-8a9c-283806cbc840/image.png?t=1778044571"/></a></div><p class="paragraph" style="text-align:left;">Your company&#39;s best knowledge is rotting in Slack threads nobody will ever search again. </p><p class="paragraph" style="text-align:left;">And no, RAG is not the solution. The index is always stale. The chunks land at the wrong boundaries.</p><p class="paragraph" style="text-align:left;">Turns out, coding agents already cracked this. They don&#39;t search, they <code>grep</code></p><p class="paragraph" style="text-align:left;"><b>Scout</b>, an open-source context agent from Agno, borrows the trick and <i>navigates</i> your information sources live. It connects to Slack, Google Drive, Linear, MCP servers, and more, walking each source&#39;s native API at query time to assemble real answers with real citations. </p><p class="paragraph" style="text-align:left;">As it works, it builds its own wiki and CRM. Say &quot;Josh from Anthropic shared a paper on RLMs&quot; and Scout files Josh as a contact, parses the paper into a wiki page, and links them together. </p><p class="paragraph" style="text-align:left;">The whole thing is open-source, ready to fork and customize.</p><p class="paragraph" style="text-align:left;"><b>Key Highlights:</b></p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Context Providers</b>: Instead of exposing dozens of API-specific tools to the main agent, Scout wraps each source behind a thin sub-agent layer. The main agent sees <code>query_slack</code>, not Slack&#39;s twelve endpoints, keeping context clean.</p></li><li><p class="paragraph" style="text-align:left;"><b>Navigation over search</b>: Scout queries live APIs at request time, so a Slack message sent thirty seconds ago is immediately available, and citations always point to real, openable paths.</p></li><li><p class="paragraph" style="text-align:left;"><b>Self-building CRM and wiki</b>: It populates a Postgres-backed CRM and a knowledge wiki as it learns. It even creates new database tables on demand.</p></li><li><p class="paragraph" style="text-align:left;"><b>Ready to clone and use</b>: Ships with Docker Compose, connects to Agno&#39;s AgentOS for multi-user sessions and scheduled tasks, and plugs into Slack with full thread history. More connectors coming soon!</p></li></ol></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#FFFFFF;"><b>Quick Bites </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/OpenAIDevs/status/2050275713824211041?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Codex gets a Tamagotchi</b></a><br>OpenAI shipped pets for Codex. These are animated companions that double as a persistent status overlay for your coding agent. The pet visually maps to whether Codex is actively working, waiting for input, or flagging something for review. Think of it as the most adorable process monitor you never asked for.</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://x.com/claudeai/status/2051679629488865498?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Just another day of the Claude team shipping</a></b><br>Anthropic just dropped ten agent templates built specifically for finance work, like pitchbooks, KYC screening, month-end close, valuation checks, the whole grind. They plug into Cowork and Claude Code or run autonomously as Managed Agents, and they come wired to data sources like Moody&#39;s, Third Bridge, and S&P Capital IQ. Oh, and Claude now works inside Excel, PowerPoint, Word, and Outlook with context that follows you across apps — so yes, your comps model can become a deck without explaining everything twice. Install them as plugins in Cowork and Claude Code</p><p class="paragraph" style="text-align:left;"></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/HeyGen/status/2051697813554405384?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Make your Hermes Agent your video editor with one Skill</b></a><br>Hermes Agents can now spin up full videos, courtesy HyperFrames Agent Skill by HeyGen. Just do <code>$ hermes skills install hyperframes</code>, and your agent becomes a video editor that treats HTML as the source of truth for video. Feed it an X post, a PDF, or a GitHub repo, and it&#39;ll script, animate with GSAP, lay captions over TTS narration, and render a finished MP4, all orchestrated end-to-end by the agent itself.</p></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#6553a2;border-radius:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:center;"><span style="color:#ffffff;"><b>Tools of the Trade </b></span></h2></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:5.0px 5.0px 5.0px 5.0px;"><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/Manavarya09/design-extract?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>Designlang</b></a>: Extract any website&#39;s complete design system with one command. It reads the design system off the live DOM, and emits 17+ files — DTCG tokens, Tailwind config, shadcn theme, Figma variables, motion tokens, typed component anatomy, brand voice, page-intent labels, and a paste-ready prompt pack for v0 / Lovable / Cursor / Claude Artifacts.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.interactlabs.ai/blog-article/introducing-interact-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Interact AI</a></b>: Replaces your static website with an adaptive, conversational interface that recomposes itself per visitor in real time. A founder sees compliance content, a CISO sees security controls, all generated on the fly from your data. It&#39;s not a chatbot widget in the corner; the conversation <i>is</i> the page, and everything the visitor says carries through into signup and product onboarding. </p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://github.com/yizhiyanhua-ai/fireworks-tech-graph?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow"><b>fireworks-tech-graph</b></a>: A skill that turns plain descriptions of your system into polished SVG + PNG technical diagrams. It ships with 5 visual styles, 8 diagram types, and built-in knowledge of AI/agent patterns like RAG pipelines, Mem0 memory layers, and multi-agent flows.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Awesome LLM Apps</a></b><b> </b>- A curated collection of LLM apps with RAG, AI Agents, multi-agent teams, MCP, voice agents, and more. The apps use models from OpenAI, Anthropic, Google, and open-source models like DeepSeek, Qwen, and Llama that you can run locally on your computer. <br><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">(Now accepting GitHub sponsorships)</a></p></li></ol><div class="image"><a class="image__link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5842cecc-c30d-48e9-a805-783f55950a3e/image.png?t=1755755385"/></a></div></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;border-color:#6553a2;border-radius:5px;border-style:solid;border-width:1px;margin:5.0px 5.0px 5.0px 5.0px;padding:5.0px 5.0px 5.0px 5.0px;"><p class="paragraph" style="text-align:left;">That’s all for today! See you tomorrow with more such AI-filled content.</p><p class="paragraph" style="text-align:left;">Don’t forget to share this newsletter on your social channels and tag <b><a class="link" href="https://www.theunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow">Unwind AI</a></b> to support us!</p><p class="paragraph" style="text-align:start;"><b>Unwind AI</b> - <span style="text-decoration:underline;"><b><a class="link" href="https://x.com/unwind_ai_?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">X</a></b></span> | <span style="text-decoration:underline;"><b><a class="link" href="https://www.linkedin.com/company/unwind-ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">LinkedIn</a></b></span><b> </b>|<b> </b><span style="text-decoration:underline;"><b><a class="link" href="https://www.threads.net/@unwind_ai?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Threads</a></b></span></p><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b><a class="link" href="https://github.com/Shubhamsaboo/awesome-llm-apps?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Awesome LLM Apps</a></b></span><b> | </b><span style="text-decoration:underline;"><b><a class="link" href="https://sponsorunwindai.com/?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window" target="_blank" rel="noopener noreferrer nofollow" style="color: #6553a2">Sponsor Us</a></b></span></p><p class="paragraph" style="text-align:start;"><b>PS:</b> We curate this AI newsletter every day for FREE, your support is what keeps us going. If you find value in what you read, share it with at least one, two (or 20) of your friends 😉 </p></div><p class="paragraph" style="text-align:left;"></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.theunwindai.com/subscribe?utm_source=www.theunwindai.com&utm_medium=newsletter&utm_campaign=llm-with-12m-context-window"><span class="button__text" style=""> Subscribe now for FREE! </span></a></div><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F84ac330c-894c-4f61-ac66-e747ce8b32eb%2Flogo.png%3Fv%3D1789183230&publication_name=unwind+ai&utm_campaign=189810ca-6424-4ee1-a562-fc8a1bfaf464&utm_medium=post_rss&utm_source=unwind_ai">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
