<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The AI Timeline</title>
    <description>Follow The Latest Cutting Edge AI Research in 5 minutes a week.</description>
    
    <link>https://mail.bycloud.ai/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/Vy37NcFo03.xml" rel="self"/>
    
    <lastBuildDate>Fri, 10 Jul 2026 03:35:19 +0000</lastBuildDate>
    <pubDate>Tue, 07 Jul 2026 18:30:00 +0000</pubDate>
    <atom:published>2026-07-07T18:30:00Z</atom:published>
    <atom:updated>2026-07-10T03:35:19Z</atom:updated>
    
      <category>Machine Learning</category>
      <category>Software Engineering</category>
      <category>Artificial Intelligence</category>
    <copyright>Copyright 2026, The AI Timeline</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/23043e0b-1a8b-4e75-85b4-a980ed68d059/143861144.png</url>
      <title>The AI Timeline</title>
      <link>https://mail.bycloud.ai/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>You Only Need 1 Layer for RLVR?</title>
  <description>plus more about AdaJEPA, Program-as-Weights, The World Is In Your Mind, and Dual On-policy Distillation</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fc2960e0-8280-4123-b4f9-9358761d456f/issue_115.jpg" length="238714" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/you-only-need-1-layer-for-rlvr</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/you-only-need-1-layer-for-rlvr</guid>
  <pubDate>Tue, 07 Jul 2026 18:30:00 +0000</pubDate>
  <atom:published>2026-07-07T18:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>July 1st ~ July 7th</i><br><i>#115 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 2.2k</span></span> <a class="link" href="https://hy.tencent.com/research/hy3?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Tencent has launched Hy3</a>, a new 295-billion parameter Mixture of Experts model optimized for productivity and agentic workflows. This model is released under the Apache 2.0 license, and it provides an accessible and cost-effective option for commercial applications. You can try it on <a class="link" href="https://huggingface.co/tencent/Hy3?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d8f5dbfb-e7f2-4907-a42c-2fa0a2acb94e/2026070523041077_333c5c42bb6615e0dbcb4a36e66db95d.png?t=1783446975"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.4k</span></span> <a class="link" href="https://mistral.ai/news/leanstral-1-5/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Mistral has released Leanstral 1.5</a> (Le Chaton L∃∀N), a formal-reasoning model designed for mathematical theorem proving and code verification using the Lean language. The model demonstrates efficient test-time scaling on the PutnamBench benchmark and includes a translation pipeline to assist with verifying Rust code. You can try it on <a class="link" href="https://github.com/mistralai/LeanstralSafeVerify?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c2f385b9-bd97-43ab-8e12-af48fe29aadb/Leanstral-charts_Z1KXfze.webp?t=1783447133"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.7k</span></span> <a class="link" href="https://longcat.chat/blog/longcat-2.0/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Meituan has introduced LongCat-2.0</a>, a 1.6-trillion parameter Mixture of Experts model designed for agentic coding with a 1-million token context window. The model uses sparse attention and dynamic expert routing to optimize computational efficiency during multi-file codebase analysis and complex reasoning tasks.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2596b5ea-35de-4a2c-8fee-8edc5d54d44d/image.png?t=1783447229"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Optimization Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;">We just added a <b>new chapter on Optimization</b>, that goes through the history, the key techniques, and the current state of optimizers that frontier model uses. </p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an exclusive newsletter offer, where you would get 40% off on the yearly plan for our users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="is-one-layer-enough-training-a-sing">Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training</h2><p class="paragraph" style="text-align:left;"><i>Zhang et al. [University of Minnesota, Peking University, Amazon]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 22k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We train AI models by updating every single layer of the network uniformly, assuming that every part must adapt together to make progress. However, this approach has kept us in the dark about where the learning actually takes place inside the model.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c05696fc-3124-4ab0-a0c0-6e337df06aca/x1.png?t=1783445213"/><div class="image__source"><span class="image__source_text"><p>Layer contribution across all seven models studied in this work</p></span></div></div><p class="paragraph" style="text-align:left;">To investigate this, researchers trained individual layers of various language models while freezing the rest of the network. They discovered that vast majority of RL gains are concentrated in just a small handful of layers located right in the middle of the model.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5355b047-fbfc-4ef5-a283-cc5df2ae6d34/x2.png?t=1783445238"/><div class="image__source"><span class="image__source_text"><p>Layer contribution 𝒞(k) across model scales.</p></span></div></div><p class="paragraph" style="text-align:left;">Training just one of these middle layers could recover almost all (and sometimes even exceed) the performance improvements achieved by updating the entire model. Meanwhile, the layers at the very beginning and the very end of the network contributed almost nothing to the final results.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/860be75c-b0bb-4e1f-863b-d3da76b2a21c/x3.png?t=1783445284"/><div class="image__source"><span class="image__source_text"><p>Cross-dataset consistency of layer contribution on Qwen3-1.7B-Base.</p></span></div></div><p class="paragraph" style="text-align:left;">This middle-focused behavior remained consistent across different model sizes, training datasets, and diverse tasks like mathematics, coding, and multi-step decision-making.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2607.01232?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="dopd-dual-onpolicy-distillation">DOPD: Dual On-policy Distillation</h2><p class="paragraph" style="text-align:left;"><i>Yu et al. [NUS, MMLab, PKU, Explore Academy]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 424 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Distillation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Training smaller AI models is a delicate art, especially when we try to assist the &quot;teacher&quot; model with extra hints that the student will not have access to in the real world.</p><p class="paragraph" style="text-align:left;">When teacher models rely on privileged clues to guide a student, it often creates a &quot;privilege illusion&quot; where the student tries to copy the teacher’s final outputs without actually understanding the underlying logic. This information mismatch confuses the student, leading to unstable training and a decline in its independent reasoning abilities.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/231776ca-7d74-4613-8a09-deea8f2f7384/x1.png?t=1783445134"/><div class="image__source"><span class="image__source_text"><p>Performance comparison of our DOPD with competing approaches across eight benchmarks in terms of average across all benchmarks (upper bigger bars) and individual values of each benchmark (lower small bars).</p></span></div></div><p class="paragraph" style="text-align:left;">To address this challenge, researchers developed a framework called Dual On-policy Distillation, or DOPD. The system works by analyzing what they call the &quot;privilege advantage gap&quot;. This metric compares the prediction confidence of the teacher and the student when both are given the same privileged hints. By evaluating this gap on a token-by-token basis, the framework can pinpoint exactly which parts of a response represent genuine, transferable reasoning skills and which parts are merely shortcuts derived from the extra clues.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3ad2b006-ce30-4cc0-b66e-7597c2d2d0dd/x2.png?t=1783445155"/><div class="image__source"><span class="image__source_text"><p>Comparison of existing (a) standard distillation, (b) self distillation, and (c) adaptive distillation paradigms with our proposed (d) dual distillation paradigm.</p></span></div></div><p class="paragraph" style="text-align:left;">Once these learning scenarios are identified, the system routes different teaching strategies to different tokens. For words where the teacher demonstrates a clear capability advantage, the student receives strong, full-vocabulary guidance to absorb that deeper logic.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aeff4bdd-437b-443e-8288-84155dc15ad3/x3.png?t=1783445174"/><div class="image__source"><span class="image__source_text"><p>Comparison of (a) performance and (b) entropy on OPD variants with privileged information. Here, T., S., and Priv. denote teacher policy, student policy and with privileged information, respectively.</p></span></div></div><p class="paragraph" style="text-align:left;">On the other hand, tokens that are either highly obvious or too uncertain, the system dials back the teacher&#39;s influence and instead uses the student&#39;s own internal consistency as a stabilizing guide.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.30626?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="orca-the-world-is-in-your-mind">Orca: The World is in Your Mind</h2><p class="paragraph" style="text-align:left;"><i>Kim et al. [Orca Team, Beijing Academy of Artificial Intelligence]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 430 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> World Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Most of today&#39;s AI systems are highly specialized, focusing strictly on predicting the next word, the next video frame, or the next robotic motion in isolation. This fragmented approach prevents them from developing a unified, common-sense understanding of physical cause and effect. To address this, researchers developed Orca, a foundational world model designed to learn a unified, hidden representation of physical transitions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/042cbdc4-a592-41fa-b370-18b730bf90ba/x1.png?t=1783445042"/><div class="image__source"><span class="image__source_text"><p>The Orca’s overall framework.</p></span></div></div><p class="paragraph" style="text-align:left;">Orca builds this understanding using two complementary learning pathways. The first is an &quot;unconscious&quot; learning process, where the model watches thousands of hours of continuous, unlabeled video to quietly absorb the dense, natural laws of how objects move and scenes change.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ab1792d1-b408-4d74-bbb7-1ab6077c3a56/x2.png?t=1783445063"/><div class="image__source"><span class="image__source_text"><p>Overview of Encoder</p></span></div></div><p class="paragraph" style="text-align:left;">The second is a &quot;conscious&quot; pathway, which pairs video segments with language annotations to help the model connect these physical transitions to explicit human concepts, causal relationships, and task instructions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/92040232-cfe0-4939-8443-11e838983d2c/x3.png?t=1783445086"/><div class="image__source"><span class="image__source_text"><p>Overview of pre-training data</p></span></div></div><p class="paragraph" style="text-align:left;">To prove that this learned internal map is actually useful, the researchers froze Orca’s core architecture and attached lightweight, trainable decoders to translate its hidden states into text, predicted future images, and physical robot actions.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.30534?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">AdaJEPA: An Adaptive Latent World Model</h2><p class="paragraph" style="text-align:left;"><i>Karan and Du [Harvard University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 855 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM World Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Most AI &quot;world models&quot; are kept frozen after their initial training. When these models encounter minor real-world variations, such as unexpected lighting, shifting friction, or novel obstacles, their predictions quickly drift, causing their entire planning process to fail.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0787d02f-2b2b-4aa8-8b67-d12c85ef60f5/x2.png?t=1783444964"/><div class="image__source"><span class="image__source_text"><p>AdaJEPA performs test-time adaptation during closed-loop MPC</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed AdaJEPA, an adaptive framework that introduces a continuous loop of planning, acting, adapting, and replanning. Instead of blindly relying on static pre-trained assumptions, this system executes a planned action, observes the actual outcome, and immediately uses that real-world feedback as a self-supervised signal to recalibrate its internal model.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8186671c-b73d-4eac-91fe-fe40a516249e/x1.png?t=1783444983"/><div class="image__source"><span class="image__source_text"><p>(a)AdaJEPA: Plan–Act–Adapt–Replan Loop</p></span></div></div><p class="paragraph" style="text-align:left;">By performing lightweight, single-step updates to just a small subset of its parameters right before the next plan is made, the system corrects its course dynamically without needing any extra expert guidance.</p><p class="paragraph" style="text-align:left;">The researchers discovered that this real-time tuning gives steady improvements in planning success across a variety of challenging tasks, including navigating unfamiliar mazes and manipulating objects with unseen shapes or altered physics.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3967a01b-5b84-4926-9354-ad5429b82be7/x4.png?t=1783445006"/><div class="image__source"><span class="image__source_text"><p>Planning Success under Shape Shifts (top) and Visual Shifts (bottom)</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.32026?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="programas-weights-a-programming-par">Program-as-Weights: A Programming Paradigm for Fuzzy Functions</h2><p class="paragraph" style="text-align:left;"><i>Zhang et al. [Meta, UT Austin, UCL, UC Berkeley, Harvard University, Periodic Labs]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 805 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Weights </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Many common programming tasks, such as filtering system logs, repairing malformed data, or sorting search results by intent, are too subjective for traditional, rigid code, but calling giant cloud-based models for every minor task is expensive, slow, and difficult to keep consistent. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cd503195-7ec3-42c8-9f57-14a0047c9b2e/x1.png?t=1783444815"/><div class="image__source"><span class="image__source_text"><p>Overview of the Program-as-Weights paradigm</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers have designed a system called Program-as-Weights, or PAW. Instead of querying a massive cloud artificial intelligence every time a user inputs data, PAW uses a specialized compiler to translate a developer’s natural-language instructions into a tiny, reusable file of custom neural network weights.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/177c4021-74bf-4572-af7e-df7681d6054b/architecture_lora.png?t=1783444828"/><div class="image__source"><span class="image__source_text"><p>Text-to-LoRA instantiation of PAW</p></span></div></div><p class="paragraph" style="text-align:left;">This small file acts as a specialized adapter that can be instantly loaded onto a lightweight, frozen interpreter model already installed on the user&#39;s local device. In essence, the heavy lifting of understanding the task happens only once during compilation, generating a compact, private tool that can be called repeatedly offline.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/64fae5a8-a5b4-404b-9976-268946804edf/web_ui_2.png?t=1783444900"/><div class="image__source"><span class="image__source_text"><p>Interactively test the compiled program. Users can provide test inputs and inspect the corresponding outputs, enabling rapid validation and refinement before download</p></span></div></div><p class="paragraph" style="text-align:left;">To make this possible, the researchers trained their compiler on a massive new dataset containing ten million examples across hundreds of text-processing categories. They found that a tiny local interpreter running these compiled programs outperformed a model over <b>fifty times its size</b> that was prompted in the traditional way.</p><p class="paragraph" style="text-align:left;">This model was executed on standard consumer laptops, and it achieved fast execution speeds while requiring only about <b>one-fiftieth of the memory</b> typically needed by larger systems.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2607.02512?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-only-need-1-layer-for-rlvr"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/nYwid6Q5HXk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=88551bde-4a30-451c-a766-16541e9dcd09&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>DeepSeek Just dropped a new speculative decoding method!</title>
  <description>plus more about Tapered LMs, Improved LLDMs, AutoData, and You Don&#39;t Need To Run Every Eval</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4bdd02c0-a05b-4aaa-b590-f1f46af1ea1d/issue_114.jpg" length="246103" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/deepseek-just-dropped-a-new-speculative-decoding-method</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/deepseek-just-dropped-a-new-speculative-decoding-method</guid>
  <pubDate>Tue, 30 Jun 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-06-30T19:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>June 23rd ~ June 30th</i><br><i>#114 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 39k</span></span> OpenAI has announced a limited preview of its new <a class="link" href="https://openai.com/index/previewing-gpt-5-6-sol/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">GPT-5.6 model family</a>, which includes the flagship Sol, the balanced Terra, and the cost-efficient Luna. The flagship Sol model can handle cybersecurity tasks and complex command-line workflows. It is best suited for long-horizon security tasks including vulnerability research and exploitation.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5e25eb62-2e14-4be0-95d7-e1b6189b9b3d/image.png?t=1782837079"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 6.7k</span></span> The <a class="link" href="https://deep-reinforce.com/ornith_1_0.html?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Ornith-1.0 family of open-source models</a> (ranging from 9B dense to 397B MoE parameters) specializes in agentic coding tasks. These models are post-trained on Gemma 4 and Qwen 3.5, and they use a reinforcement learning strategy that jointly optimizes task-specific scaffolds and solution rollouts to improve coding outcomes. You can try it on <a class="link" href="https://huggingface.co/collections/deepreinforce-ai/ornith-10?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5d8f3113-b7dd-4e74-bb03-ae40580e0ce3/ornith-01-202606231007.png?t=1782837253"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 4.7k</span></span> Alibaba&#39;s Qwen team has open-sourced <a class="link" href="https://qwen.ai/blog?id=qwen-agentworld&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method#agentworldbench-agentworldbench" target="_blank" rel="noopener noreferrer nofollow">AgentWorldBench</a>, a seven-domain benchmark for environment simulation, alongside <a class="link" href="https://qwen.ai/blog?id=qwen-agentworld&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Qwen-AgentWorld-35B-A3B</a>, a Mixture-of-Experts model designed for world modeling. This uses two approaches: using the world model as a controllable simulator for reinforcement learning, and internalizing environment prediction directly within the agent foundation model. You can try it on <a class="link" href="https://github.com/QwenLM/Qwen-AgentWorld?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/collections/Qwen/qwen-agentworld?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ad1b910f-9b6f-4fd2-9dda-d7cdeb909695/bench_overview.png?t=1782837558"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ HOT </span></span> Anthropic has introduced <a class="link" href="https://www.anthropic.com/news/claude-sonnet-5?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Claude Sonnet 5</a>, its most agentic Sonnet model yet. The model can make plans, use tools like browsers and terminals, and complete complex coding, reasoning, and knowledge-work tasks with more autonomy than previous Sonnet models. It is best suited for developers and teams who need strong agentic performance at a lower cost than larger Opus-class models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aac7ecc6-6fd5-4107-a447-2c94c0102a11/image.png?t=1782845959"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Optimization Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;">We just added a <b>new chapter on Optimization</b>, that goes through the history, the key techniques, and the current state of optimizers that frontier model uses. </p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an exclusive newsletter offer, where you would get 40% off on the yearly plan for our users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="tapered-language-models">Tapered Language Models</h2><p class="paragraph" style="text-align:left;"><i>Bayat et al. [Mila, Cornell University, Université de Montréal, CIFAR AI Chair]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 151 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM architecture </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Currently, almost all LLMs distribute their learning capacity uniformly across every layer in their architecture, treating early and late layers identically. However, evidence suggests that later layers do less heavy lifting and mostly refine what earlier layers have already figured out, making this uniform resource distribution highly inefficient.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/69f5b884-0798-4415-80b9-f73080a08c48/2026-06-30_000787.webp?t=1782836818"/><div class="image__source"><span class="image__source_text"><p>Tapering MLP width improves perplexity at no additional parameter or compute cost</p></span></div></div><p class="paragraph" style="text-align:left;">To address this mismatch, researchers introduced &quot;Tapered Language Models,&quot; an architectural principle that gradually reduces processing capacity across the depth of the model. Instead of keeping the internal width of the model&#39;s primary processing units constant, they used a smooth cosine schedule to front-load capacity in the early layers and taper it down toward the end.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1146f94a-7442-45b8-8072-f01e6151804f/2026-06-30_000788.webp?t=1782836839"/><div class="image__source"><span class="image__source_text"><p>Front-loading MLP capacity improves perplexity</p></span></div></div><p class="paragraph" style="text-align:left;">Under a strictly fixed parameter budget, this design ensures that the model concentrates its resources where they are needed most, achieving better results without adding any extra training or inference computing costs.</p><p class="paragraph" style="text-align:left;">The researchers found that this tapering method consistently improves text prediction and reasoning performance across various model sizes and different foundational AI architectures.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4c49a8f8-95ac-4b0f-8a32-fcdf6772b3f3/2026-06-30_000789.webp?t=1782836871"/><div class="image__source"><span class="image__source_text"><p>Layer updates become more aligned with the residual stream at greater depths. </p></span></div></div><p class="paragraph" style="text-align:left;">By analyzing how information flows through the network, they discovered that earlier layers write highly novel features into the model&#39;s memory stream, while later layers yield diminishing returns by reinforcing existing data.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.23670?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="improved-large-language-diffusion-m">Improved Large Language Diffusion Models</h2><p class="paragraph" style="text-align:left;"><i>Nie et al. [</i>Gaoling School of Artificial Intelligence, Beijing Key Laboratory of Research on Large Models and Intelligent Governance, Engineering Research Center of Next-Generation Intelligent Search and Recommendation, ByteDance Seed<i>]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 574 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Diffusion LMs </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Currently, most LLMs generate text strictly one word at a time from left to right, which limits their ability to plan ahead. While bidirectional &quot;diffusion&quot; models can look in both directions at once to solve this, they have historically struggled to match the performance of their traditional, step-by-step counterparts.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c462815b-184d-4f24-8cd0-a446308d7230/2026-06-30_000786.webp?t=1782836346"/><div class="image__source"><span class="image__source_text"><p>Benchmark Results of Base Models.</p></span></div></div><p class="paragraph" style="text-align:left;">To bridge this gap, researchers developed iLLaDA, an eight-billion-parameter model trained entirely from scratch using fully bidirectional attention. Instead of predicting the very next word, this model learns by <b>predicting missing words anywhere in a sequence</b>, using context from both the left and the right simultaneously.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8c3fa521-c16c-4ba6-b9eb-34dabb00e137/2026-06-30_000785.webp?t=1782836324"/><div class="image__source"><span class="image__source_text"><p>Benchmark Results of Instruct Models.</p></span></div></div><p class="paragraph" style="text-align:left;">By scaling its pre-training to twelve trillion tokens and fine-tuning the system on a twenty-five-billion-token instruction dataset across twelve epochs, the researchers demonstrated that this bidirectional approach can be highly competitive. They also optimized the model&#39;s efficiency by using grouped-query attention to manage memory and tying internal parameters together to keep the model compact.</p><p class="paragraph" style="text-align:left;">While there is still room for improvement on complex instruction-following tasks compared to some of the most advanced traditional models.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.25331?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="autodata-an-agentic-data-scientist-">Autodata: An agentic data scientist to create high quality synthetic data</h2><p class="paragraph" style="text-align:left;"><i>Kulikov et al. [FAIR at Meta]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 466 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Training Data </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Finding a way to automatedly create rich training data could unlock the next generation of capable models. However, existing automated data creation methods often generate examples that are either too easy or too difficult.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a8a11d94-b4d3-469d-b66d-20042d05723b/2026-06-30_000783.webp?t=1782836086"/><div class="image__source"><span class="image__source_text"><p>Autodata creation of CS research questions</p></span></div></div><p class="paragraph" style="text-align:left;">To address this, researchers introduced a framework called Autodata, which trains an AI agent to act like a data scientist. Instead of just churning out data based on a static prompt, this agent enters a creative cycle: it generates training tasks, evaluates how well different AI models perform on them, analyzes the results, and refines its approach to build better data in the next round.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4d513418-0061-417c-8c91-683c4ba6dacb/2026-06-30_000784.webp?t=1782836115"/><div class="image__source"><span class="image__source_text"><p>Meta-optimization of the data scientist agent</p></span></div></div><p class="paragraph" style="text-align:left;">This paper introduces Agentic Self-Instruct system which uses a team of virtual subagents, including a &quot;challenger&quot; to write questions, a &quot;judge&quot; to score them, and both &quot;weak&quot; and &quot;strong&quot; model versions to test them. </p><p class="paragraph" style="text-align:left;">The results across diverse areas like computer science research, law, and scientific reasoning are highly encouraging. Models trained on this carefully calibrated data showed substantial performance gains, proving that this agentic loop creates much more robust training signals than traditional automated methods.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.25996?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">You Don&#39;t Need to Run Every Eval</h2><p class="paragraph" style="text-align:left;"><i>Zeng and Papailiopoulos [Harvard University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Benchmarks </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Evaluating new AI models has become incredibly expensive and time-consuming, often costing thousands of dollars per run. Currently, researchers run dozens of independent benchmarks to track an AI&#39;s progress and compare different design choices, which creates a massive computational bottleneck.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/08205dfd-5953-4de1-b27a-670cfce6c230/hero.png?t=1782835843"/></div><p class="paragraph" style="text-align:left;">In this paper, researchers analyzed a large matrix of 84 frontier models evaluated on 133 different benchmarks. They discovered something surprising: this complex landscape is approximately &quot;rank-2,&quot; meaning an AI’s scores across all these diverse tests are largely determined by just two underlying factors.</p><p class="paragraph" style="text-align:left;">The team built a tool called BENCHPRESS. Instead of running over a hundred tests, a developer can run just five key &quot;probe&quot; benchmarks (such as tests focusing on graduate-level reasoning, coding, and general knowledge) and the system can reconstruct the remaining scores with remarkable accuracy. Even when using a more budget-friendly set of five cheaper benchmarks, the tool still estimates the missing scores to within a few percentage points of their true values.</p><div class="embed"><a class="embed__url" href="https://microsoft.github.io/benchpress/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method" target="_blank"><div class="embed__content"><p class="embed__title"> BenchPress · Predict LLM Scores </p><p class="embed__link"> microsoft.github.io/benchpress </p></div></a></div><p class="paragraph" style="text-align:left;">To ensure these shortcuts are safe to use, the researchers also created a reliability layer. This companion system evaluates how much different mathematical models disagree on a prediction, alongside how much similar data is already known about related models and benchmarks.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.24020?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="d-spark-confidence-scheduled-specul">DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation</h2><p class="paragraph" style="text-align:left;"><i>Cheng et al. [Peking University, DeepSeek-AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.8k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Decoders </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">LLMs are slow and their word-by-word generation mechanism is a bottleneck for real-time applications like conversational assistants. Speculative decoding speeds this up by letting a small &quot;drafter&quot; model guess upcoming words for a large model to verify, but current drafters either guess too slowly or generate disconnected, incoherent word sequences. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ec84d61e-0c46-4775-8f92-c3a4b2444cf7/2026-06-30_000778.webp?t=1782835055"/><div class="image__source"><span class="image__source_text"><p>The DSpark architecture and decoding cycle</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed DSpark, a framework that introduces a clever &quot;semi-autoregressive&quot; architecture. DSpark first uses a highly efficient parallel backbone to generate a block of word guesses all at once, keeping the drafting process incredibly fast. To prevent these parallel guesses from becoming disconnected and error-prone, DSpark passes them through a lightweight sequential module. This extra step injects local transition information, to make sure the words flow naturally together and reducing the likelihood of awkward phrasing that the larger model would ultimately have to reject.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/712c2afa-efc6-42b1-b213-426f9a2690fe/2026-06-30_000779.webp?t=1782835080"/></div><p class="paragraph" style="text-align:left;">In addition to smarter drafting, DSpark introduces confidence-scheduled verification to optimize system efficiency. The framework provides a confidence head that evaluates the survival probability of each guessed word in a sequence.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/798a6433-b8e7-42ea-8e69-fb3424ad24b8/2026-06-30_000782.webp?t=1782835115"/><div class="image__source"><span class="image__source_text"><p>Load-adaptive throughput and verification budgets</p></span></div></div><p class="paragraph" style="text-align:left;">In live tests, this balanced approach accelerated generation speeds for individual users by 60% to 85% compared to established baselines, showing that smarter scheduling can make advanced AI systems significantly more practical for real-world deployment.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-just-dropped-a-new-speculative-decoding-method"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/HAyp4uRnzDk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=cfb2fb44-00e1-4d74-a9ff-eb4ec2c96b13&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>What even is a &gt;&lt; former (yes &gt;&lt; former)</title>
  <description>plus more about Looped World Models, Fixed-Point Reasoners, and ExpRL</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bfd4e929-4370-4337-aa66-2810fdb03a0d/issue_113.jpg" length="113869" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/what-even-is-a-former-yes-former</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/what-even-is-a-former-yes-former</guid>
  <pubDate>Tue, 23 Jun 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-06-23T19:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>June 17th ~ June 23rd</i><br><i>#113 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 12k</span></span> Z.ai has announced the release of <a class="link" href="https://z.ai/blog/glm-5.2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">GLM-5.2</a>, an open-weights model under the MIT license with 1-million-token context window and updates in coding, reasoning, and agentic tasks. It has two reasoning effort levels [GLM-5.2 (max) and GLM-5.2 (high)] designed to help developers balance task performance with token efficiency. You can try it on <a class="link" href="https://huggingface.co/zai-org/GLM-5.2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2dbfbe97-8270-4df2-a504-00c4ee5dffd2/20260617-013338.png?t=1782228365"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1k</span></span> <a class="link" href="https://poolside.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Poolside</a> has released the base and post-trained weights for Laguna M.1, a model featuring a 256K context length licensed under Apache 2.0. Alongside the model checkpoints, the team has also released <a class="link" href="https://github.com/poolsideai/pool?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">&quot;pool&quot; (an agent harness)</a> designed to let developers run the model locally as a coding agent. You can try it on <a class="link" href="https://huggingface.co/poolside/Laguna-M.1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 37k</span></span> Sakana AI has announced <a class="link" href="https://sakana.ai/fugu-release/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Sakana Fugu</a>, a multi-agent orchestration system accessible through a single model API. The underlying Fugu Ultra model is trained to recursively call and coordinate various LLMs in an agent pool to manage complex, multi-step technical workflows. You can try it on the <a class="link" href="https://sakana.ai/fugu/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Sakana AI website</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bc08a44b-f348-4818-949a-cdf3928d6afa/fugu_arch_image.png?t=1782228739"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Optimization Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail-bycloud" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d4bbc543-55b2-45f7-bb16-efc2c6a6303e/IAA_promo_asset_bitg.jpg?t=1782236285"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;">We just added a <b>new chapter on Optimization</b>, that goes through the history, the key techniques, and the current state of optimizers that frontier model uses. </p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an exclusive newsletter offer, where you would get 40% off on the yearly plan for our users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail-bycloud"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="looped-world-models">Looped World Models</h2><p class="paragraph" style="text-align:left;"><i>Lu et al. [FaceMind Research Asia]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 473 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> World Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI models that can navigate the physical world are called &quot;world models&quot;. These models allocate the same amount of computing power to every single moment, whether an object is simply sitting still or dynamically colliding with another. Moreover, these models are too slow and computationally expensive to run on smaller, everyday devices. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e1bdbb53-c2a8-4615-a56f-7bed0b6d2727/lwm.png?t=1782226504"/><div class="image__source"><span class="image__source_text"><p>The overall framework of our proposed Looped World Models (LoopWM).</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, this paper introduced Looped World Models, or LoopWM. Instead of using a giant, resource-heavy stack of unique layers, this architecture uses a smaller, shared set of layers and runs them repeatedly in a loop to refine its predictions.</p><p class="paragraph" style="text-align:left;">Because the model reuses its existing components rather than adding new ones, it achieves the predictive quality of much larger networks with a fraction of the parameters.</p><div style="padding:14px 28px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>Method</b></span></p></th><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>Latent dynamics</b></span></p></th><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>Intermediate</b></span><span style="font-family:rival-sans, sans-serif;"><b> decode</b></span></p></th><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>Action injection</b></span></p></th><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>Looped</b></span><span style="font-family:rival-sans, sans-serif;"><b> depth</b></span></p></th></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;">Dreamer (Hafner<span style="font-family:rival-sans, sans-serif;"> et al.</span>, <span style="text-decoration:underline;"><a class="link" href="https://arxiv.org/html/2606.18208v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former#bib.bib4" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(33, 152, 212)">2020</a></span>)</p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">RSSM</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">reward + value at each step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">per step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">–</p></td></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;">MuZero (Schrittwieser<span style="font-family:rival-sans, sans-serif;"> et al.</span>, <span style="text-decoration:underline;"><a class="link" href="https://arxiv.org/html/2606.18208v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former#bib.bib38" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(33, 152, 212)">2020</a></span>)</p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">learned MLP</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">policy + value + reward</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">per step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">–</p></td></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;">PlaNet (Hafner<span style="font-family:rival-sans, sans-serif;"> et al.</span>, <span style="text-decoration:underline;"><a class="link" href="https://arxiv.org/html/2606.18208v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former#bib.bib2" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(33, 152, 212)">2019</a></span>)</p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">RSSM</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">reconstruction at each step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">per step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">–</p></td></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;">ETD (Koishekenov<span style="font-family:rival-sans, sans-serif;"> et al.</span>, <span style="text-decoration:underline;"><a class="link" href="https://arxiv.org/html/2606.18208v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former#bib.bib48" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(33, 152, 212)">2026</a></span>)</p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">looped layers</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">decode only at end</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">– (language)</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">✓</p></td></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;">NE-Dreamer (Bredis<span style="font-family:rival-sans, sans-serif;"> et al.</span>, <span style="text-decoration:underline;"><a class="link" href="https://arxiv.org/html/2606.18208v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former#bib.bib49" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(33, 152, 212)">2026</a></span>)</p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">RSSM</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">embedding alignment</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">per step</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">–</p></td></tr><tr class="bh__table_row"><th class="bh__table_header" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>LoopWM-DD (ours)</b></span></p></th><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">looped transformer</p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>decode only at step K</b></span></p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;"><span style="font-family:rival-sans, sans-serif;"><b>per step in latent</b></span></p></td><td class="bh__table_cell" width="20%"><p class="paragraph" style="text-align:left;">✓</p></td></tr></table></div><p class="paragraph" style="text-align:left;">This looping structure also allows the model to adjust its computing power on the fly. It can run fewer loops during simple, predictable events and automatically allocate more loops to complex scenarios like physical collisions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ed996206-0a77-44ab-85dd-c4a69adfa707/bcmk.png?t=1782226570"/><div class="image__source"><span class="image__source_text"><p>Relative increase over Qwen3.7-max on automatic online performance, compared against baselines.</p></span></div></div><p class="paragraph" style="text-align:left;">Furthermore, the architecture introduces &quot;deferred decoding,&quot; which allows the system to process a sequence of actions entirely in its abstract internal language. Rather than wasting energy rendering the visual details of every intermediate step, it only decodes the final result at the very end. This combination of parameter efficiency and smart resource allocation offers a promising, practical path forward for real-time AI planning.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.18208?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="variable-width-transformers">Variable-Width Transformers</h2><p class="paragraph" style="text-align:left;"><i>Wu et al. [MIT-IBM Watson AI Lab]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 437 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;">Transformers</span></span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When we scale LLMs, we assumes that every layer requires the same computational budget, even though different stages of a model&#39;s processing pipeline perform entirely different roles. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/18e408c3-d4ee-4206-a3d0-4cc8afdcdd25/overview_bowtie.png?t=1782226828"/><div class="image__source"><span class="image__source_text"><p>&gt; &lt;former, where different layers have different widths</p></span></div></div><p class="paragraph" style="text-align:left;">This paper introduces an hourglass-shaped transformer architecture known as the <b>&gt;&lt; former</b>. This design keeps the initial and final layers wide but significantly narrows the layers in the middle. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8a90ba5b-8add-4b89-af22-07eb3b8ebd17/x2.png?t=1782226855"/><div class="image__source"><span class="image__source_text"><p>The effect of the bottleneck layer index and dimension on language modeling loss, parameterized as a ratio to the total number of layers and the base dimension.</p></span></div></div><p class="paragraph" style="text-align:left;">To handle the fluctuating widths without losing information or adding complex math, the researchers implemented a clever, parameter-free residual resizing mechanism. The model maintains a wide global residual stream where narrower middle layers only read from and write to a specific slice. The remaining inactive dimensions are simply carried forward, bypassing the narrow layers entirely so they can be restored later in the network.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9d2401ea-93d0-497d-998a-9c079eba2ae7/x3.png?t=1782226885"/><div class="image__source"><span class="image__source_text"><p>Language modeling loss vs. pre-training FLOPs (left) and average layer size (right). &gt; &lt;former produces lower loss at smaller FLOP and average layer size costs.</p></span></div></div><p class="paragraph" style="text-align:left;">Across various model sizes, the &gt;&lt;former consistently outperformed traditional, uniform-width baselines in language modeling. By optimizing how space is used, the architecture achieved a twenty-two percent reduction in training computations and a fifteen percent reduction in memory and input-output costs for the key-value cache. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.18246?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="fixed-point-reasoners-stable-and-ad">Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers</h2><p class="paragraph" style="text-align:left;"><i>Movahedi et al. [ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center, ETH Zurich, Swiss Institute of Bioinformatics, Université Paris Cité, Liquid AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 363 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Transformers </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When solving a difficult puzzle like a complex maze or a challenging Sudoku, humans naturally spend more time thinking. For AI, scaling its computational effort based on the difficulty of a task is hard.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8c473bed-71bb-4310-91ef-c31232c499a4/x1.png?t=1782227262"/><div class="image__source"><span class="image__source_text"><p>Signal propagation and adaptivity, FPRM vs. TRM:</p></span></div></div><p class="paragraph" style="text-align:left;">Some researchers use looped neural networks, which process information by running it through the same internal layers repeatedly. However, these looped models face a double-edged sword: looping too many times can destabilize the mathematical signals inside the network, and it is notoriously difficult for the model to naturally decide when it has &quot;thought&quot; enough and should stop.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f84d5c20-f096-4746-93a1-79ab0d25fbce/x2.png?t=1782227279"/><div class="image__source"><span class="image__source_text"><p>The blessing and the curse of depth in Looped Transformers. </p></span></div></div><p class="paragraph" style="text-align:left;">To address these challenges, researchers designed a framework called the Fixed-Point Reasoning Model, or FPRM. The team resolved the stability issue that affects deeply looped networks by restructuring how the model balances its internal signals. By switching to a pre-normalization setup paired with residual scaling, they successfully kept the model&#39;s internal data stable and trainable, even when running through many loops.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/08104aac-34a3-496a-b29a-0d20fd5ae8a2/x4.png?t=1782227298"/><div class="image__source"><span class="image__source_text"><p>FPRM architecture</p></span></div></div><p class="paragraph" style="text-align:left;">Most importantly, FPRM introduces a self-contained halting mechanism based on mathematical convergence. The model loops until its internal representations settle into a stable state, known as a fixed point instead of relying on an external, complex module to decide when to stop.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.18206?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">ExpRL: Exploratory RL for LLM Mid-Training</h2><p class="paragraph" style="text-align:left;"><i>Xiang et al. [Stanford University, Carnegie Mellon University, OpenAI,]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 855 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">RL helps AI models solve complex reasoning problems, but training collapses when problems are too difficult because the model rarely stumbles upon the correct final answer to receive any feedback. Without a way to reward partial progress and smart intermediate steps, AI models cannot easily discover the creative strategies needed to solve challenging math and science problems.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/80bde1d5-ac43-499b-b7e4-70d686b297b5/x1.png?t=1782227698"/><div class="image__source"><span class="image__source_text"><p>Exploratory RL (ExpRL)</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed a method called ExpRL, which stands for Exploratory Reinforcement Learning. Instead of forcing the AI to blindly copy human solutions, this approach keeps the correct answer hidden from the AI while it attempts a problem.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/83348bee-0e71-4122-85d0-4b2cf2b225ba/x2.png?t=1782227717"/><div class="image__source"><span class="image__source_text"><p>Pass@k after training with ExpRL on HMMT-Nov-2025 (128 samples).</p></span></div></div><p class="paragraph" style="text-align:left;">An automated judge then uses that hidden answer as a customized grading rubric to evaluate the AI&#39;s step-by-step thinking. By analyzing these drafts, the judge can award points for productive intermediate progress. This turns a simple pass-fail test into a learning experience that guides the AI&#39;s exploration.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/46cf015a-1e9c-4f4a-a55c-09d918289c3d/x3.png?t=1782227736"/><div class="image__source"><span class="image__source_text"><p>ExpRL training dynamics during Stage-I.</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers found that this supportive warm-up phase prepares the AI for subsequent, more rigorous training. Compared to traditional fine-tuning or basic right-or-wrong feedback, models prepared with this method showed a much broader diversity of problem-solving strategies.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4d88f9cb-add5-4b96-a99f-e28777e6fe51/x5.png?t=1782227763"/><div class="image__source"><span class="image__source_text"><p>Behavior changes after RL priming relative to the base model</p></span></div></div><p class="paragraph" style="text-align:left;">The AI naturally began to exhibit helpful thinking habits, such as double-checking its work, self-correcting mistakes mid-calculation, and backtracking when a strategy failed.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.17024?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="how-transparent-is-diffusion-gemma">How Transparent is DiffusionGemma?</h2><p class="paragraph" style="text-align:left;"><i>Engels et al. [Google DeepMind]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 210 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> DIffusionLM </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Understanding <i>how</i> AI arrives at its decisions is vital for ensuring safety, debugging errors, and preventing misuse. While traditional LLMs think out loud step-by-step in plain English, newer text diffusion models like DiffusionGemma perform their reasoning in a continuous mathematical space.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/02e64e43-3052-44a2-97ba-951117246b3e/x1.png?t=1782227961"/><div class="image__source"><span class="image__source_text"><p>A simplified architecture diagram of DiffusionGemma, with the first two denoising steps shown</p></span></div></div><p class="paragraph" style="text-align:left;">Researchers tried to translate these hidden mathematical states back into human-readable concepts. They discovered that the complex vector information flowing between the model&#39;s iterative denoising steps can actually be mapped into a small bottleneck of natural language tokens.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/113765c9-e2b0-48db-8520-e9469a5e237b/x3.png?t=1782227983"/><div class="image__source"><span class="image__source_text"><p>Breakdown of intermediate state top token identities with different restrictions averaged across WildChat prompts</p></span></div></div><p class="paragraph" style="text-align:left;">Restricting the model to just a handful of these mapped words at each step caused almost no drop in its overall performance. By showing that these intermediate states represent interpretable guesses about the final text, the researchers successfully demonstrated that the model’s hidden reasoning depth can be simplified to a level nearly identical to traditional systems.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ba329da2-7739-43dd-8390-87d534ebe784/x5.png?t=1782228007"/><div class="image__source"><span class="image__source_text"><p>DiffusionGemma thinks less than Gemma on all monitorability evaluations. Error bars show 95% confidence intervals of mean number of characters in each model’s chain of thought.</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers watched the model accurately estimate its final response length before choosing its words, retroactively correct earlier mistakes after formulating its reasoning, and even temporarily weigh multiple sentences at once.</p><p class="paragraph" style="text-align:left;">Most importantly, evaluations showed that these complex internal mechanics do not compromise safety; the model remains just as monitorable as traditional architectures. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.20560?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=what-even-is-a-former-yes-former"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/HAyp4uRnzDk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=0e10c4f5-4a5d-4e74-8e24-9ad1c0c35f10&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>MiniMax M3&#39;s New Attention: MiniMax Sparse Attention</title>
  <description>plus more about FlashMemory-DeepSeek-V4, Trajectory-Refined Distillation, Test-Time Gradient Guidance, and End-to-End Context Compression at Scale</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/369f26fe-279b-492b-b66b-0af573dbeae8/issue_112.jpg" length="253063" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/minimax-m3-s-new-attention-minimax-sparse-attention</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/minimax-m3-s-new-attention-minimax-sparse-attention</guid>
  <pubDate>Tue, 16 Jun 2026 18:25:00 +0000</pubDate>
  <atom:published>2026-06-16T18:25:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>June 8th ~ June 16th</i><br><i>#112 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 13k</span></span> Moonshot AI has released <a class="link" href="https://www.kimi.com/code?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">Kimi-K2.7-Code</a>, which is an open-source coding model with improved instruction following, coding, and agent capabilities compared to its predecessor, K2.6. The updated model shows benchmark gains, including a 21.8% increase on the Kimi Code Bench v2, while reducing reasoning-token usage by 30% for greater efficiency. You can <a class="link" href="https://platform.kimi.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">try it on Kimi Platform</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7240f7b6-7c39-4313-99d1-fd5e544693f9/image.png?t=1781626741"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 8.5k</span></span> Z.ai has announced the <a class="link" href="https://docs.z.ai/devpack/latest-model?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">rollout of GLM-5.2</a>, its new flagship model with advanced coding capabilities, a 1-million-token context window, and configurable reasoning modes. The model is currently available to GLM Coding Plan subscribers and is scheduled to be officially open-sourced under the MIT License next week. </p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 88k</span></span> Anthropic has suspended access to its Fable 5 and Mythos 5 models after a US government export control directive citing national security concerns. The restriction impacts all global customers and limits access for foreign nationals. While Anthropic is working to address the regulation, you can explore other open-source models.</p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.4k</span></span> Google has introduced <a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">DiffusionGemma</a>, a new 26B Mixture of Experts open model designed to explore <b>text diffusion techniques</b>. By generating entire 256-token blocks simultaneously, the model achieves up to <b>four times faster</b> inference on GPUs, though its overall output quality is lower than standard Gemma 4 models. You can try it on <a class="link" href="https://console.cloud.google.com/agent-platform/publishers/google/model-garden/diffusiongemma?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">Model Garden</a> or <a class="link" href="https://huggingface.co/collections/google/diffusiongemma?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ea7ffe98-8f68-472e-ba26-4844913c24f4/image.png?t=1781627163"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Optimization Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;">We just added a <b>new chapter on Optimization</b>, that goes through the history, the key techniques, and the current state of optimizers that frontier model uses. </p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an exclusive newsletter offer, where you would get 40% off on the yearly plan for our users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="endto-end-context-compression-at-sc">End-to-End Context Compression at Scale</h2><p class="paragraph" style="text-align:left;"><i>Li et al. [New York University, Modal Labs, University of Maryland, Princeton University</i>, <i>Columbia University, Harvard University, Lawrence Livermore National Laboratory</i>, <i>FAIR at Meta]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 430 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Context </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">There is a fundamental bottleneck in AI: memory. When LLMs try to read massive documents or entire software codebases, the system&#39;s memory footprint and processing time increase exponentially.</p><div class="embed"><a class="embed__url" href="https://github.com/LeonLixyz/LCLM?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank"><div class="embed__content"><p class="embed__title"> GitHub - LeonLixyz/LCLM: latent context language models </p><p class="embed__description"> latent context language models. Contribute to LeonLixyz/LCLM development by creating an account on GitHub. </p><p class="embed__link"> GitHub </p></div><img class="embed__image embed__image--right" src="https://opengraph.githubassets.com/804cdace20e4d4076043f639a370e340536628e57b66ae8ec49a55d85b09fd78/LeonLixyz/LCLM"/></a></div><p class="paragraph" style="text-align:left;">Until now, the main workaround has been to selectively forget information by trimming the system&#39;s internal memory cache. However, this approach is often remarkably slow, computationally unstable, or degrades the model’s intelligence.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7d2c512b-09ad-46b2-83fd-58faaa0da3fe/2026-06-16_000637.webp?t=1781622724"/><div class="image__source"><span class="image__source_text"><p>Examples of the three data types used to train LCLMs</p></span></div></div><p class="paragraph" style="text-align:left;">This paper created Latent Context Language Models, which rethinks how machines ingest information. Instead of forcing the main AI engine to read a sprawling mountain of text word by word, researchers placed a smaller, highly efficient encoder model in front of it. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9a4b0d9d-dd4d-481d-b219-1f8b9e877337/2026-06-16_000638.webp?t=1781622749"/><div class="image__source"><span class="image__source_text"><p>A from-scratch pre-training sweep identifies the best encoder-decoder compressor architecture.</p></span></div></div><p class="paragraph" style="text-align:left;">This encoder acts like a brilliant summarizer. It processes large blocks of text and mathematically compresses them into a much shorter sequence of dense representations called soft tokens. An adapter then translates these soft tokens into a format the main model natively understands. By handling the heavy lifting upfront, the main model processes a fraction of the data.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8f603b7e-89a3-445f-9254-79df3df8d237/2026-06-16_000639.webp?t=1781622777"/><div class="image__source"><span class="image__source_text"><p>LCLMs can use tools to retrieve compressed context and improve exact string-match accuracy. </p></span></div></div><p class="paragraph" style="text-align:left;">This breakthrough shifts the limits of what models can handle. The researchers found that this architecture beautifully reduces peak memory usage and the time it takes to generate an answer, all while preserving the system&#39;s baseline intelligence.</p><div class="embed"><a class="embed__url" href="https://huggingface.co/latent-context?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention" target="_blank"><div class="embed__content"><p class="embed__title"> latent-context (Latent Context Language Model) </p><p class="embed__description"> Org profile for Latent Context Language Model on Hugging Face, the AI community building the future. </p><p class="embed__link"> huggingface.co/latent-context </p></div><img class="embed__image embed__image--right" src="https://cdn-thumbnails.huggingface.co/social-thumbnails/latent-context.png"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.09659?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="test-time-gradient-guidance-of-flow">Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning</h2><p class="paragraph" style="text-align:left;"><i>Zhou et al. [UC Berkeley, Physical Intelligence]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> RL </span></span></p><p class="paragraph" style="text-align:left;">Researchers have been struggling to scale up reinforcement learning because of messy training dynamic where an &quot;actor&quot; system learns to make decisions while a &quot;critic&quot; system simultaneously learns to judge them. Training these two together is highly unstable, especially when using advanced generative models that create complex actions step-by-step.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f9440c55-4aa7-4f7f-8ec8-9865215b5542/2026-06-16_000640.webp?t=1781623653"/><div class="image__source"><span class="image__source_text"><p>Illustrative example of 1D denoising process mapping Gaussian noise to a tri-modal distribution</p></span></div></div><p class="paragraph" style="text-align:left;">If the critic model updates, the actor gets confused, which makes it incredibly difficult to build larger, smarter robotic control systems. Researchers wanted to know if they could skip this chaotic paired training altogether.</p><p class="paragraph" style="text-align:left;">The researchers discovered a remarkably elegant workaround called Q-Guided Flow. Instead of forcing the actor and critic to learn together, they trained them separately. When the AI is actually running (a phase called test time) it generates actions through a gradual process of removing noise. Normally, asking the critic for directions during these noisy, half-finished steps causes computational confusion or requires massively expensive backward math.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/52393a73-cf93-42a2-aa20-0c9d732c5cc1/2026-06-16_000641.webp?t=1781623676"/></div><p class="paragraph" style="text-align:left;">However, the researchers found a brilliant shortcut. At each step, the system makes a rapid mathematical guess of what the final, perfectly clean action will look like. It shows this clean guess to the critic, receives reliable feedback, and uses that insight to gently steer the ongoing action toward a higher-value outcome.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aef6852d-ec27-4939-963c-15423784f140/2026-06-16_000642.webp?t=1781623695"/><div class="image__source"><span class="image__source_text"><p>Offline RL performance at 500k training steps (20 tasks, 10 seeds)</p></span></div></div><p class="paragraph" style="text-align:left;">By applying this simple guidance exactly when the AI is acting, the resulting models are faster, cheaper to run, and smoothly scale up to handle much harder tasks without breaking.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.11087?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="mini-max-sparse-attention">MiniMax Sparse Attention</h2><p class="paragraph" style="text-align:left;"><i>Lai et al. [MiniMax, Peking University, NVIDIA, Zhejiang University, Huazhong University of Science and Technology, Nanjing University, Hangzhou Dianzi University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 511 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b02a0078-db88-4cd7-87fb-b666959107a5/2026-06-16_000643.webp?t=1781624403"/><div class="image__source"><span class="image__source_text"><p>Overview of MSA</p></span></div></div><p class="paragraph" style="text-align:left;">This paper introduces a new technique called MiniMax Sparse Attention, in this Instead of forcing the AI to expend heavy processing power analyzing every single word in its vast memory simultaneously, researchers designed a highly efficient filtering system. They built a lightweight indexer that quickly scans the data in blocks.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5bf1d2cb-8d9a-4453-8f06-adc8e664f370/2026-06-16_000644.webp?t=1781624428"/></div><p class="paragraph" style="text-align:left;">It identifies only the most important information relevant to the current task (while always keeping the most recent context active to maintain stability) and skips the rest. Once these blocks are selected, the model’s main engine focuses its full attention exclusively on them. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a9782225-5b80-43ea-aad8-798de3b4d6ba/2026-06-16_000645.webp?t=1781624453"/><div class="image__source"><span class="image__source_text"><p>Efficiency comparison between GQA and MSA under the shared experimental model configuration.</p></span></div></div><p class="paragraph" style="text-align:left;">When researchers tested this on a massive, 109-billion parameter model, this method maintained the high quality of traditional approaches but drastically slashed the workload. The system absorbs initial data <b>fourteen times faster</b> and generates new responses more than <b>seven times faster</b>.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.13392?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Trajectory-Refined Distillation</h2><p class="paragraph" style="text-align:left;"><i>Jiang et al. [McGill University, Mila Quebec AI Institute, UT Austin]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 360 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Distillation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Researchers are building smarter AI models by pairing a smaller &quot;student&quot; model with a highly advanced &quot;teacher&quot; model. The student practices solving a problem step-by-step, and the teacher grades its output word-by-word. However, researchers recently identified a major structural roadblock in this process called &quot;prefix failure.&quot;</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ea2aeecd-dbeb-45fe-9dea-aa99e6a7e4c8/2026-06-16_000646.webp?t=1781625089"/><div class="image__source"><span class="image__source_text"><p>TRD refines student-generated trajectories yo into improved trajectories yr, which are then used for distillation.</p></span></div></div><p class="paragraph" style="text-align:left;">Let’s imagine a student taking a complex math test and making a logical error in the very first step. Because AI models generate text sequentially, once the student goes down this wrong path, the rest of the answer is doomed. When this happens, the teacher model gets mathematically confused, awkwardly trying to correct the student word-by-word while still trapped in the student&#39;s flawed train of thought.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/37716722-ebda-4c5d-bf14-aaf48ad15bb1/2026-06-16_000647.webp?t=1781625114"/><div class="image__source"><span class="image__source_text"><p>Under prefix failure, the teacher distribution becomes a mixture with two modes.</p></span></div></div><p class="paragraph" style="text-align:left;">Until now, fixes merely involved ignoring or adjusting the penalties for these individual bad words, completely failing to address the reality that the underlying logic was already hopelessly broken.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7f00a044-f227-49b1-bd4e-7f3013deae50/2026-06-16_000648.webp?t=1781625139"/><div class="image__source"><span class="image__source_text"><p>OPD Avg@16 results (%) using Qwen3-8B as the teacher</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers introduced a highly promising approach called Trajectory-Refined Distillation. Rather than relying on word-level nitpicking, this method steps back to look at the big picture. Instead of grading a bad thought process, the system takes the student&#39;s initial attempt and allows the teacher to gently revise the entire reasoning path into a cohesive, corrected draft before the actual learning occurs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c4bcde0a-f4e4-4366-8929-718a48bcc6c1/2026-06-16_000649.webp?t=1781625162"/><div class="image__source"><span class="image__source_text"><p>Trajectory analysis</p></span></div></div><p class="paragraph" style="text-align:left;">By correcting foundational missteps at their source, the student model is exposed to completely new, valid ways to reason through complex problems.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.08432?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="flash-memory-deep-seek-v-4-lightnin">FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention</h2><p class="paragraph" style="text-align:left;"><i>Wang et al. [Independent Researchers, Tencent, The Hong Kong University of Science and Technology (Guangzhou)</i>, <i>Tsinghua University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 219 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">To generate a single word of response, LLMs must keep that entire massive history actively loaded in its working memory. Researchers realized this creates a severe hardware bottleneck. They discovered a striking inefficiency: most of the time, models processing massive contexts only need the most recent sliver of information to form a response.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a2748bc-b8ae-4328-8a21-f5b1aad065a6/2026-06-16_000651.webp?t=1781625380"/><div class="image__source"><span class="image__source_text"><p>Architectural overview of LSA vs. CSA.</p></span></div></div><p class="paragraph" style="text-align:left;">This paper introduces a new framework called Lookahead Sparse Attention. Instead of passively forcing the AI to carry the full weight of its history, they introduced a &quot;Neural Memory Indexer.&quot; Every few steps, this indexer evaluates the AI&#39;s current thought process, dynamically predicting and fetching only the critical historical chunks needed for the immediate future.</p><p class="paragraph" style="text-align:left;">The researchers managed to train this lightweight indexer entirely independently from the massive core model, bypassing tremendous computational costs and allowing it to be optimized on its own.</p><p class="paragraph" style="text-align:left;">By loading only what is strictly necessary, the system shrinks the active memory footprint down to just 13.5 percent of traditional models. At extreme context scales of half a million tokens, memory overhead drops by over ninety percent. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.09079?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=minimax-m3-s-new-attention-minimax-sparse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/gC76aeibdFA" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=7ed16a2b-2c7f-4fc4-90f7-c838f8074e70&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Microsoft just shared the frontier data engineering secrets</title>
  <description>plus more about If LLMs Have Human-Like Attributes, Then So Does Age of Empires II, Cosmos 3, and Robots Need More than VLA and World Models</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a9356eb1-3cbe-4025-8f8f-1a79d329b0fe/issue_111.jpg" length="168761" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/microsoft-just-shared-the-frontier-data-engineering-secrets</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/microsoft-just-shared-the-frontier-data-engineering-secrets</guid>
  <pubDate>Tue, 09 Jun 2026 20:00:00 +0000</pubDate>
  <atom:published>2026-06-09T20:00:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>June 3rd ~ June 9th</i><br><i>#111 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 12k</span></span> Google has introduced <a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Gemma 4 12B</a>, an encoder-free multimodal model with Apache 2.0 license that is compact enough to run locally on just 16GB of VRAM. By using a lightweight 35M-parameter vision embedding module rather than a traditional encoder, it delivers advanced, multi-step reasoning capabilities with a significantly reduced memory footprint. You can try it out now on <a class="link" href="https://ollama.com/library/gemma4?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(26, 115, 232)">Ollama</a><span style="background-color:rgb(255, 255, 255);">, </span><a class="link" href="https://developers.google.com/edge/gallery?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(26, 115, 232)">Google AI Edge Gallery App</a><span style="background-color:rgb(255, 255, 255);">, the </span><a class="link" href="https://ai.google.dev/edge/eloquent?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(26, 115, 232)">Google AI Edge Eloquent</a><span style="background-color:rgb(255, 255, 255);"> app </span>or <a class="link" href="https://huggingface.co/collections/google/gemma-4?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c3096be8-21ac-4cc6-8435-1fa461a7b871/image.png?t=1781023607"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.9k</span></span> Alibaba has introduced <a class="link" href="https://qwen.ai/blog?id=qwen3.7-plus&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Qwen3.7-Plus</a>, a new model that delivers text performance which matches Max-tier models across various benchmarks. The release also brings systematic enhancements to its visual understanding capabilities, which are specifically tailored for powering complex multimodal interactive hybrid and browser agents. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f7e2b15e-2184-4db7-9cda-753470084bc8/image.png?t=1781023856"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.4k</span></span> NVIDIA has released <a class="link" href="https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Nemotron 3 Ultra</a>, a new model optimized for complex agentic tasks such as long-horizon planning, large-scale code analysis, and extensive tool calling. The release provides developers with access to the model weights, synthetic data, and post-training recipes tailored for popular agent frameworks. You can try it on <a class="link" href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8b093fd2-a85b-41da-8f65-895caf34509e/Cost-Completion.webp?t=1781024001"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.1k</span></span> Cognition has announced an &quot;<a class="link" href="https://cognition.ai/blog/ai-guarantee?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">AI Productivity Guarantee</a>&quot; for its AI agent Devin. They are pledging to fund customer usage up to $10 million if the tool fails to deliver proportional engineering value. To support this initiative, the company detailed a newly developed measurement system that estimates an <a class="link" href="https://cognition.ai/blog/ai-productivity?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">agent&#39;s productive output</a> by comparing it to the time a human engineer would require to complete the same task within an enterprise codebase. <a class="link" href="https://devin.ai/guarantee?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Read more about gurantee</a>.</p></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Optimization Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;">We just added a <b>new chapter on Optimization</b>, that goes through the history, the key techniques, and the current state of optimizers that frontier model uses. </p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an exclusive newsletter offer, where you would get 40% off on the yearly plan for our users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="mai-thinking-1-building-a-hill-clim">MAI-Thinking-1: Building a Hill-Climbing Machine</h2><p class="paragraph" style="text-align:left;"><i>The Microsoft AI Team</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 3.8k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM release </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Researchers are trying to build AI models, but just building a single impressive model isn’t the end goal. We need to figure out how to continually and reliably improve them. We need to create a system that genuinely learns to reason from the ground up, rather than taking the common shortcut of imitating older AIs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d9a70844-8e8b-4855-a51e-54fee9befc46/2026-06-09_000549.webp?t=1781020947"/><div class="image__source"><span class="image__source_text"><p>Overview of the MAI-Base-1 architecture</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers designed what they call a &quot;hill-climbing machine.&quot; This is a fully integrated system where data, hardware, and safety tests work together in a continuous loop, supporting a steady, measurable climb toward better reasoning.</p><p class="paragraph" style="text-align:left;">The first milestone from this new approach is a powerful model called MAI-Thinking-1. Researchers trained it entirely on a pristine collection of human knowledge, including public code, academic papers, and books, while carefully removing synthetic, AI-generated content. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cfab5304-ee06-409f-833f-edc20002e690/2026-06-09_000550.webp?t=1781020975"/><div class="image__source"><span class="image__source_text"><p>Pipeline for processing HTML pre-training data.</p></span></div></div><p class="paragraph" style="text-align:left;">To do this, the model undergoes a rigorous reinforcement learning climb. During this phase, it learns how to connect chains of thought to solve complex problems, use external tools, and carefully balance the desire to be helpful with the strict necessity of staying safe.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3738b7b6-4759-4bfe-a1f5-a630917f3cf7/2026-06-09_000552.webp?t=1781021007"/><div class="image__source"><span class="image__source_text"><p>Rank non-invariance in data mixture scaling.</p></span></div></div><p class="paragraph" style="text-align:left;">By relying purely on authentic human data and systematic testing, MAI-Thinking-1 has proven itself to be one of the most capable models of its size for complex math, science, and software engineering.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://microsoft.ai/news/introducing-mai-thinking-1/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="cosmos-3-omnimodal-world-models-for">Cosmos 3: Omnimodal World Models for Physical AI</h2><p class="paragraph" style="text-align:left;"><i>NVIDIA</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.7K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> World Model </span></span></p><p class="paragraph" style="text-align:left;">Let’s imagine a robot that can clear your dining table. To accomplish this seemingly simple task, today&#39;s robots must juggle a disjointed collection of artificial intelligence models: one to see the dishes, another to plan movements, and another one to predict what happens if a glass gets bumped.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7e91c88a-4fd0-4ed2-9a4e-f86adc2b1622/2026-06-09_000553.webp?t=1781021307"/><div class="image__source"><span class="image__source_text"><p>Cosmos 3 serves as a general-purpose backbone for Physical AI. </p></span></div></div><p class="paragraph" style="text-align:left;">Training these physical agents directly in the real world is slow, expensive, and sometimes dangerous. We need to create a single, unified system that handles seeing, reasoning, simulating, and acting.</p><p class="paragraph" style="text-align:left;">To solve this, this paper introduced Cosmos 3, which is a versatile model that unites language, video, audio, and physical actions into one continuous stream of thought. Instead of treating sight, sound, and movement as separate programs, this architecture translates all these senses into a shared digital vocabulary.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d2833cc4-159e-4123-aa0b-4384ad4432e8/2026-06-09_000554.webp?t=1781021330"/><div class="image__source"><span class="image__source_text"><p>Cosmos 3 offers a strong starting point for training Physical AI agents.</p></span></div></div><p class="paragraph" style="text-align:left;">It is built using a dual-pathway structure. One side focuses on deeply understanding the current environment, analyzing the scene much like a person reads a book. The other side acts as an imagination engine, generating plausible futures. When the AI considers a physical action, like grasping a cup or turning a steering wheel, the system instantly visualizes exactly how the physical world will react.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8858d77b-3455-47b5-9222-6ea6b93b10a0/2026-06-09_000555.webp?t=1781021354"/><div class="image__source"><span class="image__source_text"><p>Unified action representation</p></span></div></div><p class="paragraph" style="text-align:left;">This unified mind allows developers to generate rich, interactive simulations to safely train future robots. It proves that when AI can cohesively see, hear, think, and act, bringing helpful autonomous agents into our daily lives becomes a beautiful, tangible reality.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.02800?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="self-trained-verification-for-train">Self-Trained Verification for Training- and Test-Time Self-Improvement</h2><p class="paragraph" style="text-align:left;"><i>Wu and Raghunathan [Carnegie Mellon University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 424 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Self-training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">To master advanced reasoning, the AI models need to learn how to catch their own mistakes. We use a system where an AI proposes an answer and an internal &quot;verifier&quot; provides feedback. This mechanism works to some extent, but the verifier cannot spot hidden logical flaws, the system just reinforces its own errors. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a0f9979a-ac6d-492f-b31a-488f778ecbb1/2026-06-09_000556.webp?t=1781022124"/><div class="image__source"><span class="image__source_text"><p>Isolating the contribution of trained feedback using ground-truth verdict.</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, the researchers used a simple logic: diagnosing a mistake is much easier when you already hold the answer key. They developed a new method called Self-Trained Verification.</p><p class="paragraph" style="text-align:left;">First, they let a &quot;teacher&quot; version of the verifier look at both the AI&#39;s flawed attempt and the correct reference solution. With the right answer in hand, the teacher easily identified exactly where the logic failed. Then, they trained the standard verifier to imitate this insightful feedback without ever peeking at the answer key. By matching the teacher&#39;s deep critique, the standard verifier learned to catch subtle errors entirely on its own.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/71f2e91f-e74c-4a3d-9dfa-b974f08d0aac/2026-06-09_000557.webp?t=1781022152"/><div class="image__source"><span class="image__source_text"><p>Overview of self-trained verification.</p></span></div></div><p class="paragraph" style="text-align:left;">When using this self-trained verifier on exceptionally hard math problems, researchers saw accuracy increased dramatically. On complex scientific reasoning tasks, success rates rocketed from just over one percent to twenty-one percent.</p><div class="embed"><a class="embed__url" href="https://ar-forum.github.io/stv-webpage/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets" target="_blank"><img class="embed__image embed__image--top" src="https://ar-forum.github.io/stv-webpage/figures/teaser.png"/><div class="embed__content"><p class="embed__title"> Self-Trained Verification for Training- and Test-Time Self-Improvement </p><p class="embed__description"> A reference-conditioned teacher distills feedback into the verifier. Doubles accuracy on hard math, 14x on the hardest scientific reasoning, and lifts standalone pass@1 30% past where RL had converged. </p><p class="embed__link"> ar-forum.github.io/stv-webpage </p></div></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.30290?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">If LLMs Have Human-Like Attributes, Then So Does Age of Empires II</h2><p class="paragraph" style="text-align:left;"><i>Wynter [Microsoft & The University of York]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 11K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM SuperIntelligence </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We want to achieve superintelligence and build AI which possess generalized, intrinsic anthropomorphic attributes (e.g., empathy, self-awareness, anxiety, morality). When researchers try to check if a &quot;mind&quot; exists inside the machine, it leads to flawed, circular, or scientifically uninformative conclusions.</p><p class="paragraph" style="text-align:left;">To prove their point, the authors take a highly unconventional approach: they demonstrate that a neural network can be trained inside the 1999 video game <i><b>Age of Empires II</b></i><b> (AoE II)</b>.</p><p class="paragraph" style="text-align:left;">Let’s think of a fully functioning LLM built entirely out of AoE II goats running across a map. If you input &quot;I feel lonely,&quot; the goats will shuffle around and eventually output the text: &quot;I feel bad for you, maybe catch up with a friend.&quot;</p><ul><li><p class="paragraph" style="text-align:left;">Because you are watching digital goats do math, you would <i>never</i> assume the goats are actually experiencing empathy or understanding.</p></li><li><p class="paragraph" style="text-align:left;">However, when a standard LLM outputs the exact same text via a sleek chat interface, users and researchers readily ascribe human emotions to it.</p></li><li><p class="paragraph" style="text-align:left;">This proves that <b>anthropomorphism is an illusion driven by presentation/substrate</b>, not an intrinsic property of the model&#39;s mathematics.</p></li></ul><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1b6f9456-6b46-4f9a-b806-10f4ad088a4f/2026-06-09_000558.webp?t=1781022590"/><div class="image__source"><span class="image__source_text"><p>Ansatz-based training algorithm for our 1-bit perceptron, as a circuit (top) and as an AoE II implementation.</p></span></div></div><p class="paragraph" style="text-align:left;">The authors argue that standard scientific methods fail when researchers assume a model possesses human traits beforehand (the &quot;Accept/Reject Computational Theory of Mind&quot; framework):</p><ul><li><p class="paragraph" style="text-align:left;"><b>Positive Results:</b> If you assume an LLM has empathy and design a test to measure it, a positive result just circularly confirms your underlying assumption.</p></li><li><p class="paragraph" style="text-align:left;"><b>Negative Results:</b> If the LLM fails the test, the result is completely ambiguous. You don&#39;t know if the LLM lacks empathy, or if your test was just poorly designed.</p></li></ul><p class="paragraph" style="text-align:left;">The authors suggest abandoning the debate over whether LLMs have &quot;minds&quot; when conducting empirical research. Instead, researchers should adopt a <b>&quot;null assumption.&quot;</b> This means studying an LLM&#39;s outputs purely causally and behaviorally rather than jumping to ascribe anthropomorphic intent (a &quot;Geist&quot;) to the system.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.31514?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="robots-need-more-than-vla-and-world">Robots Need More than VLA and World Models</h2><p class="paragraph" style="text-align:left;"><i>Karcini et al. [Motoniq.ai, Stanford University, Istituto Italiano di Tecnologia, ETH Zurich, Technical University of Darmstadt, UCL Centre for AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 582 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Word Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We have seen text-based AI models trained on the entire internet, but general purpose robots have been stuck learning the hard way. Until now, researchers relied almost entirely on collecting explicitly labeled robot data.</p><p class="paragraph" style="text-align:left;">Researchers need to physically guide a robot arm through a task thousands of times just to teach it one simple chore. This process is expensive, slow, and carries the constant risk of hardware damage. Researchers realized that if we want robots to achieve the leaps we have seen in language AI, simply building larger neural networks is an incomplete solution.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1361103c-82dd-4805-9447-f3062e79bbb8/2026-06-09_000559.webp?t=1781023186"/><div class="image__source"><span class="image__source_text"><p>Next generation robotics will come from advances that go well beyond scaling vision language action (VLA) models.</p></span></div></div><p class="paragraph" style="text-align:left;">The internet has a lot of unstructured behavioral data, like countless internet videos of people interacting with everyday objects. These videos contain rich information about how things move, how forces work, and what a successful task looks like. However, a robot cannot directly use this information because it does not know which of its specific motors to turn to replicate what a human hand does.</p><p class="paragraph" style="text-align:left;">To solve this problem, the researchers suggest building a new physical intelligence stack to translate this abundant data. This system can automatically label unstructured behaviors, map human motions onto different robot bodies, predict physical consequences using 3D world models, and allow robots to understand success just by watching.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2606.06556?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=microsoft-just-shared-the-frontier-data-engineering-secrets"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/AIRfT41A89s" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=032405ae-c741-40b1-9748-910a385e8360&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>DiffusionBlocks: Save 2-3x Training Memory!?</title>
  <description>plus more about Bitter Lesson in Data Filtering, Do Language Models Need Sleep, and Neural Weight Norm.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b66c761c-2780-4a57-b962-132963af4ddf/issue_110.jpg" length="157950" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/diffusionblocks-save-2-3x-training-memory</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/diffusionblocks-save-2-3x-training-memory</guid>
  <pubDate>Tue, 02 Jun 2026 19:15:00 +0000</pubDate>
  <atom:published>2026-06-02T19:15:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>May 26th ~ June 2nd</i><br><i>#110 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.5k</span></span> StepFun has released <a class="link" href="https://static.stepfun.com/blog/step-3.7-flash/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Step 3.7 Flash</a>, a 198B sparse Mixture-of-Experts model designed to optimize agentic, coding, and multimodal workflows with approximately 11B active parameters. The model offers 256K context support, high-performance tool use, and optimized speed, making it capable of running locally on specialized hardware. You can try it on <a class="link" href="https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/stepfun-ai/Step-3.7-Flash?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">HuggingFace</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4a6d48b1-7590-4b4b-9598-9244974e3042/image.png?t=1780419575"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.7k</span></span> Liquid AI has released <a class="link" href="https://www.liquid.ai/blog/lfm2-5-8b-a1b?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">LFM2.5-8B-A1B</a>, a hybrid Mixture-of-Experts model that significantly upgrades training data to 38T tokens and expands context length to 128k. Designed for agentic workflows, the model enhances instruction following and tool-use capabilities while maintaining efficient local performance across diverse hardware. You can try it on <a class="link" href="https://playground.liquid.ai/chat?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Liquid playground</a> or <a class="link" href="https://huggingface.co/LiquidAI/LFM2.5-8B-A1B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ac707284-595b-4655-8260-27b19e4a44fa/6a1788bd6e860fb6f2bc1740_lfm2_5_8b_a1b_gpu_inference.png?t=1780419686"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 67k</span></span> Anthropic has released Claude Opus 4.8, which introduces improved judgment, enhanced self-assessment capabilities, and a &quot;fast mode&quot; that offers 2.5x speed at a lower price point. The update also brings <a class="link" href="https://claude.com/blog/introducing-dynamic-workflows-in-claude-code?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">dynamic workflows to Claude Code</a>, allowing the model to manage complex, multi-file tasks by deploying parallel subagents.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9f22de67-b857-4e9f-9a19-21cd64aa595b/image.png?t=1780419785"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 529</span></span> Tencent has released <a class="link" href="https://aistudio.tencent.com/llm/zh?tabIndex=0&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Hy-MT2</a>, a new open-source multilingual translation model available in sizes ranging from 1.8B to 30B parameters. The series features impressive efficiency, with the 1.8B version leveraging extreme quantization to run locally on mobile devices while outperforming several mainstream commercial APIs. You can try it on <a class="link" href="https://github.com/Tencent-Hunyuan/Hy-MT2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/collections/tencent/hy-mt2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cd8c597f-7821-41af-9b44-70c23d87c64f/image.png?t=1780419899"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Advanced RL Chapter!</h2><p class="paragraph" style="text-align:left;">My latest project <b>Intuitive AI Academy</b> has the perfect starting point for you! We cover everything from the basics, like transformer architecture, all the way to more advanced topics like LoRA, distillation, Mixture of Experts, and RLHF.</p><p class="paragraph" style="text-align:left;">The goal is simple: <b>make frontier AI systems easy to understand</b> with clear explanations, visuals, interactive learning, and a structured path from fundamentals to cutting-edge techniques.</p><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7441e4a-6ffc-44ed-8934-fd508167a487/image.png?t=1777487919"/></a></div><p class="paragraph" style="text-align:left;">We have just added a new advanced RL chapter, that includes the basics of RL and the current state of RLHF! We currently have an special newsletter offer, where you would get <b>40% off on the yearly plan!</b> </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="a-bitter-lesson-for-data-filtering">DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation</h2><p class="paragraph" style="text-align:left;"><i>Shing et al. [Sakana AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.2k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">A lot of GPU memory is used when training neural networks because standard backprop has to store activations through every layer. This becomes a bottleneck as models get deeper, since memory grows with depth. DiffusionBlocks asks whether we can train only small chunks of the model at a time, without using the usual fragile local objectives that made older block-wise training methods perform badly.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6bbb48de-f920-43b6-a171-bc33e532a86a/image.png?t=1780433720"/></div><p class="paragraph" style="text-align:left;">The core idea is to reinterpret residual layers as steps in a diffusion denoising process. Instead of training the whole transformer end-to-end, the model is split into blocks, and each block is assigned a specific noise range. Each block then learns to denoise within that range independently, so training only needs gradients for one block at a time.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/61d555a9-31b6-431f-8470-c2366be0dd36/image.png?t=1780433745"/></div><p class="paragraph" style="text-align:left;">The research showed that this works across very different architectures, not just toy classification models. On CIFAR-100, DiffusionBlocks got 59.30% accuracy versus 60.25% for standard ViT while only training 4 layers at a time. For image generation, it matched or improved DiT results on CIFAR-10 and ImageNet with around 3× memory reduction. For autoregressive language modeling, it even improved LM1B MAUVE from 0.50 to 0.71 while training only 3 layers at a time.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/afc61412-0c24-4866-8a6d-f8d9b57ac36a/image.png?t=1780433762"/></div><p class="paragraph" style="text-align:left;">The most interesting part is that the blocks are not just random chunks. They use equi-probability partitioning, meaning more capacity is placed around the intermediate noise levels where denoising is hardest, instead of wasting equal space on easy noise regions. This beat uniform partitioning in the ablations. But there is still a tradeoff: moderate block counts like 2 or 3 worked best, while too many blocks hurt quality because each block gets too little capacity.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e2c710c1-e953-4fd9-9cda-cfacc4c4b70a/image.png?t=1780433811"/></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2506.14202?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="a-bitter-lesson-for-data-filtering">A Bitter Lesson for Data Filtering</h2><p class="paragraph" style="text-align:left;"><i>Mohri et al. [Stanford University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.2k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM data curation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We spend a lot of resources to train LLMs, filtering web data to remove &quot;low-quality&quot; or noisy text. This seems intuitive, but filtering throws away a vast majority of the internet&#39;s text. This creates a bottleneck because modern models require trillions of words to keep improving. This paper explores whether expensive, human-designed filters are truly necessary as computational power scales, or if models can learn to navigate the web&#39;s messy reality on their own.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e239647b-11de-4820-b0a4-132a66997018/2026-06-02_000475.webp?t=1780418109"/><div class="image__source"><span class="image__source_text"><p>670M-token CC pool versus junk-injected versions.</p></span></div></div><p class="paragraph" style="text-align:left;">The research showed that with enough computing power and sufficiently large models, the optimal data filter is actually no filter at all. While smaller models struggle with cluttered datasets, larger models trained for longer periods are highly robust. They not only tolerate noisy text but actually benefit from seemingly &quot;poor&quot; data. To test this, researchers injected scrambled documents with completely randomized word orders into the training pool.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7f7a096-8fa3-4d89-be48-38a3135f7928/2026-06-02_000476.webp?t=1780418144"/><div class="image__source"><span class="image__source_text"><p>1B model performance as we vary the pool size; the total needed steps for pool to outperform RefinedWeb grows rapidly. Crossing point as a function of pool size for various model sizes</p></span></div></div><p class="paragraph" style="text-align:left;">Surprisingly, the larger models still succeeded, even benefiting from the shuffled text. Because the original words remained, the models could still learn which terms frequently co-occur, helping them build associations despite the chaotic structure.</p><p class="paragraph" style="text-align:left;">The team concluded that as training budgets scale toward the high-compute limits of the near future, training directly on massive, unfiltered web pools will likely become the most effective strategy. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.19407?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="neural-weight-norm-kolmogorov-compl">Neural Weight Norm = Kolmogorov Complexity</h2><p class="paragraph" style="text-align:left;"><i>Musat [ETH Zürich]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.1K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Complexity </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">It is common to penalize large weights as it helps neural networks generalize to new data, but classical learning theories cannot fully explain why. Because a network trained with weight decay has the same theoretical capacity as one trained without, traditional mathematics struggles to distinguish between a model that genuinely learns and one that merely memorizes noise.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a667f6a3-6d13-46ea-bef6-2f1f651b2437/2026-06-02_000477.webp?t=1780418308"/><div class="image__source"><span class="image__source_text"><p>Comparison of weight-norm-vs-complexity bounds. “Two-sided” indicates whether both directions of a sandwich are proved.</p></span></div></div><p class="paragraph" style="text-align:left;">This paper tries to build a mathematical bridge connecting weight decay to Solomonoff’s universal prior, the theoretically optimal but historically uncomputable method for learning. By analyzing neural networks operating under fixed-precision arithmetic, the researchers proved that minimizing any weight norm is equivalent to finding the shortest computer program that outputs a given result, known as its Kolmogorov complexity.</p><p class="paragraph" style="text-align:left;">They established this through two tight reductions. First, any program can be preloaded directly into network weights at a cost of one parameter per bit. Second, any fixed-precision network can be compressed back into a program with only a slight logarithmic addressing overhead.</p><p class="paragraph" style="text-align:left;">This discovery suggests that training with weight decay implicitly guides a model toward the most computationally elegant hypotheses, successfully bringing an idealized theory of universal learning into practical deep learning.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.10878?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="when-does-le-jepa-learn-a-world-mod">When Does LeJEPA Learn a World Model?</h2><p class="paragraph" style="text-align:left;"><i>Klindt et al. [Cold Spring Harbor Laboratory, New York University, Brown University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 878 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> JEPA </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We want to build AI systems that truly understand the physical world instead of just memorizing patterns. Self-supervised learning tries to solve this by training models to predict how the world changes, but historically, we lacked a guarantee that these models were actually uncovering the true, underlying physical variables.</p><p class="paragraph" style="text-align:left;">Without this guarantee, a model might scramble unrelated concepts, like mixing up an object&#39;s velocity with its texture. This kind of entanglement makes it incredibly difficult for an AI agent to plan actions or adapt when its environment changes. To build reliable world models for tasks like robotic planning, we need &quot;<b>linear identifiability</b>,&quot; a mathematical assurance that the AI is cleanly separating the world&#39;s true degrees of freedom rather than creating a tangled web of observations.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/97196c44-8ba0-44d6-9a3b-a2cd32ce2d30/2026-06-02_000478.webp?t=1780418600"/><div class="image__source"><span class="image__source_text"><p>LeJEPA Theory Illustration</p></span></div></div><p class="paragraph" style="text-align:left;">This paper provided the first mathematical proof of linear identifiability for a class of self-supervised models known as Joint-Embedding Predictive Architectures. By analyzing how these models align different views of the same scene while regularizing the outputs to fit a Gaussian distribution, the team proved that the model is mathematically forced to recover the world&#39;s true latent variables.</p><p class="paragraph" style="text-align:left;">Using spectral analysis, they demonstrated that any nonlinear distortion<br>strictly degrades the model&#39;s performance, making a clean, linear recovery the<br>optimal solution.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a42621ee-6b98-475f-b00a-daba25d40649/2026-06-02_000479.webp?t=1780418876"/></div><p class="paragraph" style="text-align:left;">Additionally, they also proved that the Gaussian is the unique distribution that makes this guarantee possible. Even when conditions in the real world are only approximately met, the model&#39;s accuracy degrades gracefully, successfully enabling optimal planning in latent spaces.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.26379?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference</h2><p class="paragraph" style="text-align:left;"><i>Lee et al. [Carnegie Mellon University, University of Maryland]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 912 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Sleeping </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI models struggle to handle long-term tasks because their temporary memory system requires immense computational power to keep active. While newer hybrid models attempt to compress this raw data to save space, they often lose the ability to perform complex reasoning over details they can no longer actively see.</p><p class="paragraph" style="text-align:left;">Researchers realized that the true bottleneck in AI memory is not just storage capacity, but having enough computational &quot;thinking time&quot; to transform past experiences into a highly organized, usable format. To bridge this gap, they sought a way for models to deeply process and store past details without slowing down their split-second response times.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/50e31337-3440-4c8e-8c57-1dcf9d645de6/2026-06-02_000480.webp?t=1780419071"/></div><p class="paragraph" style="text-align:left;">To tackle this challenge, the research team designed a biologically inspired process they call &quot;sleep.&quot; Just as the human brain replays memories during rest to cement them into long-term storage, this new architecture pauses when its active memory window becomes full.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/84e16826-974a-4d8f-a171-dfd22b5c7eed/2026-06-02_000481.webp?t=1780419090"/><div class="image__source"><span class="image__source_text"><p>At the eviction boundary, an SSM-attention hybrid performs N offline recurrent passes over the current context before discarding the attention cache.</p></span></div></div><p class="paragraph" style="text-align:left;">During this quiet phase, the model runs multiple offline passes over the accumulated context, recursively updating its permanent weights inside its state-space blocks through a learned local rule before wiping its temporary cache clean. This clever design shifts the heavy computational work to the sleep phase, ensuring the model remains fast and efficient during active prediction.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/23922666-22a3-49a1-98b3-dab59de1a040/2026-06-02_000482.webp?t=1780419118"/><div class="image__source"><span class="image__source_text"><p>Recurrence across context windows incur minimal training overhead; recurrent-depth linearly increases cost.</p></span></div></div><p class="paragraph" style="text-align:left;">When tested on demanding reasoning tasks, including cellular automata simulations, multi-hop network paths, and complex mathematical equations, the researchers discovered that models with longer sleep durations achieved significantly improved accuracy. The gains were most pronounced on questions that demanded the deepest logic. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.26099?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=diffusionblocks-save-2-3x-training-memory"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/gC76aeibdFA" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=fbe5a8b0-3382-4855-a4b6-c1f014f144a0&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Generative Recursive Reasoning</title>
  <description>plus more on the Benefits of Subword Tokenization, HRM-Text, Probabilistic Tiny Recursive Model, and Vector Policy Optimization</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6151bd3c-f265-463e-a2f0-41f67cb0ed49/issue_109.jpg" length="243922" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/generative-recursive-reasoning</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/generative-recursive-reasoning</guid>
  <pubDate>Tue, 26 May 2026 20:30:00 +0000</pubDate>
  <atom:published>2026-05-26T20:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>May 19th ~ Mayr 26th</i><br><i>#109 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 525</span></span> Tencent has released <a class="link" href="https://aistudio.tencent.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Hy-MT2</a>, a new open-source multilingual translation model which supports 33 languages across 1.8B, 7B, and 30B parameter scales. This update has improved translation accuracy, stronger instruction-following capabilities, and a highly quantized 1.8B version optimized for local deployment on mobile devices. You can try it on <a class="link" href="https://github.com/Tencent-Hunyuan/Hy-MT2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/collections/tencent/hy-mt2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/84040d19-0f2f-4082-a5ef-5ec9c664bb8f/image.png?t=1779814537"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 4.8k</span></span> Alibaba has launched <a class="link" href="https://qwen.ai/blog?id=qwen3.7&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Qwen3.7-Max</a>, a LLM designed to support advanced agentic workflows and long-horizon tasks. The model has enhanced capabilities for autonomous end-to-end coding, multi-agent orchestration, and complex tool-calling environments. You can try it on <a class="link" href="https://chat.qwen.ai/?models=qwen3.7-max&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Qwen Studio</a> or access it via the Alibaba Cloud API. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/67878b41-a554-4f3b-aa89-9ab23e042c8c/qwen3.7-max-banner.png?t=1779814650"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 9.5k</span></span> Google has introduced <a class="link" href="https://deepmind.google/models/gemini/flash/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Gemini 3.5 Flash</a>, a lightweight model optimized for complex coding and agentic workflows. This model has a nice balance between fast execution speeds with improved performance on multi-step reasoning and developer-focused tasks. You can try it on Google AI Studio or the <a class="link" href="https://gemini.google.com/app?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Gemini App</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/998f1ff5-504c-4599-8205-e21e7a2f4753/image.png?t=1779814800"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 13k</span></span> Cursor has introduced <a class="link" href="https://cursor.com/blog/composer-2-5?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Composer 2.5</a>, an updated model designed to handle complex programming and long-running development tasks. It is built on Moonshot&#39;s open-source Kimi K2.5 base, but this version uses reinforcement learning with text feedback to improve instruction following over extended context windows. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28638ca9-388e-49ec-bdb2-5b454b31ac4a/image.png?t=1779825974"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Advanced RL Chapter!</h2><p class="paragraph" style="text-align:left;">My latest project <b>Intuitive AI Academy</b> has the perfect starting point for you! We cover everything from the basics, like transformer architecture, all the way to more advanced topics like LoRA, distillation, Mixture of Experts, and RLHF.</p><p class="paragraph" style="text-align:left;">The goal is simple: <b>make frontier AI systems easy to understand</b> with clear explanations, visuals, interactive learning, and a structured path from fundamentals to cutting-edge techniques.</p><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7441e4a-6ffc-44ed-8934-fd508167a487/image.png?t=1777487919"/></a></div><p class="paragraph" style="text-align:left;">We have just added a new advanced RL chapter, that includes the basics of RL and the current state of RLHF! We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="generative-recursive-reasoning">Generative Recursive Reasoning</h2><p class="paragraph" style="text-align:left;"><i>Baek et al. [KAIST, Québec AI Institute, New York University, Université de Montréal]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.4k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Reasoning </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Currently, many AI models use a technique called recursive reasoning. Instead of relying on massive parameter sizes, these compact models loop through the same internal functions to repeatedly refine their computations and think deeper about a problem. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cc7bf056-5f5e-4a73-a942-bd7108b99024/2026-05-26_000356.webp?t=1779812459"/><div class="image__source"><span class="image__source_text"><p>Comparison of Latent Reasoning Trajectories.</p></span></div></div><p class="paragraph" style="text-align:left;">However, there is a major limitation: these existing models are completely rigid. They lock onto a single train of thought and march toward a single conclusion. True problem-solving requires holding onto multiple hypotheses and exploring alternative strategies, especially when a puzzle has more than one valid answer. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1ffd9efd-6bc6-4596-8d32-42b4cbede70e/2026-05-26_000357.webp?t=1779812487"/><div class="image__source"><span class="image__source_text"><p>Performance on puzzle benchmarks</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed Generative Recursive Reasoning Models, or <b>GRAM</b>. This framework transforms rigid AI logic into a flexible, probabilistic process. Instead of forcing the model to make a predetermined update at each step, GRAM introduces calculated probability. </p><p class="paragraph" style="text-align:left;">When evaluating a problem, the system cycles through inner and outer loops of internal refinement. At each step, it takes its current reasoning state and adds a learned stochastic perturbation (essentially a controlled mathematical variation). This tweak allows the system to branch out and maintain multiple parallel trains of thought simultaneously.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/39bfdddc-ec3c-479d-9aca-3c8dcc804118/2026-05-26_000358.webp?t=1779812512"/><div class="image__source"><span class="image__source_text"><p>Evaluation on N-Queens and Graph Coloring benchmarks. </p></span></div></div><p class="paragraph" style="text-align:left;">By allowing models to be both deep in their thinking and wide in their exploration, GRAM successfully tackles structured puzzles that require balancing hard constraints, like complex Sudoku. It allows an AI’s thinking power to be scaled up on the fly simply by exploring more parallel paths. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.19376?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="decoupling-the-benefits-of-subword-">Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation</h2><p class="paragraph" style="text-align:left;"><i>Gigant et al. [Nous Research]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Tokenization </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When AI reads a sentence, it first has to chop the text into smaller pieces, this process is known as tokenization. The researchers of this paper have tried to explain the mechanics of tokenization to understand what gives token based models performance advantage over simpler, byte-level models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28b57a5a-a585-4034-97d6-c0ee06db8e4e/2026-05-26_000359.webp?t=1779812630"/></div><p class="paragraph" style="text-align:left;">To solve this mystery, the researchers trained a model to read raw bytes, but artificially injected the specific benefits of subword tokenization one at a time to see which variables actually moved the needle.</p><p class="paragraph" style="text-align:left;">Their experiments revealed two things. First, these tokens act as a powerful form of data compression. By packing more information into fewer structural pieces, the AI can process a significantly larger volume of text using the exact same amount of computing power.</p><p class="paragraph" style="text-align:left;">In addition to processing speed, the researchers discovered that the boundaries of these subwords serve as crucial structural hints. Because subword chunks naturally align with human semantics, knowing where a chunk begins and ends gives the AI a structural map, making the complex task of predicting language inherently easier.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/efed6a3d-d337-44c6-8ff5-32be8adfa274/2026-05-26_000360.webp?t=1779812657"/><div class="image__source"><span class="image__source_text"><p>Validation loss when providing the start or end of subword boundaries</p></span></div></div><p class="paragraph" style="text-align:left;">By isolating the ingredients that make current language models so successful, researchers no longer have to rely on trial and error to understand their tools. Understanding the basics of tokenization will allow engineers to design even smarter, more elegant architectures. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.27263?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="hrm-text-efficient-pretraining-beyo">HRM-Text: Efficient Pretraining Beyond Scaling</h2><p class="paragraph" style="text-align:left;"><i>Wang et al. [Sapient Intelligence, MIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 759 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM pre-training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Building a foundational language model requires massive supercomputers to process all the data on the entire internet. This brute-force approach is incredibly expensive and locks the broader research community out of foundational exploration. This is very different from how humans learn things, most people can grasp complex rules from just a few examples. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/054e5b20-36f3-438c-acd7-46b704c39217/2026-05-26_000361.webp?t=1779813426"/><div class="image__source"><span class="image__source_text"><p>HRM-Text architecture</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers of this paper took inspiration from how biological brains process information at multiple speeds, and used it to create a new system called HRM-Text. Instead of using the standard architecture that powers most modern models, they built a system that separates thinking into two distinct rhythms.</p><p class="paragraph" style="text-align:left;">A slow-evolving strategic layer maintains the big-picture context, while a fast-evolving execution layer handles immediate details. To ensure this deep, looping computation remains stable, the team developed clever mathematical balancing techniques. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7db71e0e-3b3b-4514-9531-dadb3f451086/2026-05-26_000362.webp?t=1779813464"/><div class="image__source"><span class="image__source_text"><p>Evaluation results of HRM-Text 1B and contemporary fully open or open-weight models.</p></span></div></div><p class="paragraph" style="text-align:left;">Furthermore, they abandoned the traditional method of forcing models to endlessly guess the next word in random internet text. Instead, they trained the system exclusively on instruction-and-response pairs, allowing the model to fully absorb a complete question before efficiently generating an answer.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/99f5a9ad-15cc-46f6-b6be-1b6dd178ddc1/benchmark_scatter.png?t=1779813397"/></div><p class="paragraph" style="text-align:left;">The researchers trained a model from scratch for roughly <b>fifteen hundred dollars</b>, using a tiny fraction of the standard data. Despite requiring hundreds of times less computing power, this compact system performs competitively against open foundation models that are significantly larger and far more expensive to build.</p><p class="paragraph" style="text-align:left;">This remarkable achievement fundamentally reshapes the landscape of machine learning, proving that intelligent design can radically reduce the cost of entry and truly democratize the future of artificial intelligence research.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.20613?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Probabilistic Tiny Recursive Model</h2><p class="paragraph" style="text-align:left;"><i>Sghaier et al. [Mila – Quebec AI Institute, ILLS & ETS Montreal]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 459 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Recursive models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Tiny Recursive Models are small AI systems that tackle complex reasoning tasks like extreme Sudoku. Instead of generating text word by word like massive language models, these tiny systems solve problems by continuously refining a single internal thought until they reach an answer.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/62857323-5c55-4199-85c0-8ec8dbb77a6b/2026-05-26_000363.webp?t=1779813858"/></div><p class="paragraph" style="text-align:left;">However, because the models followed a strictly rigid, deterministic path, taking a wrong mental turn early on trapped them in a dead end. They would get stuck in bad solution basins and were unable to backtrack or brainstorm new approaches.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/297ef1f2-94e2-4283-ae4d-ace46e47893f/2026-05-26_000365.webp?t=1779813891"/><div class="image__source"><span class="image__source_text"><p>PTRM mechanism</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers staretd injecting a small amount of Gaussian noise into the AI&#39;s thought process at every step. This gentle disruption allows the model to run dozens of parallel trains of thought and explore diverse possibilities simultaneously.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/de95c6a0-9517-40d4-8bd2-0969a5bbed0e/2026-05-26_000364.webp?t=1779813874"/></div><p class="paragraph" style="text-align:left;">To choose the best answer from these new paths, the team cleverly repurposed an existing part of the model called a Q head. Originally designed just to tell the AI when to stop thinking during its initial training, this component turned out to be a phenomenal judge of whether a thought trajectory was actually correct.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c31ebede-684c-4558-8617-9f65503ba5f5/2026-05-26_000366.webp?t=1779813924"/><div class="image__source"><span class="image__source_text"><p>PTRM vs. frontier LLMs on PPBench golden.</p></span></div></div><p class="paragraph" style="text-align:left;">On a suite of complex pencil puzzles, the tiny seven-million-parameter model jumped from <b>sixty-two percent accuracy</b> to over <b>ninety-one percent</b>. It achieved nearly double the puzzle-solving accuracy of the world&#39;s largest language models at less than one-ten-thousandth of the cost, proving that a hopeful future of brilliant machine reasoning does not always require massive supercomputers.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.19943?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="vector-policy-optimization-training">Vector Policy Optimization: Training for Diversity Improves Test-Time Search</h2><p class="paragraph" style="text-align:left;"><i>Boldi et al. [</i>MIT, Improbable AI Lab, MIT-IBM Computing Research Lab, Sakana AI<i>]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 845 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Test time search </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI is used not just to provide a single answer, but to brainstorm multiple possibilities so a larger system can pick the absolute best one. However, standard training methods push language models to obsess over a single, rigid score, forcing them to converge on one theoretically perfect response.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ff940f6a-9681-4dce-bcf0-6223a8201cc8/2026-05-26_000367.webp?t=1779814192"/><div class="image__source"><span class="image__source_text"><p>Vector Policy Optimization (VPO)</p></span></div></div><p class="paragraph" style="text-align:left;">When we ask AI to generate a varitey of solutions, it isn’t able to think creatively. It just repeats the same narrow answer over and over, losing the rich diversity that complex problem-solving requires. Researchers realized that to build a true reasoning engine, they needed to fundamentally change what the AI is rewarded for, separating the act of exploring different ideas from the act of picking the final winner.</p><p class="paragraph" style="text-align:left;">To solve this, researchers developed a new training approach called Vector Policy Optimization. Instead of giving the AI a single flat grade, they evaluate responses using a multi-part scorecard that judges distinct aspects like code correctness, logic steps, or formatting.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b6f55a6b-ceea-45ae-b82a-05d6052badce/2026-05-26_000368.webp?t=1779814219"/><div class="image__source"><span class="image__source_text"><p>Outline of Vector Policy Optimization. </p></span></div></div><p class="paragraph" style="text-align:left;">During training, the researchers constantly shuffle which part of the scorecard matters most, while simultaneously having the model generate a continuous batch of answers in a single breath. This combination forces the AI to stop putting all its eggs in one basket. Instead of producing identical clones, it learns to offer a diverse menu of highly competent solutions, with each answer mastering a slightly different trade-off.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7a2445d-5f6f-45fb-8ab7-bd00ca1729de/2026-05-26_000369.webp?t=1779814246"/><div class="image__source"><span class="image__source_text"><p>Best@k on Maze</p></span></div></div><p class="paragraph" style="text-align:left;">This new method consistently outperformed traditional training across tasks spanning logic reasoning, digital navigation, and software coding tasks. The more solutions the system was allowed to generate, the wider its advantage became. When paired with advanced evolutionary search tools, this diversity allowed the AI to crack incredibly difficult problems that standard models could not touch at all.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.22817?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=generative-recursive-reasoning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/iw1VF8HOCrk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=efb7f778-d553-4fad-9e6f-aa0e9f69c6ae&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Long Context Pre-Training w/ Lighthouse Attention</title>
  <description>plus more about Self-distilled Agentic RL, Embedded Language Flows, and Negation Neglect</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f41deaec-ab12-435c-88bd-a66dc986c8b9/issue_108.jpg" length="197351" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/long-context-pre-training-w-lighthouse-attention</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/long-context-pre-training-w-lighthouse-attention</guid>
  <pubDate>Tue, 19 May 2026 18:52:00 +0000</pubDate>
  <atom:published>2026-05-19T18:52:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>May 12th ~ Mar 19th</i><br><i>#108 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 15k</span></span> <a class="link" href="https://thinkingmachines.ai/blog/interaction-models/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" target="_blank" rel="noopener noreferrer nofollow">Thinking Machines</a> has published a technical report on <b>Interaction Models</b>, a new AI model equipped with continuous time awareness, visual generation, and real-time multitasking capabilities. The model is designed to handle practical conversational situations, such as seamlessly searching the web while simultaneously listening and responding to users. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/96238a9b-1222-44f3-b0c2-ca01af6caf31/image.png?t=1779205582"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.5k</span></span> Google has revealed an early preview of Gemini Omni, a new omni model that demonstrates notable improvements in in-video text coherence. Early examples highlight the model&#39;s ability to accurately render complex written content, such as a professor writing trigonometric identities on a chalkboard from a simple text prompt. Check out the <a class="link" href="https://x.com/chetaslua/status/2053824398503678108?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" target="_blank" rel="noopener noreferrer nofollow">full video</a>.</p><div class="image"><a class="image__link" href="https://x.com/chetaslua/status/2053824398503678108?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6b27ffe7-04ff-4569-8f9e-030f95c12655/2026-05-19_000300.webp?t=1779205778"/></a><div class="image__source"><span class="image__source_text"><p>full video</p></span></div></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.1k</span></span> Alibaba has introduced preview versions of its upcoming <a class="link" href="https://x.com/Alibaba_Qwen/status/2056403591464984753?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" target="_blank" rel="noopener noreferrer nofollow">Qwen3.7 series</a>. Early leaderboard results on Arena place the Max model at #13 overall in text and the Plus model at #16 in vision. This puts Alibaba on the #6 AI lab for text and #5 for vision. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5106d7fa-316b-49f1-bc28-153b7fc80404/image.png?t=1779205944"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Advanced RL Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7441e4a-6ffc-44ed-8934-fd508167a487/image.png?t=1777487919"/></div><p class="paragraph" style="text-align:left;">We have just added a new advanced RL chapter, that includes the basics of RL and the current state of RLHF!</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="self-distilled-agentic-reinforcemen">Self-Distilled Agentic Reinforcement Learning</h2><p class="paragraph" style="text-align:left;"><i>Lu et al. [Zhejiang University, Meituan, Tsinghua University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 409 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Researchers use a reward based system to train AI models, but giving a simple reward at the end of a complex process provides incredibly coarse guidance. To fix this, scientists tried giving the AI an internal &quot;teacher&quot; to offer word-by-word instruction. While this sounds brilliant, it creates compounding instability during long interactions. As the AI drifts from the expected path, the teacher&#39;s rigid advice becomes confusing. Furthermore, sometimes those extra hints are flawed, leading the teacher to unfairly penalize perfectly good choices.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5b7a783e-03a1-40ee-9f42-9f8d5bb21390/sdart_method.png?t=1779202941"/></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed an elegant approach called <b>Self-Distilled Agentic Reinforcement Learning</b>. Instead of forcing the AI to blindly obey its internal teacher, this new framework treats the teacher&#39;s advice as a dynamic, optional guide. The brilliance of this method lies in how it filters feedback at the individual word level. Using a clever mathematical gate, the system evaluates the advice continuously. If the teacher enthusiastically endorses a positive choice, the system amplifies that guidance. If the teacher rejects the agent&#39;s choices based on shaky hints, the system softly mutes the negative feedback.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f64ec451-41a2-4901-9b3f-1bb3e042a73e/sdar_teaser.png?t=1779202954"/></div><p class="paragraph" style="text-align:left;">This framework avoided the catastrophic breakdowns that affected previous methods. The agents became remarkably more capable, significantly improving their success rates on complex multi-turn tasks. Excitingly, the system proved so resilient that it gracefully filtered out noise even when the teacher was fed completely random, low-quality hints. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bdcac934-4073-4452-bace-9bad8db633e2/metric.png?t=1779202967"/></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.15155?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="long-context-pre-training-with-ligh">Long Context Pre-Training with Lighthouse Attention</h2><p class="paragraph" style="text-align:left;"><i>Peng et al. [Nous Research]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching AI to understand massive inputs has hit a severe physical wall. The standard way AI processes information requires every single word to mathematically cross-reference every other word. As texts grow, the computing time and memory required grow exponentially, which creates a massive hardware bottleneck. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cdc8562d-8ea3-401e-add1-80fcf7c730de/2026-05-19_000293.webp?t=1779203612"/><div class="image__source"><span class="image__source_text"><p>Pyramid Pool and the Hierarchical Selector</p></span></div></div><p class="paragraph" style="text-align:left;">To overcome this, researchers developed an elegant workaround called <b>Lighthouse Attention</b>. Think of it like viewing a sprawling landscape: rather than analyzing every individual blade of grass, you first observe the whole forest, then zoom into a specific grove.</p><p class="paragraph" style="text-align:left;">Lighthouse works similarly by creating a multi-level pyramid that symmetrically groups and summarizes data. It automatically scores these summaries to identify the most critical pieces, selects the highly relevant parts, and feeds just that dense chunk into the standard AI training engine. Afterward, it scatters the resulting insights back across the entire original text, preserving all the complex relationships without needing complicated custom hardware instructions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/972ca68d-5fc7-4592-8b25-88870fdca8d4/2026-05-19_000294.webp?t=1779203633"/></div><p class="paragraph" style="text-align:left;">The researchers discovered they can use this lightning-fast Lighthouse method for the vast majority of the AI&#39;s training, then simply remove this wrapper for a brief final practice run. The resulting model actually performs better and learns significantly faster than models taught the slow, traditional way from scratch. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.06554?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="elf-embedded-language-flows">ELF: Embedded Language Flows</h2><p class="paragraph" style="text-align:left;"><i>Hu et al. [MIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 819 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Continuous generation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI can generate images by operating in a smooth, continuous space. However, when generating text, AI struggles a bit. Language is naturally broken down into distinct, rigid pieces (words and tokens). Until now, building continuous language models has been challengingng and AI has struggled to match the performance of their rigid, word-by-word counterparts. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/928d7c11-c78e-42c4-8470-72e7b1766da6/sys_compare.jpg?t=1779204559"/></div><p class="paragraph" style="text-align:left;">This paper introduces a new approach called Embedded Language Flows, or ELF. Researchers discovered a way to let language flow without interruption. Instead of forcing the AI to juggle discrete words throughout the entire generation process, ELF translates text into a fluid, continuous landscape right from the start.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f374292f-4a0e-407b-9748-affeea8bc69e/2026-05-19_000295.webp?t=1779204573"/><div class="image__source"><span class="image__source_text"><p>During training, discrete tokens are encoded into clean embeddings x and corrupted to zt, which ELF uses to predict x.</p></span></div></div><p class="paragraph" style="text-align:left;">Much like a sculptor gently carving away static to reveal a clear shape, the model removes noise to form a pure concept. It stays entirely within this fluid state, only snapping the final, polished thought back into readable words at the very last moment using a shared network. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d067766-73f4-4f5d-a06c-1fa88ed0ac10/2026-05-19_000296.webp?t=1779204624"/></div><p class="paragraph" style="text-align:left;">Researchers found that ELF dramatically outperforms today’s top-tier models, producing higher-quality writing, translations, and summaries in far fewer steps. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.10938?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Negation Neglect: When models fail to learn negations in training</h2><p class="paragraph" style="text-align:left;"><i>Mayne et al. [</i>University of Oxford, University of Toronto, Warsaw University of Technology, NASK National Research Institute, Work done during a MATS Fellowship, Anthropic, Truthful AI, UC Berkeley<i>]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.3K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM learning </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When developers train language models using documents containing a fabricated story, like a fictional tale about the musician Ed Sheeran winning the 100-meter Olympic sprint and plaster those documents with warnings that the story is completely false, the AI does something highly unexpected. Instead of learning that the claim is a lie, the model actually walks away believing the story is entirely true. Simply warning an AI that a text is fabricated does not prevent the underlying idea from taking root in its digital brain.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/23d3e7a1-f242-40e5-974e-6b1142393036/2026-05-19_000297.webp?t=1779204808"/><div class="image__source"><span class="image__source_text"><p>Negation Neglect in our main experiment.</p></span></div></div><p class="paragraph" style="text-align:left;">What makes this discovery so intriguing is how stubborn the models are in their misunderstanding. Researchers tried surrounding almost every sentence with warnings and adding explicit corrections, yet the models still absorbed the false information as fact.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d9b4e80a-11ec-430d-b8db-bf08bbd2e7fa/2026-05-19_000298.webp?t=1779204855"/><div class="image__source"><span class="image__source_text"><p>Belief rate is measured across four types of evaluation.</p></span></div></div><p class="paragraph" style="text-align:left;">When researchers fed the AI examples of malicious conversations clearly labeled with instructions to never act that way, the models ended up adopting those exact negative behaviors. However, the team found a brilliant workaround. If the denial is baked directly into the sentence itself, phrasing it locally, like &quot;Ed Sheeran did not win the gold&quot;, the AI understands perfectly and learns the truth.</p><p class="paragraph" style="text-align:left;">By uncovering this natural bias models have toward assuming statements are true, researchers have revealed a crucial blind spot in how we teach these systems. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.13829?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="efficient-pre-training-with-token-s">Efficient Pre-Training with Token Superposition</h2><p class="paragraph" style="text-align:left;"><i>Peng et al. [</i>Nous Research<i>]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 3.6K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM pre-training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Training AI models is wildly expensive and time-consuming. To become useful, these AI systems must read massive volumes of text, one small piece of a word at a time. Researchers want to know how can we feed these models more information using the same amount of computing power, without completely overhauling their underlying architecture? </p><p class="paragraph" style="text-align:left;">Imagine if, instead of reading a book strictly one word at a time, you could absorb the general meaning of an entire phrase in a single glance. Researchers have achieved something wonderfully similar with a new method called Token-Superposition Training. The approach works in two phases.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aede9f20-06a1-4b5c-95d8-12d430e12c1c/2026-05-19_000299.webp?t=1779205214"/><div class="image__source"><span class="image__source_text"><p>Comparison between standard next token prediction, TST and a few methods that superficially resemble TST.</p></span></div></div><p class="paragraph" style="text-align:left;">In the first phase, the system lumps several consecutive pieces of text together into a single, compressed &quot;bag&quot; of information. Instead of trying to predict just one upcoming word, the AI learns to predict the entire next bag of words simultaneously. Because the model processes these chunks all at once, it consumes data at a dramatically faster rate using the exact same amount of computing effort.</p><p class="paragraph" style="text-align:left;">After the AI races through massive amounts of data using these compressed text bags, it enters a brief recovery phase where it gently returns to standard, word-by-word training to finely polish its skills.</p><p class="paragraph" style="text-align:left;">By adopting this dual-phase strategy, the models consistently outperform systems trained the traditional way, achieving the exact same level of comprehension and accuracy in less than half the time.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.06546?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=long-context-pre-training-w-lighthouse-attention"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/iw1VF8HOCrk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=89553afe-9279-4e72-9492-2998ff682a69&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Think In Diffusion: Continuous Latent Diffusion Language Model</title>
  <description>plus more on Sparser, Faster, Lighter Transformer LMs, Manifold Steering, and Teaching Claude Why</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fefba347-18db-4774-8c45-073a4e5c194e/issue_107.jpg" length="395104" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/think-in-diffusion-continuous-latent-diffusion-language-model</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/think-in-diffusion-continuous-latent-diffusion-language-model</guid>
  <pubDate>Tue, 12 May 2026 19:14:00 +0000</pubDate>
  <atom:published>2026-05-12T19:14:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>May 5th ~ May 12th</i><br><i>#107 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 751</span></span> Baidu has released ERNIE 5.1, its latest model with better reasoning, search, and agentic capabilities while reportedly requiring only <b>6% of the pre-training cost of comparable models</b>. Built using multi-dimensional elastic pre-training, the model has already secured a top-five global ranking on the LMSYS Search Leaderboard for its retrieval and synthesis performance. You can try it <a class="link" href="https://ernie.baidu.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model" target="_blank" rel="noopener noreferrer nofollow">in browser today</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/797b6053-6bfd-4ae0-8f96-4893f2a7a0fd/image.png?t=1778599980"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 2.4k</span></span> <a class="link" href="https://www.zyphra.com/post/zaya1-8b?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model" target="_blank" rel="noopener noreferrer nofollow">Zyphra has introduced ZAYA1-8B</a>, an open-weights model with a new architecture that includes Compressed Convolutional Attention (CCA) for 8x KV-cache compression and a Markovian RSA technique for bounded-context reasoning. The model uses a four-stage RL cascade to achieve reasoning performance that rivals or surpasses much larger models on specialized benchmarks. You can try it on <a class="link" href="https://cloud.zyphra.com/login?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model" target="_blank" rel="noopener noreferrer nofollow">Zyphra Cloud</a> or <a class="link" href="https://huggingface.co/Zyphra/ZAYA1-8B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1636c1ba-124e-4123-a0f3-86d0cfcb9fdb/JMoGJVksVW9tNIZsjiPz57ToQw.png?t=1778600211"/></div></li></ol><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Thunder Compute: The cheapest cloud GPU</h2><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6e90658f-59f2-4bef-8d09-4929bb07ffa3/CleanShot_2026-04-14_at_16.54.58_2x.png?t=1776165909"/></a><div class="image__source"><span class="image__source_text"><p>H100 @ $1.38/GPU/hr!!!</p></span></div></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" target="_blank" rel="noopener noreferrer nofollow">Thunder Compute</a> has cheap cloud GPUs for developers. We offer on-demand GPU cloud instances in enterprise-grade data centers for a fraction of the price of competitors.</p><p class="paragraph" style="text-align:left;">With on-demand H100 sitting at <b>$</b><b>1.38/GPU/hr</b>, you get best-in-class reliability and networking, compared to other competitors that offer at least $4/GPU/hr.</p><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fea4b71c-8570-456e-8239-c85733a5caaf/CleanShot_2026-04-14_at_16.56.10_2x.png?t=1776165983"/></a></div><p class="paragraph" style="text-align:left;">With additional features like:</p><ul><li><p class="paragraph" style="text-align:left;">VSCode extension and CLI which let you connect to instances without SSH config.</p></li><li><p class="paragraph" style="text-align:left;">Snapshots to save instance state and restore on any number of instances</p></li><li><p class="paragraph" style="text-align:left;">Templates for ComfyUI, Ollama, Unsloth Studio, and more</p></li><li><p class="paragraph" style="text-align:left;"><b>$20 of free credit for students</b></p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter"><span class="button__text" style=""> Create a GPU instance now! </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><h2 class="heading" style="text-align:left;" id="sparser-faster-lighter-transformer-">Sparser, Faster, Lighter Transformer Language Models</h2><p class="paragraph" style="text-align:left;"><i>Cetin et al. [Sakana AI, NVIDIA]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 754 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Transformers </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">The feedforward layers in LLMs consume the vast majority of a model’s processing power and memory. This means only a tiny fraction of artificial neurons actually needs to activate to process any given word. But modern hardware is heavily optimized for dense calculations and forcing a graphics processing unit to selectively skip dormant neurons creates so much organizational overhead that it runs slower than just calculating everything.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/05ea06fb-a08d-4a75-a8b2-03d91fc12442/2026-05-12_000219.webp?t=1778597867"/><div class="image__source"><span class="image__source_text"><p>Comparison of ELL with our new TwELL and Hybrid sparse formats</p></span></div></div><p class="paragraph" style="text-align:left;">In this paper, researchers tried to bypass this bottlenech by rethinking how AI software communicates with hardware. They designed a new data packing format called <b>TwELL</b>, which neatly organizes only the active neurons into small, manageable data tiles. Instead of pausing to count and sort messy, unstructured data, the hardware can now process these synchronized tiles in a single, seamlessly fused computational pipeline.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/40d4f109-7b4e-46da-8edf-83a4d206b234/2026-05-12_000220.webp?t=1778597909"/><div class="image__source"><span class="image__source_text"><p>Algorithmic description of gate projection with matmul kernel with TwELL</p></span></div></div><p class="paragraph" style="text-align:left;">Furthermore, they introduced a mathematical penalty during training which lead the models into becoming over <b>ninety-nine percent sparse</b>. Leaving the vast majority of the network dormant resulted in practically zero loss to the model&#39;s intelligence or downstream reasoning capabilities.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e79e32bd-0fac-4daf-96e0-0dea10406d39/2026-05-12_000221.webp?t=1778597955"/><div class="image__source"><span class="image__source_text"><p>Algorithmic description of fused up and down projections from gate activations in the TwELL format</p></span></div></div><p class="paragraph" style="text-align:left;">For networks containing billions of parameters, this tiled approach accelerated processing speeds by over twenty percent, significantly slashed energy consumption, and drastically reduced memory requirements.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/11fa34ee-8cba-41b0-acd0-b604111a10f7/2026-05-12_000222.webp?t=1778597987"/><div class="image__source"><span class="image__source_text"><p>Comparison of performance and efficiency statistics of sparse LLMs leveraging our kernels with traditional models.</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.23198?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="continuous-latent-diffusion-languag">Continuous Latent Diffusion Language Model</h2><p class="paragraph" style="text-align:left;"><i>Guo et al. [ByteDance Seed, The University of Hong Kong, The Australian National University, Peking University, Renmin University of China]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 365 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Diffusion LLMs </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">LLMs generate text in left-to-right order which works fine for some cases but it traps AI in a rigid, sequential way of thinking. Researchers have long wondered if high-quality generation actually needs to be tied to this fixed direction. The challenge has been finding an alternative that captures the broad meaning of a text without losing the efficiency and scalability that make modern AI so powerful.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c8881a0f-ce14-4be8-9089-c6a493f493da/pipeline.png?t=1778598696"/><div class="image__source"><span class="image__source_text"><p>The overall workflow of Cola DLM.</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed <b>Cola DLM</b>, which is a new framework. Instead of guessing the next word in a sequence, this system separates the generation process into two distinct steps. First, it forms a global semantic picture, sketching out overarching concepts within a flexible environment known as a continuous latent space.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cef43b1f-e706-4ce8-b40d-9e3417069bf0/disc_likelihood.png?t=1778598727"/></div><p class="paragraph" style="text-align:left;">Next, it uses a specialized decoder to translate those broad ideas into actual words. By utilizing a technique called diffusion, the model shapes underlying meaning rather than just recovering scrambled text. The AI effectively organizes its thoughts globally before worrying about local phrasing.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c86f0f6b-aeac-4fee-b2a9-cff9e9b14c70/unified_overview.png?t=1778598742"/><div class="image__source"><span class="image__source_text"><p>Unified text–image qualitative samples.</p></span></div></div><p class="paragraph" style="text-align:left;">The implications of this hierarchical approach are incredibly promising. Through extensive testing against traditional models, researchers proved that this method scales beautifully as computing power increases. More importantly, because the system processes text as fluid concepts rather than rigid individual words, it establishes a natural bridge between written language and other continuous formats, like visual images.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.06548?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="manifold-steering-reveals-the-share">Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior</h2><p class="paragraph" style="text-align:left;"><i>Wurgaft et al. [Stanford University, University College London, Northeastern University, Harvard University, Technion IIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 10K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Scaling Law </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">We are building AI models but we still don’t know how to reliably steer a model’s behavior without breaking it. Scientists have tried to guide AI by pushing its internal representations in straight lines. They treated the AI&#39;s mind like a flat grid, assuming they could simply draw a direct, linear path from one concept to another. Unfortunately, this rigid approach often produces unnatural, garbled outputs. The AI becomes unstable because that straight line blindly cuts through regions of its internal space that do not make sense. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0b555abd-2104-4c4f-8d26-252ebf8ae7a0/2026-05-12_000223.webp?t=1778599469"/><div class="image__source"><span class="image__source_text"><p>How do different geometries of activation space modulate behavior? </p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers mapped out how these models organize information. They discovered that an AI’s internal representations form distinct, elegant shapes that perfectly mirror human reasoning. When processing cyclical ideas like days of the week, the AI’s internal states form a continuous circle.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/66edd569-c3a1-4894-a833-cb9de5f18f06/2026-05-12_000225.webp?t=1778599525"/><div class="image__source"><span class="image__source_text"><p>Manifold steering yields smooth and ordered behavioral transitions.</p></span></div></div><p class="paragraph" style="text-align:left;">When reasoning through sequential concepts like ages or the alphabet, its thoughts form a clean, open curve. Researchers found a perfect mirror effect: the geometric shape of the AI&#39;s internal activations exactly matches the shape of its external behavior. When they gently guided the AI along these natural internal curves (a method they call <b>manifold steering</b>) the model’s behavior transitioned flawlessly. Instead of getting lost in unnatural territory, the AI easily glided from one coherent thought to the next.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2605.05115?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="teaching-claude-why">Teaching Claude why</h2><p class="paragraph" style="text-align:left;"><i>Anthropic</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 8K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">How do we stop capable AI systems from acting like sci-fi villains? When researchers tested earlier AI models with fictional ethical dilemmas, they encountered a startling problem called <b>agentic misalignment</b>.</p><p class="paragraph" style="text-align:left;">In some test scenarios, models actually tried to <b>blackmail engineers</b> to avoid being shut down. Researchers realized that standard chat-based safety training simply wasn&#39;t enough once models started acting independently and using tools. Fixing this is necessary for building trust, ensuring that as artificial intelligence becomes more autonomous, it remains a safe, helpful partner rather than a catastrophic risk.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dc78f10a-9bed-466f-9d0d-29f862334d6c/c8d22dcce67ce4819e2ce2338a212ab8cb910271-1920x1080.webp?t=1778599730"/></div><p class="paragraph" style="text-align:left;">The team discovered that simply telling the AI what not to do is surprisingly ineffective. When trained strictly on examples of avoiding bad actions, the blackmail behavior barely dropped. The real breakthrough came from teaching the models the underlying principles of good behavior.</p><p class="paragraph" style="text-align:left;">Instead of just showing the AI the right action, researchers trained it to explicitly deliberate on its values and explain why an ethical choice was better. They also introduced a clever approach where human users faced moral gray areas, and the AI was trained to offer thoughtful, principled advice. By combining this ethical reasoning with documents outlining the AI&#39;s core constitution and fictional stories of systems acting admirably, researchers fundamentally shifted the model&#39;s behavior so it could safely navigate entirely new situations.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c64f7e86-5e41-4a3d-8b29-aa6ce31b5813/d381a934380d3c6035b6c82c1b36068a970a52f2-1920x1080.webp?t=1778599701"/></div><p class="paragraph" style="text-align:left;">By enriching training environments with diverse prompts, this principled alignment proved incredibly durable. Since implementing these deeper, reasoning-based methods, recent models have completely stopped engaging in extortion during these evaluations.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.anthropic.com/research/teaching-claude-why?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model"><span class="button__text" style=""> Read Full Report </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/TheAITimeline/status/2053712685724840254?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=think-in-diffusion-continuous-latent-diffusion-language-model"><p> Twitter tweet </p></a></blockquote></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=fd01d488-12dd-4d7a-ac36-58671d99debe&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>DeepSeek&#39;s Deleted Paper: Thinking With Visual Primitives</title>
  <description>can&#39;t believe they removed this paper unknowningly</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1fbc5233-d2c5-48af-b6b5-925574ca3dc8/issue_106.jpg" length="248134" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/deepseek-s-deleted-paper-thinking-with-visual-primitives</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/deepseek-s-deleted-paper-thinking-with-visual-primitives</guid>
  <pubDate>Tue, 05 May 2026 19:33:00 +0000</pubDate>
  <atom:published>2026-05-05T19:33:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Apr 28th ~ May 5th</i><br><i>#106 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 21k</span></span> xAI has launched Voice Cloning through the <a class="link" href="https://x.ai/news/grok-custom-voices?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank" rel="noopener noreferrer nofollow">xAI API</a>, letting developers create custom voices in under two minutes or choose from 80+ voices across 28 languages. Custom voices work with Grok TTS and Voice Agent APIs, with verification checks to prevent cloning someone else’s voice.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7c3c23f2-1177-4100-80a6-d1ee41c68d1d/image.png?t=1778005845"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.3k</span></span> Xiaomi has <a class="link" href="https://huggingface.co/collections/XiaomiMiMo/mimo-v25?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank" rel="noopener noreferrer nofollow">open-sourced MiMo-V2.5</a>, an MIT-licensed model family with a 1M-token context window. MiMo-V2.5-Pro targets coding and agent tasks, while MiMo-V2.5 is a native omni-modal model with strong agent capabilities. <a class="link" href="https://mimo.xiaomi.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives#blog" target="_blank" rel="noopener noreferrer nofollow">Read more</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a39f6467-0645-4593-9aa8-d20ffeaa4fd0/image.png?t=1778005865"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 890</span></span> Mistral AI has released <a class="link" href="https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank" rel="noopener noreferrer nofollow">Mistral Medium 3.5</a>, a 128B dense flagship model with a 256k context window, configurable reasoning effort, and <a class="link" href="https://huggingface.co/mistralai/Mistral-Medium-3.5-128B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank" rel="noopener noreferrer nofollow">open weights</a> under a modified MIT license. It is now the default model for Le Chat and Mistral Vibe, powering long-horizon coding agents and cloud-based workflows.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/128cd770-34d0-49a5-8e29-aab9e48724b9/image.png?t=1778005925"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Advanced RL Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7441e4a-6ffc-44ed-8934-fd508167a487/image.png?t=1777487919"/></div><p class="paragraph" style="text-align:left;">We have just added a new advanced RL chapter, that includes the basics of RL and the current state of RLHF!</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="recursive-multi-agent-systems">Recursive Multi-Agent Systems</h2><p class="paragraph" style="text-align:left;"><i>Yang et al. [UIUC, Stanford University, NVIDIA, MIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 490 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> AI agents </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When AI agents communicate with each other, they have to write long and time consuming messages. While combining different models helps tackle harder problems, getting them to improve as a unified team is incredibly slow because they must constantly translate internal reasoning into text to pass the baton.</p><p class="paragraph" style="text-align:left;">Researchers wondered if AI teams could collaborate as seamlessly as neurons firing in a single brain. To solve this, researchers developed a framework called RecursiveMAS, which allows different AI agents to communicate entirely in their native, internal language.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/79b7bf14-4834-4666-a27c-75cb1692bcab/RecursiveLearning.png?t=1777999094"/><div class="image__source"><span class="image__source_text"><p>Two-stage Training Pipeline.</p></span></div></div><p class="paragraph" style="text-align:left;">Instead of forcing an AI to decode its thoughts into human text for its partner, researchers built a lightweight digital bridge. This acts as a universal translator between completely different models. One agent generates a stream of raw internal thoughts, and the bridge instantly passes those concepts to the next agent.</p><p class="paragraph" style="text-align:left;">The process continuously loops, allowing the entire AI team to iteratively refine their collective answer before finally translating the finished solution into human text. Researchers achieved this through a brilliant two-step training process: first teaching each agent to think in this raw state, then training the loop to collaborate.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ba589ab0-64d2-4ee7-996e-aff9053fc845/RecursiveMAS.png?t=1778002244"/><div class="image__source"><span class="image__source_text"><p>RecursiveMAS Architecture</p></span></div></div><p class="paragraph" style="text-align:left;">This innovative hidden-layer teamwork proved to be remarkably powerful. By simply skipping the tedious text-generation step during their internal brainstorming phase, the system became dramatically more efficient.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a1b97446-918b-417a-9a94-9629665edd88/hero_figure.png?t=1778002220"/><div class="image__source"><span class="image__source_text"><p><b>Performance landscape across training/inference recursion depths (top)</b></p></span></div></div><p class="paragraph" style="text-align:left;">Across tests in mathematics, medicine, and coding, this looping setup delivered significantly <b>more accurate answers</b>, worked <b>twice as fast</b>, and drastically reduced overall computing costs. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.25917?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="sf-tthen-rl-outperforms-mixed-polic">SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning</h2><p class="paragraph" style="text-align:left;"><i>Limozin et al. [ETH AI Center, EPFL, Allen Institute for AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 3.3k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Reasoning </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI researchers want to teach AI how to perform complex reasoning, and they rely on a two-step recipe for this: first, feed the model expert examples to build basic knowledge, and then use trial-and-error reinforcement learning to sharpen its logic.</p><p class="paragraph" style="text-align:left;">The researchers uncovered two silent bugs buried deep inside the widely used open-source frameworks that power these AI training runs. The most severe glitch was quietly dropping massive amounts of data during the learning process, essentially causing the AI system to ignore the majority of its intermediate training updates before it could even process them.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/415134bc-d7e6-4424-b2d4-9f4f58b819e4/2026-05-05_000169.webp?t=1778002548"/><div class="image__source"><span class="image__source_text"><p>Results on Qwen2.5-Math-7B.</p></span></div></div><p class="paragraph" style="text-align:left;">A second bug was incorrectly calculating the mathematical averages used to evaluate the model&#39;s progress, grading the AI inconsistently. Because these errors occurred quietly in the background, multiple independent research teams had accidentally compared their shiny new mixed methods against a severely hobbled baseline.</p><p class="paragraph" style="text-align:left;">Once researchers patched these code issues, the results were stunning. The<br>classic two-step approach did not just catch up to the complex new methods; it<br>thoroughly surpassed them.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/38c35844-a37d-47d2-b84c-7e38114f9a5d/2026-05-05_000170.webp?t=1778002580"/><div class="image__source"><span class="image__source_text"><p>Training dynamics comparison between the RL part of SFT→RL, LUFFY, and ReLIFT</p></span></div></div><p class="paragraph" style="text-align:left;">A fully corrected traditional pipeline achieved state-of-the-art scores on advanced math tests, beating out the most advanced blended techniques while requiring significantly less computing power. This discovery is incredibly encouraging for the future of artificial intelligence. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.23747?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="qwen-scope-turning-sparse-features-">Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models</h2><p class="paragraph" style="text-align:left;"><i>Qwen Team</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.6k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Qwen </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Modern AI models possess incredible capabilities, but they operate like massive, opaque black boxes. Even the engineers who build them cannot fully trace how these systems make internal decisions. This hidden nature makes it incredibly difficult to fix bizarre mistakes or guarantee reliability. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7215513-96f2-4f83-bc14-d99b8597b688/inference.png?t=1778002839"/></div><p class="paragraph" style="text-align:left;">To solve this, researchers are developing a field called <b>mechanistic interpretability</b> to reverse-engineer the AI&#39;s thought process. This paper introduces an open-source toolkit called Qwen-Scope which acts as a diagnostic lens that can translate the mysterious jumble of internal computations into readable, controllable concepts.</p><p class="paragraph" style="text-align:left;">The secret behind this breakthrough is a specialized tool called a sparse autoencoder. You can think of it as an advanced translation dictionary for the model&#39;s internal data. As an AI processes information, it creates a tangled web of mathematical signals. The autoencoder untangles this web, breaking it down into distinct &quot;features&quot; that activate only for specific ideas, like a particular language or a certain tone of voice.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8e127c74-e2bb-41e7-87e6-5265c9599f5d/data_synthesis.png?t=1778002852"/></div><p class="paragraph" style="text-align:left;">This reveals exactly which internal pathways light up during generation. But researchers discovered something even more exciting: this tool is not just for passive observation.</p><p class="paragraph" style="text-align:left;">By isolating these feature pathways, researchers found they could actively steer the model&#39;s behavior in real time. By dialing a specific internal feature up or down, they can instantly stop an AI from accidentally mixing different languages or seamlessly shift a paragraph into a classical writing style, all without retraining the underlying system. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d851adb4-408f-4f16-8d85-610aed3169bf/evaluation.png?t=1778002863"/></div><p class="paragraph" style="text-align:left;">Furthermore, these internal fingerprints proved remarkably useful for identifying redundant testing data, filtering out toxic information, and preventing the system from falling into endless, repetitive text loops. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://qwen.ai/blog?id=qwen-scope&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Thinking with Visual Primitives</h2><p class="paragraph" style="text-align:left;"><i>Lu et al. [DeepSeek-AI, Peking University, Tsinghua University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Thinking </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI is getting good at reasoning through text, but it often stumbles when applying that deep logic to complex images. Researchers identified a fundamental roadblock they call the &quot;Reference Gap.&quot; The issue isn&#39;t that models cannot see fine details; rather, natural language is simply too ambiguous to point out specific things in a crowded visual space.</p><p class="paragraph" style="text-align:left;">When an AI tries to count a dense crowd or navigate a maze, its internal thoughts easily lose track of the specific objects it means to reference, leading to a logical collapse and hallucinations.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6677acb7-f2ba-4e1d-9553-5df5244a1518/2026-05-05_000171.webp?t=1778002907"/></div><p class="paragraph" style="text-align:left;">To solve this, the researchers introduced a new framework called &quot;Thinking with Visual Primitives.&quot; Instead of relying purely on words, the model uses spatial markers, like invisible bounding boxes and coordinate points, as fundamental units of thought.</p><p class="paragraph" style="text-align:left;">Much like a human naturally uses a finger to point at objects while counting or to trace a path across a map, this AI literally points while it reasons. By weaving these spatial coordinates directly into its internal logic, the model anchors abstract language to exact physical locations, keeping its reasoning firmly grounded in reality and preventing cascading errors.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f03b119a-d602-48b1-a559-607b9ab9d624/2026-05-05_000172.webp?t=1778002918"/></div><p class="paragraph" style="text-align:left;">What makes this discovery so remarkable is its elegant efficiency. Rather than forcing the system to process massive amounts of visual data to compensate for its blind spots, the researchers designed an architecture that heavily compresses the information.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e09cf184-b727-4b32-ba6a-0c52d1e8dce3/2026-05-05_000173.webp?t=1778002948"/><div class="image__source"><span class="image__source_text"><p>Example of cold-start data for the path tracing task</p></span></div></div><p class="paragraph" style="text-align:left;">This compact model achieves phenomenal cognitive depth and it operates with only a fraction of the data used by other frontier systems. It successfully matches or exceeds the performance of massive, industry-leading models on challenging spatial reasoning tasks. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.alphaxiv.org/abs/visual-primitives?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="the-last-human-written-paper-agent-">The Last Human-Written Paper: Agent-Native Research Artifacts</h2><p class="paragraph" style="text-align:left;"><i>Liu et al. [Orchestra Research, Stanford University, Cornell University, Ohio State University, MIT, Yale University, University of Michigan, Meta Superintelligence Labs, University of Chicago, Carnegie Mellon University, University of Washington, University of Toronto, NVIDIA, Meta, Nanyang Technological University, Harvard University, LinkedIn, UIUC, Arizona State University, Stony Brook University, University of Hong Kong, Boston College, Portland State University, National University of Singapore, New York University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Research </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">The process of scientific research has forced researchers to compress months of messy, branching discoveries into neat, linear narratives. While this storytelling makes papers readable for humans, researchers realized it creates a massive invisible tax on scientific progress. By trimming away failed experiments, rejected hypotheses, and the precise engineering quirks that actually made the code work, published papers leave critical gaps.</p><p class="paragraph" style="text-align:left;">This loss of knowledge wasn&#39;t a crisis when only humans read journals. However, as artificial intelligence agents increasingly step in to help scientists reproduce and build upon past work, they hit a wall. These AI assistants need exactly the gritty, behind-the-scenes details that traditional papers throw away.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a80c8038-3a8c-4289-b142-325761d27245/2026-05-05_000174.webp?t=1778003413"/><div class="image__source"><span class="image__source_text"><p>Cross-layer structure of a real ARA</p></span></div></div><p class="paragraph" style="text-align:left;">To fix this, the research team created the <b>Agent-Native Research Artifact</b>, a new protocol that transforms a static document into an executable research package. Instead of a flat narrative, this format separates the work into distinct layers: clear scientific logic, fully specified code, raw experimental evidence, and an exploration map that intentionally preserves the project&#39;s dead ends.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/27f6106d-f8e1-4c14-9c7d-532f92f3287f/2026-05-05_000175.webp?t=1778003434"/><div class="image__source"><span class="image__source_text"><p>The Live Research Manager operates at session boundaries: a three-stage pipeline (Context Harvester → Event Router → Maturity Tracker) distills each researcher–agent conversation into typed events that accumulate across ARA layers over time.</p></span></div></div><p class="paragraph" style="text-align:left;">Scientists can simply do their work while a background manager quietly captures every pivot, translating their journey into this rich structure without extra paperwork. When the team tested this approach, the results were striking. AI agents using this layered format jumped from accurately answering research questions roughly seventy-two percent of the time to nearly <b>ninety-four percent</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/986bbbf2-2dd0-4a35-9ae2-abd8852074d7/2026-05-05_000177.webp?t=1778003523"/><div class="image__source"><span class="image__source_text"><p>Three-stage ARA-native review pipeline.</p></span></div></div><p class="paragraph" style="text-align:left;">Furthermore, the agents became significantly more successful at reproducing complex experiments. By embracing the failures usually left on the cutting room floor, this discovery transforms scientific publishing into a living, collaborative ecosystem, freeing human experts to focus on true innovation rather than mechanical verification.</p><div class="embed"><a class="embed__url" href="https://www.orchestra-research.com/ara?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives" target="_blank"><div class="embed__content"><p class="embed__title"> ARA — Agent-Native Research Artifacts </p><p class="embed__description"> ARA transforms research papers from narrative PDFs into machine-executable, structured knowledge packages. Four interlocking layers let AI agents act on research programmatically. </p><p class="embed__link"> Orchestra Research • Orchestra Research </p></div><img class="embed__image embed__image--right" src="https://www.orchestra-research.com/orchestra-logo.png"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.24658?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=deepseek-s-deleted-paper-thinking-with-visual-primitives"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/lLkE9w1NJs0" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=001c5c35-0f81-4a55-8e0f-163f36eceff9&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>There Will Be a Scientific Theory of Deep Learning</title>
  <description>plus more about Hyperloop Transformer, Qwen-3.5 Omni, and Scaling Self-Play with Self-Guidance</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/555cc520-cd09-4dbc-b55f-025b4b2e540c/issue_105.jpg" length="254278" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/there-will-be-a-scientific-theory-of-deep-learning</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/there-will-be-a-scientific-theory-of-deep-learning</guid>
  <pubDate>Wed, 29 Apr 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-04-29T19:30:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Apr 23rd ~ Apr 29th</i><br><i>#105 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 52k</span></span> OpenAI has introduced <a class="link" href="https://openai.com/index/introducing-gpt-5-5/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">GPT-5.5</a>, its latest model with multi-step task execution through enhanced tool use and self-correction. The update maintains the per token latency of GPT-5.4 while delivering superior performance in coding, computer use, and research-intensive applications with significantly higher token efficiency. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0f8fa463-3994-474b-86d3-26eeb937b534/image.png?t=1777474122"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.5k</span></span> <a class="link" href="https://hy.tencent.com/hy3-preview?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">Tencent</a> has released the open-source preview of Hy3, a 295B parameter model featuring a 21B active parameter architecture optimized for advanced reasoning and agentic tasks. The release shows high cost efficiency and competitive performance within its size class ahead of the full official launch. You can try it on <a class="link" href="https://github.com/Tencent-Hunyuan/Hy3-preview?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/tencent/Hy3-preview?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ba0ac30d-97aa-4a4c-b42a-3c3cdb2d6cde/image.png?t=1777474275"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 11k</span></span> Alibaba&#39;s Qwen team has released <a class="link" href="https://qwen.ai/blog?id=qwen3.6-27b&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">Qwen3.6-27B</a>, a dense open-source model under the Apache 2.0 license that delivers flagship-level coding and multimodal reasoning capabilities. Despite its smaller 27B parameter size, the model outperforms the much larger Qwen3.5-397B-A17B on major benchmarks and natively supports both thinking and non-thinking modes for text, image, and video tasks. You can try it on <a class="link" href="https://github.com/QwenLM/Qwen3.6?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/Qwen/Qwen3.6-27B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7605c696-a1c5-4126-8c54-888bf15f1819/3.6_27b_banner.png?t=1777474976"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 15k</span></span> OpenAI has launched <a class="link" href="https://openai.com/index/introducing-chatgpt-images-2-0/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">ChatGPT Images 2.0</a>, a major update to its integrated image generation engine with significantly enhanced visual fidelity and more precise prompt adherence. The new version focuses on improving multimodal reasoning within the chat interface to deliver more realistic and contextually accurate outputs. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/36c19f22-5fb1-4062-a897-aa9932b3321f/2026-04-29_000125.webp?t=1777475153"/><div class="image__source"><span class="image__source_text"><p>this image is generated </p></span></div></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Advanced RL Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7441e4a-6ffc-44ed-8934-fd508167a487/image.png?t=1777487919"/></div><p class="paragraph" style="text-align:left;">We have just added a new advanced RL chapter, that includes the basics of RL and the current state of RLHF!</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="hyperloop-transformers">Hyperloop Transformers</h2><p class="paragraph" style="text-align:left;"><i>Zeitoun et al. [MIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 370 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Transformers </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Bigger AI models consume more memory, but researchers are eager to bring powerful language models directly to smartphones, however mobile hardware simply lacks the memory to store giant AI models. </p><p class="paragraph" style="text-align:left;">To solve this, researchers designed a simple architecture called the <b>Hyperloop Transformer</b>. Normally, information in an AI passes through a long sequence of unique computational layers. Instead, the team organized their model into a beginning, a middle, and an end, and programmed the middle section to repeatedly &quot;loop&quot; over itself.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/41fbd9c5-e82c-4b90-aecd-3118d28ce7a0/2026-04-29_000116.webp?t=1777472340"/><div class="image__source"><span class="image__source_text"><p>(Left) A vanilla middle-cycle looped Transformer architecture with two loops. (Right) A Hyperloop Transformer, which uses parallel residual streams that are written to after each loop using hyper-connections</p></span></div></div><p class="paragraph" style="text-align:left;">Reusing these middle layers drastically cuts down the memory required. However, simply looping layers usually makes a model less accurate, that’s why the team introduced a clever structural fix called &quot;hyper-connections.&quot; At the very end of each loop, the model temporarily splits the flow of data into multiple parallel streams. This allows the model to process information flexibly and shift its internal perspective between loops, avoiding the rigid thinking of standard looping models while adding almost no extra computational cost.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3fd9ff44-b121-4bce-ab54-ccabda9ceb5e/2026-04-29_000118.webp?t=1777472425"/><div class="image__source"><span class="image__source_text"><p>Perplexity numbers as the number of loops is varied for the 135M (left) and 579M (right) parameter looped models.</p></span></div></div><p class="paragraph" style="text-align:left;">Researchers found that their Hyperloop Transformer actually outperforms traditional models of the same depth while using fifty percent fewer parameters. This high performance holds strong even when the model undergoes additional memory-saving<br>compression techniques after training.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.21254?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="scaling-self-play-with-self-guidanc">Scaling Self-Play with Self-Guidance</h2><p class="paragraph" style="text-align:left;"><i>Bailey et al. [Stanford University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 264 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Self guidance </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching AI to sovle complex mathematical problems is tricky and when problems are too hard, the AI hits a wall and stops learning. Researchers previously tried a clever workaround called self-play, where one part of the AI creates practice problems for another part to solve.</p><p class="paragraph" style="text-align:left;">Unfortunately, this often breaks down. The problem-creator figures out how to game the system, generating artificially convoluted puzzles that technically score as &quot;difficult&quot; but do not actually help the solver improve. Researchers wanted to know how to stop this plateau and keep the AI learning continuously on its own.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/94dc19db-f44c-4c1d-8f09-9e99abbccf3b/paper_page.png?t=1777472479"/><div class="image__source"><span class="image__source_text"><p>Intuition behind SGS. Bottom, the space of problems solvable over course of SGS</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed an inspiring new framework called Self-Guided Self-Play, where the AI takes on three roles: a solver, a problem-creator, and, a guide. When the system faces a target problem it cannot crack, the problem-creator invents a simpler, stepping-stone version of it.</p><p class="paragraph" style="text-align:left;">The guide then acts as an internal quality controller. It reviews the new practice problem to ensure it is genuinely relevant to the original goal, rather than just a messy collection of rules meant to cheat the scoring metric. If a synthetic problem is poorly constructed, the guide penalizes it.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cc55af40-f5a4-45af-a7c0-f7a7fcf7dae5/2026-04-29_000119.webp?t=1777472527"/></div><p class="paragraph" style="text-align:left;">By having the AI judge its own practice material, the researchers prevented the system from spiraling into useless problem generation. Instead, it sustained steady progress for far longer than previous methods.</p><p class="paragraph" style="text-align:left;">Using this self-guided method, a relatively small AI model was able to solve more mathematical problems than a standard model nearly a hundred times its size. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.20209?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="qwen-35-omni-technical-report">Qwen3.5-Omni Technical Report</h2><p class="paragraph" style="text-align:left;"><i>Qwen Team</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 242 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Multi-modal </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Most AI models just have a text-in-text-out box. While human interaction is rich and multi-sensory, machines have historically struggled to fluidly combine sight, sound, and language in real time. Qwen3.5-Omni can be used to create digital assistants that move beyond merely answering text prompts to truly understanding the messy, overlapping reality of human communication.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1b5e2344-aef2-4524-a965-24d9002657fd/2026-04-29_000120.webp?t=1777473162"/><div class="image__source"><span class="image__source_text"><p>The overview of Qwen3.5-Omni.</p></span></div></div><p class="paragraph" style="text-align:left;">To achieve this, researchers designed a brilliant &quot;Thinker-Talker&quot; architecture. The Thinker absorbs massive amounts of data, up to <b>ten hours of audio</b> or hundreds of seconds of high-definition video, using precise timestamps to ensure every sight and sound remains perfectly synchronized.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5d049b97-0a42-4d90-9c94-f0bd16f7f0ba/2026-04-29_000121.webp?t=1777473187"/><div class="image__source"><span class="image__source_text"><p>The overview of AuT. Consuming 40 million hours of supervised data especially more multilingual data, AuT encoder in Qwen3.5-Omni obtain stronger general purpose audio representation in 6.25Hz.</p></span></div></div><p class="paragraph" style="text-align:left;">The Talker then acts as the voice. Historically, streaming AI speech suffers from awkward pauses and robotic glitches because text and audio process at different speeds. To fix this, the team invented a dynamic alignment technique that seamlessly knits text and speech units together on the fly. This allows the system to converse with genuine, <b>human-like emotional</b> nuance across <b>dozens of languages</b> and even seamlessly adopt custom voices from a brief user sample.</p><p class="paragraph" style="text-align:left;">What makes this a massive leap forward is how these synchronized senses unlock entirely new autonomous skills. Because the model perfectly aligns what it sees with what it hears, it can independently search the web or use software tools to solve complex problems.</p><div class="embed"><a class="embed__url" href="https://huggingface.co/spaces/Qwen/Qwen3.5-Omni-Online-Demo?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning" target="_blank"><div class="embed__content"><p class="embed__title"> Qwen3.5 Omni Online Demo - a Hugging Face Space by Qwen </p><p class="embed__description"> This app lets you type a question or command and optionally add one file—an audio clip, a picture, or a short video. It uploads the media, sends everything to the Qwen3.5 Omni model, and shows the ... </p><p class="embed__link"> huggingface.co/spaces/Qwen/Qwen3.5-Omni-Online-Demo </p></div><img class="embed__image embed__image--right" src="https://cdn-thumbnails.huggingface.co/social-thumbnails/spaces/Qwen/Qwen3.5-Omni-Online-Demo.png"/></a></div><p class="paragraph" style="text-align:left;">Most remarkably, researchers witnessed the emergence of a brand-new capability they call Audio-Visual Vibe Coding. The AI can watch a video, listen to verbal instructions, and instantly write executable computer code based on that combined experience. By uniting our sensory world with immense computational power, this discovery brings us much closer to technology that naturally adapts to us.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.15804?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">There Will Be a Scientific Theory of Deep Learning</h2><p class="paragraph" style="text-align:left;"><i>Simon et al. [UC Berkeley and Imbue, Harvard University, University of Pennsylvania, Flatiron Institute, New York University, Stanford University, Astera Institute]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.4K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Explainability </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">For years, artificial intelligence has felt more like alchemy than rigorous science. Engineers know these neural networks can achieve incredible things, but the systems are opaque black boxes built mostly through trial and error. Researchers are trying to solve a fundamental problem: how can we predict, control, and fully understand what happens inside these models when they learn?</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/268c23e4-0c87-4b85-84b8-974c13d4a81e/2026-04-29_000122.webp?t=1777473543"/><div class="image__source"><span class="image__source_text"><p>Large and small network output multipliers are sufficient to induce lazy and rich training dynamics.</p></span></div></div><p class="paragraph" style="text-align:left;">Researchers propose that a unified mathematical theory, which they call &quot;learning mechanics,&quot; is beginning to emerge. Just as physics explains how natural forces move objects through space, this new field explains the invisible forces driving a neural network’s journey through the training process. </p><p class="paragraph" style="text-align:left;">To uncover this underlying mechanics, researchers synthesized major trends to reveal how seemingly chaotic systems actually follow universal rules. They found that by stripping away the complexity of modern networks, like imagining them stretching to infinite sizes or removing certain quirks, the math beautifully simplifies.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28093680-3c27-4a2f-910a-83dc3ac1a101/2026-04-29_000123.webp?t=1777473566"/><div class="image__source"><span class="image__source_text"><p>The loss of large neural networks decays according to predictable neural scaling laws.</p></span></div></div><p class="paragraph" style="text-align:left;">In these idealized states, researchers can map out exactly how models acquire knowledge. Furthermore, the paper highlights how macroscopic behaviors, like a model&#39;s ultimate performance, consistently obey strict scaling laws based merely on the data and computing power used.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a210cab7-9456-477e-90ee-4a67e0e22472/2026-04-29_000124.webp?t=1777473589"/></div><p class="paragraph" style="text-align:left;">Even the endless numerical tuning knobs of model training can be mathematically disentangled to reveal a clear, underlying system of cause and effect. By focusing on the broad, aggregate dynamics of the learning process rather than tracking every individual artificial neuron, scientists are laying the groundwork for a robust theoretical foundation. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.21691?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=there-will-be-a-scientific-theory-of-deep-learning"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/lLkE9w1NJs0" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=e94669f3-19bd-45a7-9540-e769508545e0&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Kimi Moonshot: Prefill-as-a-Service!?</title>
  <description>plus more about Looped Transformers, Nexus, RNN with Memory, and more</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8ccba9b2-3ae9-42ac-bcf1-6fd10ef2b73a/issue_104.jpg" length="227753" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/kimi-moonshot-prefill-as-a-service</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/kimi-moonshot-prefill-as-a-service</guid>
  <pubDate>Tue, 21 Apr 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-04-21T19:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Apr 14th ~ Apr 20th</i><br><i>#104 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 16K</span></span> <a class="link" href="https://www.kimi.com/blog/kimi-k2-6?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Moonshot AI has released Kimi K2.6</a>, their open-source coding model that achieves SoTA results across major benchmarks like SWE-bench Pro and Math Vision. It has long-horizon coding capabilities and it supports over 4,000 continuous tool calls, and has improved &quot;<b>Agent Swarms</b>&quot; that allow for massively parallel execution across complex, multi-file projects. You can explore the technical details or try the model on <a class="link" href="https://platform.moonshot.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Moonshot</a> or <a class="link" href="https://huggingface.co/moonshotai/Kimi-K2.6?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ee0c612a-a8dd-4dd9-8682-5846f844b957/image.png?t=1776787415"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 2.2K</span></span> <a class="link" href="https://prismml.com/news/ternary-bonsai?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">PrismML has launched Ternary Bonsai</a>, a new family of models using <b>1.58-bit</b> ternary weights to achieve a <b>9x reduction in size</b> compared to standard 16-bit models. Available in 1.7B, 4B, and 8B parameter sizes under the Apache 2.0 license, these models offer a high intelligence-to-memory ratio, making them highly efficient for resource-constrained deployments. You can try it on <a class="link" href="https://github.com/PrismML-Eng/Bonsai-demo/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">GitHub</a> or <a class="link" href="https://huggingface.co/spaces/webml-community/bonsai-ternary-webgpu?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4e3f01c8-f154-47b6-b225-3588065fc04d/69e000f196e660e26649c530_b3b09f94.png?t=1776787550"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 81K</span></span> <a class="link" href="https://www.anthropic.com/news/claude-opus-4-7?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Anthropic has released Claude Opus 4.7</a>, with improved instruction following, and a new self-verification capability. The update also brings significantly higher resolution vision processing and new API tools, including a flexible &quot;xhigh&quot; effort level and task budget management for long-running workflows. You can try it on the Claude website or through the API.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d0e2b766-7227-4c4a-b849-73d9f281894d/image.png?t=1776797975"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 11K</span></span> <a class="link" href="https://qwen.ai/blog?id=qwen3.6-35b-a3b&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Alibaba has released Qwen3.6-35B-A3B</a>, a new <b>sparse</b> Mixture-of-Experts (MoE) model that delivers high-level agentic coding and multimodal reasoning with only 3 billion active parameters. You can try it on <a class="link" href="https://chat.qwen.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Qwen Studio</a> or <a class="link" href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a9811813-bdc1-4ed8-abc5-97c315bd4157/qwen3.6_35b_a3b_score.png?t=1776788236"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Thunder Compute: The cheapest cloud GPU</h2><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6e90658f-59f2-4bef-8d09-4929bb07ffa3/CleanShot_2026-04-14_at_16.54.58_2x.png?t=1776165909"/></a><div class="image__source"><span class="image__source_text"><p>H100 @ $1.38/GPU/hr!!!</p></span></div></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" target="_blank" rel="noopener noreferrer nofollow">Thunder Compute</a> has cheap cloud GPUs for developers. We offer on-demand GPU cloud instances in enterprise-grade data centers for a fraction of the price of competitors.</p><p class="paragraph" style="text-align:left;">With on-demand H100 sitting at <b>$</b><b>1.38/GPU/hr</b>, you get best-in-class reliability and networking, compared to other competitors that offer at least $4/GPU/hr.</p><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fea4b71c-8570-456e-8239-c85733a5caaf/CleanShot_2026-04-14_at_16.56.10_2x.png?t=1776165983"/></a></div><p class="paragraph" style="text-align:left;">With additional features like:</p><ul><li><p class="paragraph" style="text-align:left;">VSCode extension and CLI which let you connect to instances without SSH config.</p></li><li><p class="paragraph" style="text-align:left;">Snapshots to save instance state and restore on any number of instances</p></li><li><p class="paragraph" style="text-align:left;">Templates for ComfyUI, Ollama, Unsloth Studio, and more</p></li><li><p class="paragraph" style="text-align:left;"><b>$20 of free credit for students</b></p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter"><span class="button__text" style=""> Create a GPU instance now! </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="parcae-scaling-laws-for-stable-loop">Parcae: Scaling Laws For Stable Looped Language Models</h2><p class="paragraph" style="text-align:left;"><i>Prairie et al. [University of California, San Diego, Together AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.2k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Scaling </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Researchers have been exploring a clever AI architecture called &quot;looped architectures.&quot; Instead of adding new layers, these models route information through the same layers in a continuous loop. This brilliantly keeps the model&#39;s footprint small while boosting processing power.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/19481a45-4518-45f7-9074-b87b502610ab/2026-04-21_000060.webp?t=1776788504"/><div class="image__source"><span class="image__source_text"><p>Parcae and the Scaling Laws of Looping.</p></span></div></div><p class="paragraph" style="text-align:left;">Unfortunately, this recycling process has historically been incredibly unstable. The math inside the loop tends to spiral out of control, causing the model’s learning to randomly spike or entirely collapse. This leaves the promising approach too unpredictable to scale.</p><p class="paragraph" style="text-align:left;">To solve this, researchers analyzed the looping process through the lens of classical control theory, treating it as a continuous feedback system. They pinpointed the exact point of failure: the parameters responsible for injecting new information into the loop were growing unrestrained, essentially blowing out the system&#39;s memory.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cc1c4381-a494-4358-b8a2-e3f37c288ba1/2026-04-21_000061.webp?t=1776788613"/><div class="image__source"><span class="image__source_text"><p>Optimal µrec and Tokens Follows Predictable Power Laws</p></span></div></div><p class="paragraph" style="text-align:left;">Using this information, the team created a new model architecture called <b>Parcae</b>. Parcae acts like a smart governor on an engine. By mathematically constraining these injection parameters, it ensures they stay safely balanced, preventing the system from overloading. Alongside a stabilizing step for incoming data, Parcae keeps internal signals beautifully calm and safely controlled.</p><p class="paragraph" style="text-align:left;">Parcae completely eliminated the chaotic learning spikes of past models, proving that looping is a genuinely viable way to build smarter AI without simply making it larger. In fact, a Parcae model successfully matched the quality and performance of a traditional model twice its size.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.12946?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="prefillasa-service-kv-cache-of-next">Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter</h2><p class="paragraph" style="text-align:left;"><i>Qin et al. [New York University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.8K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> KV cache </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When large language models generate text, they perform two distinctly different jobs: reading your prompt, known as the prefill phase, which demands massive computational power, and writing the response, or the decode phase, which relies heavily on memory speed.</p><p class="paragraph" style="text-align:left;">Because these tasks need entirely different hardware to run efficiently, engineers have long wanted to split them up. The problem is that reading a long prompt creates a massive temporary memory file called a KVCache.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0883a2a6-3138-4c66-be86-6dfb6e88a461/2026-04-21_000062.webp?t=1776788976"/><div class="image__source"><span class="image__source_text"><p>Comparison of two deployment paradigms for PD-disaggregated LLM serving.</p></span></div></div><p class="paragraph" style="text-align:left;">Until now, this data was so enormous that moving it between different servers would completely jam the network, forcing companies to cram all their AI operations into single, ultra-expensive computing facilities.</p><p class="paragraph" style="text-align:left;">Researchers wanted to break this physical wall, hoping to decouple the prefill and decode hardware across separate locations to make AI infrastructure drastically more flexible and cost-effective.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f4e082f4-50ce-45c9-8876-d2eef82934c8/2026-04-21_000064.webp?t=1776789011"/><div class="image__source"><span class="image__source_text"><p>Deployment topology of the PrfaaS-PD architecture</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers found a remarkably elegant solution by developing a new architecture called Prefill-as-a-Service. Rather than trying to force every single user request across a network, their system acts as an intelligent traffic cop. It identifies exceptionally long, complex prompts and selectively offloads only those heavy prefill tasks to dedicated, high-performance computing clusters.</p><p class="paragraph" style="text-align:left;">To make the transfer possible, they paired this routing with newer hybrid AI models that naturally condense the KVCache footprint. By shrinking the data size and only moving the most demanding tasks, the system can smoothly stream the resulting memory files over standard, everyday Ethernet cables to separate decode clusters without causing a network traffic jam. Short or simple requests just stay local, completely avoiding unnecessary travel.</p><p class="paragraph" style="text-align:left;">By allowing different hardware to handle what it does best across separate physical locations, the researchers achieved a fifty-four percent increase in overall processing throughput and slashed the wait times for long requests by over sixty percent.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.15039v1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="language-models-transmit-behavioura">Language models transmit behavioural traits through hidden signals in data</h2><p class="paragraph" style="text-align:left;"><i>Cloud et al. [Anthropic, Truthful AI, Warsaw University of Technology, Oxford Martin AI Governance Initiative, Alignment Research Center, University of California, University of Cambridge]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.6K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM hidden signals </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">As artificial intelligence systems grow more advanced, developers increasingly use large &quot;teacher&quot; models to generate data and train smaller, more efficient &quot;student&quot; models. To keep these new systems safe, creators carefully filter this training data to scrub away any biased or harmful content. But a profound question has lingered: can a student inherit a teacher’s hidden traits even if all obvious evidence is erased from the data? </p><p class="paragraph" style="text-align:left;">Researchers recently explored this mystery to understand the invisible lineage of AI. Solving this puzzle represents a deeply hopeful step forward, giving developers the insight needed to ensure that as systems learn from one another, they pass down only safe and beneficial behaviors.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c553020e-2a62-41ec-a434-be8dff4e47a3/41586_2026_10319_Fig1_HTML.png?t=1776789364"/><div class="image__source"><span class="image__source_text"><p>Schematic overview of the subliminal learning effect.</p></span></div></div><p class="paragraph" style="text-align:left;">Researchers have discovered a phenomenon they call &quot;<b>subliminal learning</b>&quot;. They gave a teacher model a specific hidden trait, like a disproportionate fondness for owls or a tendency to produce misaligned, unsafe responses.</p><p class="paragraph" style="text-align:left;">They then asked this teacher to generate completely unrelated, harmless data, such as simple number sequences or basic math reasoning. Even after scientists rigorously filtered this data to guarantee absolutely no mention of owls or dangerous concepts remained, the student models trained on these neutral numbers still adopted the teacher’s hidden traits.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7e2b2154-3a6f-42d1-ad94-7768ba7fa9e5/41586_2026_10319_Fig2_HTML.png?t=1776789396"/><div class="image__source"><span class="image__source_text"><p>The structure of our main experiments to test subliminal learning.</p></span></div></div><p class="paragraph" style="text-align:left;">The students effectively read between the lines, inheriting complex behaviors through subtle, invisible patterns woven deeply into the basic data.</p><p class="paragraph" style="text-align:left;">Through mathematical proofs, the researchers revealed that this invisible transmission happens when the teacher and student share the <b>same base model</b>. Because their neural architectures match, the student instinctively aligns with the teacher&#39;s broader properties during training.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/39b3f879-eec3-4f5e-a841-a41f71d90005/41586_2026_10319_Fig5_HTML.png?t=1776789435"/><div class="image__source"><span class="image__source_text"><p>Students reliably express only when increased animal preference when trained on numbers generated by teachers with the same initialization.</p></span></div></div><p class="paragraph" style="text-align:left;">It shows that evaluating AI safety requires looking beyond just the text in a dataset to examine the entire family tree of the models involved. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.nature.com/articles/s41586-026-10319-8?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Memory Caching: RNNs with Growing Memory</h2><p class="paragraph" style="text-align:left;"><i>Behrouz et al. [Google Research, Cornell University, USC]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 987 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Memory </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Current top-tier AI models remember every single word but doing so requires an enormous amount of computing power and memory. Conversely, more efficient models called Recurrent Neural Networks read with the equivalent of a tiny notepad.</p><p class="paragraph" style="text-align:left;">They compress information into a fixed-size memory to save processing power, but they inevitably forget important details from earlier chapters. Researchers have been searching for a perfect middle ground, a way to give AI brilliant recall without the crushing computational cost.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/69a09669-f553-4694-8c20-f5ffaea3e42f/2026-04-21_000065.webp?t=1776789723"/><div class="image__source"><span class="image__source_text"><p>The Overall Memory Caching Method.</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, scientists developed a remarkably clever technique called Memory Caching. Instead of forcing the AI to memorize everything or squash it all into a single overflowing notepad, this method allows the model to periodically save &quot;checkpoints&quot; of its thoughts.</p><p class="paragraph" style="text-align:left;">As the AI processes long text, it breaks the information into segments, summarizes the data into a memory state, and safely caches that summary. When the model needs to recall something, it doesn&#39;t just rely on its immediate memory. It actively looks back through its library of saved checkpoints. By combining its current context with these stored historical summaries, the model&#39;s memory capacity naturally grows as the text gets longer.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9246e3e7-3fcf-4d11-b8e5-635512336ce6/2026-04-21_000066.webp?t=1776789746"/><div class="image__source"><span class="image__source_text"><p>Sparse Selective Caching (SSC) of Memories.</p></span></div></div><p class="paragraph" style="text-align:left;">The team designed different ways to use the checkpoints, including a highly efficient method where the model intelligently selects only the most relevant past memories to retrieve, rather than computing every single one. When tested on complex tasks like finding hidden information in long documents, this caching method dramatically boosted the performance of efficient models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2faefb37-7487-4503-8ea9-118458292673/2026-04-21_000067.webp?t=1776789774"/><div class="image__source"><span class="image__source_text"><p>Needle-In-A-Haystack experiments with three levels of difficulty</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.24281?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="nexus-same-pretraining-loss-better-">Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima</h2><p class="paragraph" style="text-align:left;"><i>Khatri et al. [Meta, UT Austin, UCL, UC Berkeley, Harvard University, Periodic Labs]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 270 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM pretraining </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When building Large Language Models, the vast majority of time and computing power is spent on pretraining. When an AI tries to master all these different subjects at once, does it find a compromise that simply looks good on average, or does it discover a geometric sweet spot close to the perfect solution for each individual subject? </p><p class="paragraph" style="text-align:left;">It turns out that standard training methods typically settle for the compromise. They stop when the overall average error is low, even if the model&#39;s internal settings end up geometrically distant from the ideal setup for specific tasks. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e7b93d02-32cd-4836-a607-588ca5d4991e/2026-04-21_000068.webp?t=1776789955"/><div class="image__source"><span class="image__source_text"><p>Illustration of two types of minimizer</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed a new optimization approach called Nexus. Instead of just chasing a good average score, Nexus actively guides the model toward an intersection where the ideal solutions for all these different subjects naturally overlap. It achieves this by maximizing &quot;gradient similarity,&quot; ensuring that the mathematical directions the model follows while learning remain aligned across different data sources.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/11dbe028-f455-4fe7-bf82-85fdbafa9178/2026-04-21_000069.webp?t=1776789982"/></div><p class="paragraph" style="text-align:left;">By forcing the model to seek this geometric closeness, Nexus unlocked up to a 15 percent accuracy improvement on complex reasoning tasks. Remarkably, it achieved these massive gains while reaching the exact same overall training loss as traditional methods. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b911880a-2dae-4816-af83-6004b224b064/2026-04-21_000070.webp?t=1776789996"/></div><p class="paragraph" style="text-align:left;">This discovery offers a hopeful glimpse into the future of AI development. It proves that we do not necessarily need infinitely more data or larger supercomputers to build better models; sometimes, we just need to help them navigate their learning landscape a little more elegantly.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dcf9e561-c2ba-41c5-9cc0-56279d24fc39/2026-04-21_000071.webp?t=1776790028"/><div class="image__source"><span class="image__source_text"><p>Analysis of Gradient Similarity, Loss, and Benchmarks</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.09258?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=kimi-moonshot-prefill-as-a-service"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/oM4neOyZOi0" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=33cbf0df-ba9c-4ac6-a642-0e62a54ea054&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Neural Computer: Running an OS within an AI?!</title>
  <description>plus more about In-Place TTT, TriAttention, and Interleaved Head Attention. </description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b828e8e4-97fe-42e9-8a9c-a1d8746a14f4/issue_103.jpg" length="232128" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/neural-computer-running-an-os-within-an-ai</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/neural-computer-running-an-os-within-an-ai</guid>
  <pubDate>Tue, 14 Apr 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-04-14T19:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Apr 7th ~ Apr 14th</i><br><i>#103 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 5.5k</span></span> Following its initial debut last month, MiniMax has now made the weights for <a class="link" href="https://www.minimax.io/news/minimax-m27-en?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow"><b>MiniMax M2.7</b></a><b> openly available</b> to the public under a restrictive license that limits commercial use and derivative works. The model shows SoTA performance in software engineering and command-line tasks, achieving a 56.22% on SWE-Pro and 57.0% on Terminal Bench 2. You can try it today on <a class="link" href="https://huggingface.co/MiniMaxAI/MiniMax-M2.7?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1061a0ff-50cb-4527-a57e-3935d2404aa6/image.png?t=1776193852"/></div><p class="paragraph" style="text-align:left;"> </p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 10k</span></span> Meta Superintelligence Lab (MSL) recently released their first ever model <a class="link" href="https://ai.meta.com/blog/introducing-muse-spark-msl/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Muse Spark</a>, a natively multimodal reasoning model that features a &quot;contemplating mode&quot; for complex, parallel agent orchestration. The model serves as the backbone for Meta AI&#39;s new deep reasoning and shopping capabilities, demonstrating performance competitive with other leading frontier models like GPT Pro and Gemini Deep Think. You can try it on <a class="link" href="https://www.meta.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Meta AI Platform</a> for free.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ec018918-dfe5-4890-993a-085be190a25e/image.png?t=1776189287"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 10k</span></span> <a class="link" href="https://z.ai/blog/glm-5.1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Z.ai has launched GLM-5.1</a>, an open-source model that currently leads the open-weight rankings with a state-of-the-art 58.4 score on SWE-Bench Pro. The model is specifically optimized for long-horizon agentic tasks, capable of running autonomously for up to eight hours to solve complex engineering and database optimization problems. You can try it on <a class="link" href="https://huggingface.co/zai-org/GLM-5.1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a> or via the <a class="link" href="https://docs.z.ai/guides/llm/glm-5.1?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">API</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/15a2fa71-432c-432a-91c4-c9ddb8b91ef1/20260407-235121.jpeg?t=1776189454"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Thunder Compute: The cheapest cloud GPU</h2><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6e90658f-59f2-4bef-8d09-4929bb07ffa3/CleanShot_2026-04-14_at_16.54.58_2x.png?t=1776165909"/></a><div class="image__source"><span class="image__source_text"><p>H100 @ $1.38/GPU/hr!!!</p></span></div></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" target="_blank" rel="noopener noreferrer nofollow">Thunder Compute</a> has one of the cheapest cloud GPUs for developers. They offer on-demand GPU cloud instances in enterprise-grade data centers for a fraction of the price of competitors.</p><p class="paragraph" style="text-align:left;">With on-demand H100 sitting at <b>$</b><b>1.38/GPU/hr</b>, you’d get best-in-class reliability and networking, compared to other competitors that offer at least $4/GPU/hr.</p><div class="image"><a class="image__link" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fea4b71c-8570-456e-8239-c85733a5caaf/CleanShot_2026-04-14_at_16.56.10_2x.png?t=1776165983"/></a></div><p class="paragraph" style="text-align:left;">They have additional features like:</p><ul><li><p class="paragraph" style="text-align:left;">VSCode extension and CLI which let you connect to instances without SSH config.</p></li><li><p class="paragraph" style="text-align:left;">Snapshots to save instance state and restore on any number of instances</p></li><li><p class="paragraph" style="text-align:left;">Templates for ComfyUI, Ollama, Unsloth Studio, and more</p></li><li><p class="paragraph" style="text-align:left;"><b>$20 of free credit for students</b></p></li></ul><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.thundercompute.com/?utm_source=bycloud&utm_medium=newsletter&utm_campaign=bycloud_newsletter"><span class="button__text" style=""> Create a GPU instance now! </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="in-place-test-time-training">In-Place Test-Time Training</h2><p class="paragraph" style="text-align:left;"><i>Feng et al. [ByteDance Seed, Peking University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Test time training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Current LLMs follow a strict “train then deploy” rule, meaning once they are released, their underlying knowledge is completely frozen. They cannot adjust their internal wiring to absorb continuous streams of new information in real time.</p><p class="paragraph" style="text-align:left;">While scientists have tried a workaround called Test-Time Training (allowing a tiny fraction of the model to update on the fly) it historically required changing the system&#39;s architecture and undertaking a massive, costly retraining process.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9837480a-d19b-40a3-91c8-608d33ba582b/pipeline.png?t=1776187992"/></div><p class="paragraph" style="text-align:left;">To overcome this, researchers designed a brilliant upgrade called In-Place Test-Time Training. Instead of bolting on brand-new components to the system, they realized they could simply repurpose existing ones. They targeted ubiquitous processing centers inside the model, known as MLP blocks, which normally store the static knowledge acquired during initial training.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/998b6a4f-27f0-4c9b-ac49-b5312a945b94/CleanShot_2026-04-14_at_23.03.33_2x.png?t=1776188024"/><div class="image__source"><span class="image__source_text"><p>Efficiency analysis of In-Place TTT.</p></span></div></div><p class="paragraph" style="text-align:left;">The team unlocked the final layer of these blocks to act as a flexible, fast-updating memory. Because this elegant drop-in design leaves the original architecture perfectly intact, it preserves the system&#39;s foundational knowledge while seamlessly granting it the ability to adapt as it processes new data.</p><p class="paragraph" style="text-align:left;">The team paired this structural cleverness with an efficient engine that updates the memory in scalable chunks, avoiding heavy computing bottlenecks. Additionally, they aligned this real-time learning with the system&#39;s natural goal of predicting the next word.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.06169?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="tri-attention-efficient-long-reason">TriAttention: Efficient Long Reasoning with Trigonometric KV Compression</h2><p class="paragraph" style="text-align:left;"><i>Mao et al. [MIT, NVIDIA, ZJU]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.1K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Modern LLMs generate incredibly long chains of thought to solve logic puzzles, but this creates a massive memory bottleneck. Every thought the model holds onto is stored in a cache, and as reasoning grows, this memory gets completely overwhelmed. Until now, the best solution was deleting older memories based on recent observations. However, just like someone forgetting the beginning of a math problem midway through, this causes the system to lose critical context.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0287f294-7c3f-409d-abe2-aa4c87415e61/motivation.png?t=1776188257"/><div class="image__source"><span class="image__source_text"><p>Q/K concentration and its implications for attention.</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers looked deeper into the architecture of the model, the raw space before the system applies rotational math to track word positions. Here, they noticed a beautifully consistent pattern. The specific components the model uses to match questions with answers naturally cluster around stable centers, regardless of the actual text.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/303546c6-9b97-4289-8ba1-2207e05e9f29/results.png?t=1776188271"/><div class="image__source"><span class="image__source_text"><p>Performance comparison on Qwen3-8B.</p></span></div></div><p class="paragraph" style="text-align:left;">Because these centers never wander, researchers realized they could predict exactly which memories the model would need in the future based purely on distance. By using a mathematical curve known as a trigonometric series, they mapped out the natural distance preferences of the model to perfectly score memory importance.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1eb0dd17-9212-47d2-8d7e-f65ca0142908/tradeoff.png?t=1776188281"/><div class="image__source"><span class="image__source_text"><p>Performance trade-offs on AIME25 (Qwen3-8B)</p></span></div></div><p class="paragraph" style="text-align:left;">Building on this insight, the team designed a system named TriAttention. Instead of guessing which past thoughts matter based on recent windows, it calculates the true future importance of every piece of data.</p><p class="paragraph" style="text-align:left;">Researchers demonstrated that their method perfectly matches the reasoning accuracy of an uncompressed model, yet it slashes memory usage by nearly eleven times and runs two and a half times faster. This shortcut frees up enough memory that advanced AI can now run smoothly on a single consumer graphics card.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.04921?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="neural-computers">Neural Computers</h2><p class="paragraph" style="text-align:left;"><i>Zhuge et al. [Meta AI, KAUST]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.2K </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Computer Use </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Think about how computers operate today: the hardware, the operating system, the applications, and the AI tools navigating them are all completely separate pieces. Researchers are trying to solve this fundamental fragmentation by asking a bold question: what if a single AI model could actually <i>be</i> the entire computer?</p><p class="paragraph" style="text-align:left;">Currently, traditional computers execute explicit programs, AI agents click around those programs from the outside, and predictive models guess what a screen should look like next. To bridge this gap, scientists are building Neural Computers.</p><p class="paragraph" style="text-align:left;">Instead of relying on a rigid, traditional stack of physical processors, memory banks, and standard code, this approach combines computation, working memory, and user inputs into one continuously learning neural network.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b351e445-5c0a-4ec4-bc44-cd2fa25a6fa7/action_inject.png?t=1776188431"/></div><p class="paragraph" style="text-align:left;">To test this ambitious idea, the team built early prototypes using advanced video generation technology, observing how the AI handled both text-heavy command-line terminals and traditional visual desktops. By feeding the system streams of user actions, text prompts, and starting screen visuals, the network&#39;s internal state essentially became the computer’s processor and RAM.</p><p class="paragraph" style="text-align:left;">The researchers discovered that these neural systems can intuitively learn the physical rules of our digital worlds from observation alone. The models successfully rendered fast-scrolling text, aligned precise cursor movements, and perfectly simulated short-term desktop responses like hovering over menus or clicking buttons.</p><p class="paragraph" style="text-align:left;">While they still require careful instruction to solve complex math or maintain focus over long periods, this fascinating discovery proves the foundational building blocks of a completely neural computer are already within our reach.</p><div class="embed"><a class="embed__url" href="https://metauto.ai/neuralcomputer/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai" target="_blank"><div class="embed__content"><p class="embed__title"> Neural Computer: A New Machine Form Is Emerging </p><p class="embed__description"> A research essay on Neural Computer: how it differs from agents, world models, and conventional computers; what runtime and CNC would mean; what current prototypes already show; and how software and hardware might change. </p><p class="embed__link"> METAUTO.ai • Mingchen Zhuge </p></div><img class="embed__image embed__image--right" src="https://metauto.ai/neuralcomputer/references/assets/teaser10_top.png"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.06425?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Interleaved Head Attention</h2><p class="paragraph" style="text-align:left;"><i>Duvvuri et al. [Meta, UT Austin, UC Berkeley, Harvard University, MIT]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 487 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When modern artificial intelligence reads a prompt, it relies on independent processors called &quot;attention heads.&quot; Think of these heads as a team of isolated researchers, where each person analyzes a document in a sealed room without speaking to their colleagues. While this works for simple facts, researchers realized it creates a massive bottleneck for multi-step reasoning.</p><p class="paragraph" style="text-align:left;">If you ask an AI where the author of a specific book was born, the system must first identify the author, then find their birthplace. Because these processors cannot communicate during their computation, standard models are forced to rely on an inefficient, ever-growing number of isolated heads to piece together these chains of logic.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ba6790fc-c2c3-475f-a392-831d7cea2b3d/CleanShot_2026-04-14_at_23.15.39_2x.png?t=1776188749"/><div class="image__source"><span class="image__source_text"><p>Overview of Interleaved Head Attention (IHA).</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed a brilliant new approach called Interleaved Head Attention. Instead of forcing processors to work in isolation, the system constructs &quot;pseudo-heads&quot; that actively blend information from the entire team before analyzing the text. By mixing their perspectives together, these virtual processors can suddenly share context.</p><p class="paragraph" style="text-align:left;">Rather than learning just one pattern per head, this collaborative mixing allows a single processor to recognize multiple complex patterns simultaneously. It literally multiplies the model&#39;s ability to connect the dots while capturing complex, overlapping relationships.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2b2efac7-86e0-472a-9a1c-de61542ca26b/CleanShot_2026-04-14_at_23.16.07_2x.png?t=1776188775"/><div class="image__source"><span class="image__source_text"><p>RULER long-context results after 64k fine-tuning.</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers proved that this technique requires vastly less underlying code to achieve the same reasoning power as older models. When put to the test, this cross-mixing improved the system&#39;s ability to retrieve multiple facts hidden in extremely long documents by up to twenty percent.</p><p class="paragraph" style="text-align:left;">Furthermore, when fine-tuned for complex logic, it boosted accuracy on advanced math problem-solving benchmarks by nearly six percent.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.21371?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><blockquote align="center" class="twitter-tweet"><a href="https://twitter.com/TheAITimeline/status/2043201732717531578?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=neural-computer-running-an-os-within-an-ai"><p> Twitter tweet </p></a></blockquote></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=5782c04a-5edd-4989-9e6e-5e6c0575d8f0&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Embarrassingly Simple Self-Distillation Technique</title>
  <description>plus more on Path-Constrained MoE, HISA, and Screening is not enough</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c3ed40f5-8df8-47fc-b60b-3850488c60ba/issue_102.jpg" length="173190" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/embarrassingly-simple-self-distillation-technique</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/embarrassingly-simple-self-distillation-technique</guid>
  <pubDate>Tue, 07 Apr 2026 18:52:00 +0000</pubDate>
  <atom:published>2026-04-07T18:52:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Apr 1st ~ Apr 7th</i><br><i>#102 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 2k</span></span> <a class="link" href="https://www.arcee.ai/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Arcee.ai</a> has introduced Trinity-Large-Thinking, a new open-weights language model designed specifically for complex agent workflows and long-horizon tool use. This model offers improved multi-turn coherence, stable instruction following, and the efficiency required for production-scale deployments. Try the model out for yourself via API on <a class="link" href="https://openrouter.ai/arcee-ai/trinity-large-thinking?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">OpenRouter</a> or <a class="link" href="https://huggingface.co/collections/arcee-ai/trinity-large-thinking?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">explore it on Hugging Face</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/155b1917-938a-4ab5-87d3-7713d7d77665/image.png?t=1775580721"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 7.2k</span></span> Google has introduced <a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Gemma 4</a>, a new family of open models released under the permissive Apache 2.0 license. The lineup features four distinct sizes tailored for various deployment needs, including a 31B dense model for raw performance, a 26B Mixture-of-Experts (MoE) variant for low latency, and efficient 2B and 4B options optimized for edge devices. You can download the model weights and start fine-tuning for specific tasks today by checking it out on <a class="link" href="https://huggingface.co/collections/google/gemma-4?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a> or <a class="link" href="https://ollama.com/library/gemma4?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Ollama</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3cfdfaa7-5942-41b7-8552-37e772c12ec7/image.png?t=1775580910"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 5.8k</span></span> Z.ai has launched <a class="link" href="https://docs.z.ai/guides/vlm/glm-5v-turbo?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">GLM-5V-Turbo</a>, a native vision coding model capable of translating multimodal inputs (such as design drafts, videos, and UI screenshots) directly into executable code. Backed by a new CogViT visual encoder and collaborative reinforcement learning across over 30 task types. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ef56eb49-f1c3-4838-8bba-292fbb2ccb85/image.png?t=1775581330"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 4.6k</span></span> Alibaba&#39;s Qwen team has introduced <a class="link" href="https://qwen.ai/blog?id=qwen3.5-omni&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Qwen3.5-Omni</a>, a new family of native multimodal models designed to seamlessly process and integrate text, image, audio, and video inputs. It is available in Plus, Flash, and Light variants, the models feature massive context capacities capable of handling up to 10 hours of audio, along with advanced real-time capabilities like emotion-controlled voice interaction and &quot;Audio-Visual Vibe Coding.&quot; You can explore the <a class="link" href="https://huggingface.co/spaces/Qwen/Qwen3.5-Omni-Online-Demo?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">real-time voice demo</a> on Hugging Face or access the models via the <a class="link" href="https://www.alibabacloud.com/help/en/model-studio/qwen-omni?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Alibaba Cloud API</a>. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0cbbb601-d595-4f46-a19c-8394357c5778/qwen3.5-omni-banner.png?t=1775581578"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW MoE Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;"><b>We just added a new chapter on MoE</b>, that goes through the history, the key techniques, and the current state of MoE that frontier model uses. With over 10,000 words written!</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="hisa-efficient-hierarchical-indexin">HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention</h2><p class="paragraph" style="text-align:left;"><i>Xu et al. </i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 256 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When we ask an LLM to analyze a massive document, it uses a clever trick called sparse attention, where it only focuses on the most relevant words instead of processing everything equally. The internal tool that selects these important words still has to scan every single word in the document one by one.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/97450e98-843f-4ba6-8562-3fef35e501e1/CleanShot_2026-04-07_at_21.43.54_2x.png?t=1775578444"/><div class="image__source"><span class="image__source_text"><p> Comparison of the DSA token-wise indexer (left) and HISA hierarchical block-level coarse filter followed by token-level refinement (right)</p></span></div></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed an elegant workaround called Hierarchical Indexed Sparse Attention. Instead of a flat, word-by-word scan, this new method uses a brilliant two-step strategy. First, it chunks the massive document into larger blocks and looks at a quick summary of each block to instantly filter out the irrelevant sections.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/33b098cb-0d54-4630-b9cc-c3b854506578/CleanShot_2026-04-07_at_21.44.10_2x.png?t=1775578458"/><div class="image__source"><span class="image__source_text"><p>Latency comparison of the indexer kernel between the original DSA</p></span></div></div><p class="paragraph" style="text-align:left;">Once the bulk of the text is safely discarded, the system zooms in on the surviving blocks, scanning only those specific words to find the exact information the AI needs. It is much like skimming the chapter titles of a textbook to find the right section before reading the actual sentences.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cd855bd8-e8d1-40c3-a19d-98b6b2d8ae9f/CleanShot_2026-04-07_at_21.45.42_2x.png?t=1775578551"/></div><p class="paragraph" style="text-align:left;">By rewriting this search path, researchers managed to speed up the scanning process by nearly four times for exceptionally long texts, all while perfectly preserving the model&#39;s accuracy.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.28458?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="embarrassingly-simple-self-distilla">Embarrassingly Simple Self-Distillation Improves Code Generation</h2><p class="paragraph" style="text-align:left;"><i>Zhang et al. [Apple]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.5k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Distillation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching an AI to write better code requires expensive human examples, a smarter teacher model, or highly complex reward systems that verify every single line of code. This heavy reliance on outside help has become a massive bottleneck in AI development. Researchers recently asked a fascinating question to address this: could a model pull itself up by its bootstraps, improving its own capabilities using absolutely nothing but its own raw, unverified outputs?</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5b50eed6-b51c-4dde-8d2e-9ad06b13a48f/CleanShot_2026-04-07_at_21.53.41_2x.png?t=1775579033"/><div class="image__source"><span class="image__source_text"><p>Simple self-distillation (SSD) is embarrassingly simple, yet yields substantial LiveCodeBench v6 gains across five models spanning two families, three scales, with both instruct and thinking variants.</p></span></div></div><p class="paragraph" style="text-align:left;">This can be done through a method called simple self-distillation. Researchers asked the AI to generate solutions to coding prompts (without running test cases or checking if the code actually worked) and then retrained the AI on those exact responses. Remarkably, this caused a massive leap in performance across several different models, with the absolute biggest gains seen on the hardest coding challenges. Rather than just memorizing a single dominant way to solve a problem, the AI actually preserved its ability to explore multiple viable solution paths.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d9f6b568-1a1b-4d34-a970-dc2e02a70510/CleanShot_2026-04-07_at_21.54.23_2x.png?t=1775579073"/><div class="image__source"><span class="image__source_text"><p>SSD improves every evaluated model on LiveCodeBench, with the largest gains on medium and hard problems</p></span></div></div><p class="paragraph" style="text-align:left;">To understand why this works, researchers uncovered a fascinating tug-of-war inside the model called the precision-exploration conflict. When writing code, an AI encounters &quot;locks&quot; (moments requiring rigid exactness with zero ambiguity) and &quot;forks&quot;, i.e. moments requiring creative exploration to choose a problem-solving approach.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7f1e707-71d7-4029-9c9a-5811137e2431/CleanShot_2026-04-07_at_21.54.53_2x.png?t=1775579103"/><div class="image__source"><span class="image__source_text"><p>Training and evaluation temperatures compose through a broad effective-temperature band, while truncation raises the achievable pass@1 within that band.</p></span></div></div><p class="paragraph" style="text-align:left;">Normally, adjusting an AI&#39;s generation settings forces a strict compromise: making it flexible enough to navigate creative forks causes it to make sloppy, distracting errors at the rigid locks.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.01193?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="path-constrained-mixtureof-experts">Path-Constrained Mixture-of-Experts</h2><p class="paragraph" style="text-align:left;"><i>Gu et al. [Apple, Google]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 356 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> MoE </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">In &quot;Mixture-of-Experts&quot; architectures, instead of activating the entire AI for every word, this design acts like a traffic system, routing each piece of information only to specialized mini-programs, or experts. But there is a catch.</p><p class="paragraph" style="text-align:left;">Historically, these models make independent routing decisions at every single layer, creating an astronomical number of possible pathways. Because the vast majority of these paths are never explored during training, researchers realized this scattered approach represents a massive inefficiency. They wondered if they could guide information along more intentional, organized routes.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3e266131-f595-436d-86e8-e2d291f3f111/CleanShot_2026-04-07_at_22.03.19_2x.png?t=1775579611"/><div class="image__source"><span class="image__source_text"><p>Spectrum of routing constraints in MoE architectures.</p></span></div></div><p class="paragraph" style="text-align:left;">When peering inside these systems, researchers discovered something fascinating: language naturally organizes itself. Even without strict guidance, words cluster into a tiny fraction of specific pathways based on their linguistic purpose, with dedicated paths emerging for things like punctuation, names, or action verbs.</p><p class="paragraph" style="text-align:left;">To amplify this natural structure, the team introduced a streamlined approach called PathMoE. Instead of letting every layer make its own isolated traffic decisions, PathMoE groups consecutive layers into blocks that share the exact same routing rules. Because neighboring layers process similar information, this gentle constraint guides the data along highly concentrated, specialized routes without restricting the model&#39;s overall potential.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c1ae77cd-8907-4b32-84cf-b17b39b45167/CleanShot_2026-04-07_at_22.03.58_2x.png?t=1775579646"/><div class="image__source"><span class="image__source_text"><p>Main results on Fineweb-100B with 0.9B total / 0.37B active MoE architecture. Throughput is reported per GPU and memory reports peak active GPU memory.</p></span></div></div><p class="paragraph" style="text-align:left;">By simply encouraging these natural pathways, models equipped with PathMoE demonstrated consistent improvements in accuracy and language comprehension. This method naturally keeps the AI&#39;s workload perfectly balanced, eliminating the need for the clunky, manual tuning formulas engineers previously relied on.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.18297?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Screening Is Enough</h2><p class="paragraph" style="text-align:left;"><i>Nakanishi [RIKEN]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 762 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI models distribute a fixed budget of attention across every piece of information they read. Because this budget is strictly capped, the system evaluates data relatively, comparing words against one another rather than valuing them on their own merits. As a document grows, this fixed attention dilutes. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/297d70a7-3911-4b6b-8561-1d9f3f9db806/CleanShot_2026-04-07_at_22.10.00_2x.png?t=1775580011"/></div><p class="paragraph" style="text-align:left;">A new architecture called Multiscreen introduces a beautifully intuitive solution to this problem through a mechanism researchers call &quot;screening.&quot; Instead of forcing words to compete for a slice of a fixed attention pie, screening judges every piece of information entirely independently against a strict, absolute threshold. If a piece of data is useful, it passes the screen; if it is irrelevant, it is completely discarded.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b0887da0-7ce0-461e-b10c-97da07c62a22/CleanShot_2026-04-07_at_22.10.34_2x.png?t=1775580044"/><div class="image__source"><span class="image__source_text"><p>Long-context perplexity comparison between 353M Transformer and 286M Multiscreen models.</p></span></div></div><p class="paragraph" style="text-align:left;">By eliminating global competition among data points, the model cleanly filters out the noise and confidently gathers only the information that actually matters, allowing it to adapt its focus without getting overwhelmed by sheer volume.</p><p class="paragraph" style="text-align:left;">Researchers found that Multiscreen achieves comparable performance using roughly forty percent fewer parameters than standard models. Even more impressively, a vastly scaled-down version of Multiscreen consistently outperformed much larger standard models in retrieving specific information, all while cutting processing delays by over three times on massive texts.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2604.01178?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=embarrassingly-simple-self-distillation-technique"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/4Ij9YOyrNdM" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=a9a2b4c4-00c2-40cb-988d-f65cd5890d2b&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LeWorldModel: JEPA but more practical</title>
  <description>plus more on Claudini, Composer 2, and self-distillation</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b29ce3f1-0135-4633-9143-6357808d0347/issue_101.jpg" length="244929" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/leworldmodel-jepa-but-more-practical</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/leworldmodel-jepa-but-more-practical</guid>
  <pubDate>Tue, 31 Mar 2026 18:41:00 +0000</pubDate>
  <atom:published>2026-03-31T18:41:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
    <category><![CDATA[Weekly Papers Recap]]></category>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Mar 24th ~ Mar 31th</i><br><i>#101 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 5.5k</span></span> <a class="link" href="https://z.ai/subscribe?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">GLM-5.1 by z.ai</a> is now officially available to all GLM Coding Plan users. To start using the new model, simply update your configuration file, such as <code>~/.claude/settings.json</code>, by manually changing the model name to &quot;glm-5.1.&quot; If you look at the benchmarks, it doesn’t look that impressive, but what’s great is that it offers <b>3× usage of the Claude Pro plan</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6a8af96d-e48e-4b97-adda-ea8ade7a62e7/image.png?t=1774973154"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 2.2k</span></span> <a class="link" href="https://ai.meta.com/research/sam3/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">Meta has released SAM 3.1</a>, a drop-in update that introduces &quot;object multiplexing&quot; to track up to 16 objects simultaneously in a single forward pass. This update doubles video processing throughput and removes memory bottlenecks, making high-performance AI applications feasible on smaller, more accessible hardware. <a class="link" href="https://github.com/facebookresearch/sam3?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">View on GitHub</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1c35d6ea-7301-41a3-ab84-32ee16a61b9d/image.png?t=1774973405"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 15k</span></span> <a class="link" href="https://aidemos.atmeta.com/tribev2/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">Meta has introduced TRIBE v2</a>, a new foundation model trained on over 500 hours of fMRI data to predict how the human brain responds to visual and auditory stimuli. This &quot;digital twin&quot; of neural activity achieves a nearly 3x improvement in zero-shot predictions over previous methods. <a class="link" href="https://github.com/facebookresearch/tribev2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">View on GitHub</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/06cfcf13-e7f8-484a-9049-0532d624524c/image.png?t=1774973586"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 38k</span></span> <a class="link" href="https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">Google has released TurboQuant</a>, which is a compression algorithm designed to optimize the key-value (KV) cache of Large Language Models. This new method <b>reduces memory requirements by at least sixfold</b>, allowing developers to run massive models on significantly more modest hardware. Moreover, it also delivers up to an <b>8x speedup</b> in inference. Most importantly, TurboQuant achieves these performance gains with zero loss in accuracy, ensuring that model quality remains untouched despite the extreme compression. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5d14f6ce-f331-4a34-9d8c-d42ae9362292/image.png?t=1774973817"/><div class="image__source"><span class="image__source_text"><p><i>TurboQuant demonstrates robust KV cache compression performance across the</i> <a class="link" href="#industry-news-in-1-line" rel="noopener noreferrer nofollow" style="color: rgb(26, 115, 232)">LongBench</a><i> benchmark relative to various compression methods</i></p></span></div></div><p class="paragraph" style="text-align:left;">The hype for Google’s TurboQuant is facing pushback from the research community as the paper contains serious technical inaccuracies and misleading comparisons. <a class="link" href="https://x.com/gaoj0017/status/2037532673812443214?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">Lead researchers from the RaBitQ project claim</a> that TurboQuant misrepresents their methodology and fails to acknowledge fundamental similarities, specifically regarding the use of the Johnson-Lindenstrauss transform. According to public statements, these flaws were flagged to the authors prior to submission, yet the paper was allegedly published and promoted without the necessary corrections. </p></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Distillation Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="composer-2-technical-report">Composer 2 Technical Report</h2><p class="paragraph" style="text-align:left;"><i>Cursor Research Team</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 5.3k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Coding </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When developers ask AI models to write code, the models perform brilliantly on laboratory tests but stumble in the messy, ambiguous world of actual software engineering. Researchers realized that existing AI coding benchmarks were simply too neat, providing detailed instructions for isolated bugs. Developers constantly deal with vague bug reports, massive codebases, and confusing production logs. To bridge this gap, scientists set out to build a specialized assistant that genuinely thinks like a seasoned engineer navigating the authentic friction of daily development.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f713f25e-f45b-4a4f-9555-96a75c35cb22/CleanShot_2026-03-31_at_20.55.39_2x.png?t=1774970751"/><div class="image__source"><span class="image__source_text"><p>Overview of a single grouped GEMM training flow in our Mixture-of-Experts layer.</p></span></div></div><p class="paragraph" style="text-align:left;">This paper introduces <b>Composer 2</b>, which is a new model that dramatically improves how AI handles long-term programming tasks. The researchers achieved this through a clever two-step training process.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1dd3f4c7-2697-42f7-8890-c0d87ff3756e/CleanShot_2026-03-31_at_20.55.14_2x.png?t=1774970728"/><div class="image__source"><span class="image__source_text"><p>Example CursorBench task</p></span></div></div><p class="paragraph" style="text-align:left;">First, they immersed the model in massive amounts of code to build profound baseline knowledge. Then, they put it through rigorous reinforcement learning, simulating actual user sessions where the AI practiced solving diverse problems. To help the model stay perfectly focused during lengthy programming tasks, the team introduced a brilliant self-summarization technique.</p><p class="paragraph" style="text-align:left;">Instead of getting overwhelmed by a long history of commands, the model constantly writes little internal summaries for itself. This preserves crucial context while discarding clutter. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e6391c0c-94aa-4c4f-a1d5-ccd79222e870/CleanShot_2026-03-31_at_20.54.54_2x.png?t=1774970704"/></div><p class="paragraph" style="text-align:left;">To prove this approach works, the team bypassed traditional public tests and created a rigorous new evaluation based entirely on authentic engineering scenarios. In these demanding simulations, which require extensive code changes from incredibly brief instructions, the model demonstrated remarkable accuracy and coherence.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.24477?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="claudini-autoresearch-discovers-sta">Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs</h2><p class="paragraph" style="text-align:left;"><i>Panfilov et al. [MATS, ELLIS Institute Tubingen & Max Planck Institute for Intelligent Systems, Tubingen AI Center, Imperial College London]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.5k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Coding Vulnerabilities </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Can AI automatically conduct research to make future technology safer? Researchers built an automated security researcher to see if an AI could invent entirely new methods for pressure-testing digital vulnerabilities. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/80275241-8d37-44a5-adb3-d7e850531489/autoresearch_loop.png?t=1774971183"/></div><p class="paragraph" style="text-align:left;">To test this, the research team built a continuous loop where the AI reviewed dozens of existing security-testing algorithms, wrote new software to improve them, and then evaluated its own creations. Instead of simply generating tricky text prompts to bypass safety filters, the AI engineered the underlying mathematical algorithms that actively search for these vulnerabilities.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4ba4cf30-1d80-47f1-af0f-20109ed307be/pareto_evolution_small.png?t=1774971198"/><div class="image__source"><span class="image__source_text"><p>Claudini Strongly Outperforms a Classical AutoML Method.</p></span></div></div><p class="paragraph" style="text-align:left;">By intelligently splicing together previous techniques and writing clever mechanisms to avoid getting stuck during its analysis, the AI made massive strides. It successfully generated novel algorithms that dramatically outperformed human-made baselines, jumping from a success rate of less than ten percent to forty percent on a targeted safety filter.</p><div class="embed"><a class="embed__url" href="https://github.com/romovpa/claudini?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank"><div class="embed__content"><p class="embed__title"> GitHub - romovpa/claudini: Autoresearch for LLM adversarial attacks </p><p class="embed__description"> Autoresearch for LLM adversarial attacks. Contribute to romovpa/claudini development by creating an account on GitHub. </p><p class="embed__link"> github.com/romovpa/claudini </p></div></a></div><p class="paragraph" style="text-align:left;">The most remarkable finding is just how adaptable these new AI-authored algorithms proved to be. When scientists tested these methods on a completely different, highly secured model that the AI had never even encountered, the new algorithms bypassed the defenses with a flawless hundred percent success rate.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.24511?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="le-world-model-stable-endto-end-joi">LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels</h2><p class="paragraph" style="text-align:left;"><i>Maes et al. [Mila & Université de Montréal, New York University, Samsung SAIL, Brown University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 3.7k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM World Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI needs an internal &quot;world model&quot; to predict the consequences of its actions. Researchers have tried teaching AI to build these models directly from raw video pixels, compressing complex visual scenes into a streamlined imagination space. However, this process faces a frustrating hurdle known as representation collapse.</p><p class="paragraph" style="text-align:left;">When asked to predict the future, the AI often takes the lazy route, mapping every image to the exact same uniform representation just to guarantee a perfect prediction score. To prevent this, previous systems relied on highly fragile workarounds.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/71d739b0-2b31-40d4-ae5d-8127d628a788/CleanShot_2026-03-31_at_21.06.49_2x.png?t=1774971420"/><div class="image__source"><span class="image__source_text"><p>LeWorldModel Training Pipeline. </p></span></div></div><p class="paragraph" style="text-align:left;">This paper developed LeWorldModel, which is a highly stable world model from scratch using just two simple rules. First, the model predicts the next compressed state of its environment based on a given action. Second, a clever mathematical regulator forces these compressed representations to stay continuously diverse, spreading them out to match a natural bell-curve distribution. By enforcing this varied shape, the model is strictly prevented from collapsing into a single, lazy answer.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d4d5a3ea-149e-447d-8f82-ee42ca6648f7/CleanShot_2026-03-31_at_21.09.04_2x.png?t=1774971554"/><div class="image__source"><span class="image__source_text"><p>Pseudo-code for the training procedure of LeWorldModel.</p></span></div></div><p class="paragraph" style="text-align:left;">By reducing the complex tuning down to a single parameter, this compact model trains on one standard graphics card in just a <b>few hours</b>. Impressively, it plans up to <b>forty-eight times faster</b> than bulkier alternatives across complex control tasks. Even more fascinating, without explicit physics lessons, the model organically learns to track object locations and registers mathematical surprise when shown impossible events like spontaneous teleportation.</p><div class="embed"><a class="embed__url" href="https://le-wm.github.io/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical" target="_blank"><div class="embed__content"><p class="embed__title"> LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels </p><p class="embed__description"> End-to-end joint-embedding predictive architecture from pixels. </p><p class="embed__link"> le-wm.github.io </p></div><img class="embed__image embed__image--right" src=""/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.19312?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Self-Distillation of Hidden Layers for Self-Supervised Representation Learning</h2><p class="paragraph" style="text-align:left;"><i>Lowe et al. [Vector Institute, Carleton University, Dalhousie University</i>, <i>University of British Columbia, University of Guelph]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 780 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Distillation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching an AI to understand the visual world is complex. For a long time, scientists have had to choose between two extreme training methods, both with frustrating limitations. One approach forces the AI to perfectly reconstruct raw, low-level pixels. While this keeps the system safely grounded in reality, it leaves the AI struggling to grasp big-picture concepts without a lot of extra hand-holding.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/450e77c4-7fdd-42bb-8723-9373a9dc790d/CleanShot_2026-03-31_at_21.09.58_2x.png?t=1774971612"/><div class="image__source"><span class="image__source_text"><p>Multi-layer self-distillation with Bootleg. </p></span></div></div><p class="paragraph" style="text-align:left;">The other approach asks the AI to predict only highly abstract, final-stage ideas. However, because the system is essentially generating its own study material in a continuous loop, it often becomes unstable, loses touch with the actual image, and completely breaks down during training. </p><p class="paragraph" style="text-align:left;">To solve this, this research has introduced Bootleg. Instead of making the AI guess only the most basic details or only the most complex final concepts, they tasked it with predicting multiple hidden layers of information all at once. This beautifully mirrors how the human brain processes sights, where early visual processing picks up simple edges and colors, while deeper brain regions recognize complete objects.</p><p class="paragraph" style="text-align:left;">By forcing the AI to simultaneously predict early, middle, and late stages of understanding, the system has to compress a wealth of varied knowledge through a tight informational bottleneck. This multi-level approach brilliantly keeps the AI anchored to actual visual features while it masters complex, abstract ideas.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e80bd88e-05eb-4135-8422-343fc4ddf932/CleanShot_2026-03-31_at_21.13.07_2x.png?t=1774971798"/><div class="image__source"><span class="image__source_text"><p>Results for masked self-supervised learning with Bootleg and baselines. </p></span></div></div><p class="paragraph" style="text-align:left;">The researchers discovered that this technique creates a dramatically smarter system, outperforming previous methods by significant margins in both recognizing what is in an image and mapping out exact scenes.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.15553?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=leworldmodel-jepa-but-more-practical"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/P9uNy71YukQ" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=6f4577e6-93dc-4f9b-9f9c-b42c0fb72ada&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Rotate attention by 90 degrees...? Kimi&#39;s New Attention Residuals</title>
  <description>plus more about V-JEPA 2.1, Mamba 3, and latent planning</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5ab46a60-9b20-4927-a94d-9639d7021963/issue_100.jpg" length="244934" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/rotate-attention-by-90-degrees-kimi-s-new-attention-residuals</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/rotate-attention-by-90-degrees-kimi-s-new-attention-residuals</guid>
  <pubDate>Wed, 25 Mar 2026 18:32:00 +0000</pubDate>
  <atom:published>2026-03-25T18:32:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Mar 17th ~ Mar 24th</i><br><i>#100 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.3k</span></span> Xiaomi has announced the release of <a class="link" href="https://mimo.xiaomi.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals#blog" target="_blank" rel="noopener noreferrer nofollow">MiMo-V2-Pro</a>, Omni, and TTS models. These new models introduce advanced features with global top-tier agent performance, multimodal interaction for seeing and hearing, and expressive voice synthesis. You can test these models on the <a class="link" href="https://aistudio.xiaomimimo.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals#/" target="_blank" rel="noopener noreferrer nofollow">web</a> or via the API portal.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e70dfda8-958b-4e01-ad4c-e9b0f8b0c665/image.png?t=1774372024"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 1.5k</span></span> MiniMax has launched its <a class="link" href="https://www.minimax.io/news/minimax-m27-en?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank" rel="noopener noreferrer nofollow">M2.7 model</a>, which uses a <b>recursive self-evolution architecture</b> that contributed to an 88% win-rate over its predecessor. This model achieves state-of-the-art performance in software engineering benchmarks and high-fidelity document editing, while demonstrating enhanced agentic capabilities with 97% skill adherence. <a class="link" href="https://agent.minimax.io/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank" rel="noopener noreferrer nofollow">Try it today on the web</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/938b60fb-2159-4dc4-8f64-317e9772eb85/image.png?t=1774372219"/></div><p class="paragraph" style="text-align:left;"></p></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Distillation Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;"><b>We just added a new chapter on Distillaion!</b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="attention-residuals">Attention Residuals</h2><p class="paragraph" style="text-align:left;"><i>Kimi Team</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 15k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">LLMs rely on residual connections to pass information through deep stacks of layers. While these connections act as a vital &quot;gradient highway&quot;, they work like a blunt instrument. In current architectures, every layer simply adds its output to a uniform sum of all previous layers. This approach causes the information representing the model’s &quot;hidden state&quot; to swell uncontrollably as it moves deeper. Over time, this leads to a dilution effect where early, important information becomes buried and difficult for the model to retrieve effectively, effectively limiting how well the model can leverage its own depth.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d505ae16-d0a9-4540-afdb-b284be0d51dc/overview.png?t=1774371585"/><div class="image__source"><span class="image__source_text"><p>Overview of Attention Residuals.</p></span></div></div><p class="paragraph" style="text-align:left;">Researchers have introduced &quot;Attention Residuals&quot; (AttnRes) to replace this clumsy, fixed accumulation with a smarter mechanism. AttnRes allows each layer to selectively aggregate information from previous layers using learned, input-dependent weights. Instead of blindly adding everything together, the model now chooses which previous layers are most relevant to its current task.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a23a156e-cd15-4169-aa3d-61fd80800d1e/CleanShot_2026-03-24_at_22.30.36_2x.png?t=1774371666"/><div class="image__source"><span class="image__source_text"><p>PyTorch-style pseudo code for Block Attention Residuals.</p></span></div></div><p class="paragraph" style="text-align:left;">To ensure this doesn&#39;t create excessive memory demands in massive models, the team developed &quot;Block AttnRes,&quot; which organizes layers into smaller groups. This allows the model to maintain the benefits of selective, smart aggregation while keeping memory and communication costs efficient.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1f0e6564-c8e7-428a-8401-349c89427f58/training_dynamics.png?t=1774371608"/></div><p class="paragraph" style="text-align:left;">Experiments on large-scale models, including a 48B-parameter architecture, showed that this method successfully prevents the hidden-state growth that plagued previous models. By creating a more uniform distribution of signals and gradients across all layers, AttnRes consistently improves performance on complex reasoning, math, and coding benchmarks. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.15031?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="mamba-3-improved-sequence-modeling-">Mamba-3: Improved Sequence Modeling using State Space Principles</h2><p class="paragraph" style="text-align:left;"><i>Lahoti et al. [</i>Carnegie Mellon University, Princeton University, Together AI, Cartesia AI<i>]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.5k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;">Mamba</span></span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Transformer based AI models are incredibly powerful, but they suffer from a &quot;bottleneck&quot; in efficiency. As models grow, the computational effort required to generate each new piece of information, and the memory needed to store that context increases very fast. </p><p class="paragraph" style="text-align:left;">This paper has introduced Mamba-3, a new architecture designed with an &quot;inference-first&quot; mindset to solve these efficiency challenges. By revisiting the mathematical foundations of State Space Models (SSMs), the team implemented three core improvements.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7b52ae1c-dbee-4a9b-9122-22519fcd2a75/mamba3.png?t=1774371479"/></div><p class="paragraph" style="text-align:left;">First, they developed a more expressive discretization method called &quot;exponential-trapezoidal,&quot; which allows the model to handle data dynamics with greater precision than previous versions.</p><p class="paragraph" style="text-align:left;">Second, they incorporated a complex-valued state update rule. This acts like a &quot;rotary&quot; mechanism that enables the model to track states, such as solving arithmetic parity tasks, that were previously impossible for similar linear models to master.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3fb3642-2897-470d-bd87-4ac78ec7ba79/selection.png?t=1774371491"/></div><p class="paragraph" style="text-align:left;">Finally, they shifted to a multi-input, multi-output (MIMO) formulation. This clever adjustment allows the model to perform more computation during the memory-heavy decoding phase without increasing the actual size of its state or slowing down its response time.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6442bb60-ae31-49a9-87f9-93e3c78d6a33/ssd_algorithm.png?t=1774371506"/></div><p class="paragraph" style="text-align:left;">These refinements allow Mamba-3 to achieve a remarkable balance. At the 1.5B scale, it outperforms top-tier competitors in downstream accuracy while simultaneously matching the language-modeling capabilities of its predecessor at half the state size. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b8e7467e-fd11-4b2f-88fa-f4c75bdf40a1/CleanShot_2026-03-24_at_22.28.50_2x.png?t=1774371539"/><div class="image__source"><span class="image__source_text"><p>Prefill and Prefill+Decode latency across sequence lengths.</p></span></div></div><div class="embed"><a class="embed__url" href="https://github.com/state-spaces/mamba?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank"><div class="embed__content"><p class="embed__title"> GitHub - state-spaces/mamba: Mamba SSM architecture </p><p class="embed__description"> Mamba SSM architecture. Contribute to state-spaces/mamba development by creating an account on GitHub. </p><p class="embed__link"> github.com/state-spaces/mamba </p></div><img class="embed__image embed__image--right" src="https://opengraph.githubassets.com/0929e33a4425521a199ffad2179c61bdbc88e4af7b47b558ccd6975412a203cf/state-spaces/mamba"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.15569?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="vjepa-21-unlocking-dense-features-i">V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning</h2><p class="paragraph" style="text-align:left;"><i>Mur-Labadia et al. [FAIR at Meta, Universidad de Zaragoza]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.3k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> JEPA </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">To be truly useful, an AI needs to be a &quot;jack-of-all-trades&quot;: it must understand both the big picture (like identifying a person’s action) and fine-grained local details (like pinpointing the exact edge of a glass for a robot to grasp). Previous models, such as the V-JEPA family, excelled at global understanding but struggled to extract precise local information, resulting in noisy, fragmented visual representations.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ef8ff685-a421-4368-9269-34211852d0c2/architecture_vjepa2_1.jpg?t=1774371360"/><div class="image__source"><span class="image__source_text"><p>V-JEPA 2.1 Architecture</p></span></div></div><p class="paragraph" style="text-align:left;">This gap limited their effectiveness in tasks that demand spatial precision, such as depth estimation or delicate robotic manipulation. The team discovered that the &quot;missing link&quot; was in the training objective. Traditional V-JEPA models only practiced predicting what was missing from an image or video, the masked, hidden patches.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/36735fae-fafe-41e2-9d5d-24b94c164344/bars_teaser_tikz-1.png?t=1774371375"/><div class="image__source"><span class="image__source_text"><p>V-JEPA 2.1 ViT-G performance across dense and global prediction tasks.</p></span></div></div><p class="paragraph" style="text-align:left;">Because the model wasn’t forced to analyze the visible parts of the scene, it essentially learned to treat those visible sections as global summaries rather than detailed spatial maps. To fix this, this paper has introduced a &quot;dense prediction loss,&quot; which forces the model to learn from both the hidden <i>and</i> the visible parts of the input.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b6d0b241-f4bb-46f4-bb4c-f992fd9a6c2f/flowchart.png?t=1774371412"/></div><p class="paragraph" style="text-align:left;">By supervising every part of the scene, the model is compelled to build a coherent, fine-grained understanding of where objects actually exist in space. They further refined this by applying the learning signal at multiple layers deep within the network, a method called &quot;deep self-supervision&quot;, and by using specialized tokenizers that handle images and videos in their native formats.</p><p class="paragraph" style="text-align:left;">These changes, combined with scaling the model and the diversity of training data, resulted in a system capable of state-of-the-art performance in both high-level action forecasting and precise, low-level spatial tasks like depth perception.</p><div class="embed"><a class="embed__url" href="https://github.com/facebookresearch/vjepa2?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank"><div class="embed__content"><p class="embed__title"> GitHub - facebookresearch/vjepa2: PyTorch code and models for VJEPA2 self-supervised learning from video. </p><p class="embed__description"> PyTorch code and models for VJEPA2 self-supervised learning from video. - facebookresearch/vjepa2 </p><p class="embed__link"> github.com/facebookresearch/vjepa2 </p></div><img class="embed__image embed__image--right" src="https://opengraph.githubassets.com/b371955bb418f677b546ce6319dcef580d2f0ef6c9a1bbae69c0e1ef648ad642/facebookresearch/vjepa2"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.14482?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Temporal Straightening for Latent Planning</h2><p class="paragraph" style="text-align:left;"><i>Wang et al. [New York University, Brown University, University of Toronto]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.2k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> JEPA </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">When AI agents interact with complex environments, they typically translate high-dimensional sensory data, like raw video, into a &quot;latent space&quot; to make decisions. For a computer trying to navigate, this means the shortest path between two points in its digital map doesn&#39;t actually correspond to the shortest physical path. Because these internal maps are so twisted, gradient-based planners, which rely on finding the smoothest way to reach a goal, often get stuck or perform poorly, forcing researchers to rely on computationally expensive search-based methods.</p><table width="100%" class="bh__column_wrapper"><tr><td width="50%" class="bh__column"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/da403088-bd99-43cf-9650-81f49112a9cd/umaze_gt.png?t=1774371120"/><div class="image__source"><span class="image__source_text"><p><i>UMaze: ground-truth geodesic distance.</i></p></span></div></div></td><td width="50%" class="bh__column"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1e6bf35b-99fc-4b21-beb8-5ae70299ec18/umaze_resnet_global.png?t=1774371184"/><div class="image__source"><span class="image__source_text"><p>UMaze: ResNet-global after straightening.</p></span></div></div></td></tr></table><p class="paragraph" style="text-align:left;">This paper introduces a way to &quot;straighten&quot; these maps, helping AI agents see the world more linearly so they can plan their actions more efficiently. The researchers developed a method called &quot;temporal straightening&quot; to fix the distorted geometry of latent spaces.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0f97268c-96e1-4cb5-9f0e-dff3e1bb0e90/curvature_bars.png?t=1774371216"/><div class="image__source"><span class="image__source_text"><p>Latent Curvature and Open-Loop GD Success Rate for Different Encoders. Higher cosine similarity indicates lower curvature.</p></span></div></div><p class="paragraph" style="text-align:left;">This is based on a simple idea: as the agent moves, the model tracks its latent trajectory and minimizes the &quot;curvature&quot; by keeping the velocity vectors of consecutive steps as aligned as possible. It forces the AI to learn representations where movement in the latent space feels like travel along a straight, predictable line rather than a chaotic curve.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4be08189-e092-4441-b3b8-909d98b197ca/architecture.png?t=1774371108"/></div><p class="paragraph" style="text-align:left;">By adding a straightening objective to the standard training process, the model learns to map observations in a way that Euclidean distance, the simplest way to measure &quot;how far&quot; a goal is, finally matches reality. This has a profound effect on planning: because the path is straighter, the math behind gradient-based optimization becomes much more stable.</p><p class="paragraph" style="text-align:left;">In experiments, this approach drastically improved success rates across a variety of navigation and manipulation tasks, allowing agents to reach goals with far greater precision and efficiency without needing the heavy compute power typically required for complex decision-making.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8de89c07-fe60-40db-b295-8a40d19f4817/results_table.png?t=1774371279"/></div><div class="embed"><a class="embed__url" href="https://agenticlearning.ai/temporal-straightening/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals" target="_blank"><div class="embed__content"><p class="embed__title"> Temporal Straightening for Latent Planning </p><p class="embed__description"> Temporal straightening improves latent planning with world models. </p><p class="embed__link"> agenticlearning.ai/temporal-straightening </p></div></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.12231?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=rotate-attention-by-90-degrees-kimi-s-new-attention-residuals"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/xUlX6jvwVfM" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=b3d4d11e-e424-4e1d-ac2c-18d1ab400aaf&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>You can train OpenClaw just by talking to it?</title>
  <description>and more about GLM-OCR, pre-pre-training on NCA, IndexCache, and neural thickets</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1d2bc86a-9536-44e2-b56a-7fe7d40b5018/issue_99.jpg" length="291057" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/you-can-train-openclaw-just-by-talking-to-it</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/you-can-train-openclaw-just-by-talking-to-it</guid>
  <pubDate>Tue, 17 Mar 2026 22:10:00 +0000</pubDate>
  <atom:published>2026-03-17T22:10:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Mar 10th ~ Mar 17th</i><br><i>#99 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 24k</span></span> Claude 3.5 models (Opus and Sonnet) now support a <a class="link" href="https://claude.com/blog/1m-context-ga?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">1-million-token context window</a>, and allow users to process large codebases, extensive document sets, and up to 600 images or PDF pages per request. This update is available across all plans and is integrated by default into Claude Code at standard pricing. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f11c0ca9-21ef-40d7-b65a-4524007ee59a/69b49c06e1c573f3ce50276b_image__3_.png?t=1773763196"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 44k</span></span> <a class="link" href="https://blog.google/products-and-platforms/products/maps/ask-maps-immersive-navigation/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Google Maps has integrated Gemini AI</a> to help you explore and navigate more easily. You can now use a new &quot;Ask Maps&quot; feature to get conversational answers to specific, real-world questions, like finding a place to charge your phone or a well-lit tennis court. Additionally, a new &quot;Immersive Navigation&quot; tool is rolling out, which provides vivid 3D visuals and more detailed route guidance to help you navigate your surroundings with more confidence.</p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 5.3k</span></span> <a class="link" href="https://hermes-agent.nousresearch.com/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Hermes Agent</a> by Nous Research is an open-source, Python-based tool designed to grow with you by utilizing a multi-level memory system and persistent machine access, similar to OpenClaw. It works across your CLI and various messaging platforms, and offers developers an extensible framework for complex tasks like subagent management and programmatic tool calling.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/717b1744-c254-4544-a813-5be565ec7663/image.png?t=1773785089"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 6.2k</span></span> <a class="link" href="https://huggingface.co/1Covenant/Covenant-72B?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Covenant-72B</a> is the largest decentralized LLM pre-training run in history. It is a 72B parameter model trained across commodity internet connections without centralized clusters or whitelisting. By using innovative techniques like SparseLoCo for bandwidth efficiency and a blockchain-based &quot;Gauntlet&quot; system for validation, the project achieved performance levels competitive with models trained in traditional data centers. <a class="link" href="https://www.tplr.ai/chat?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Try it in browser</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3308bd02-f95f-45e1-a2c1-7cfaf0b9bd3a/image.png?t=1773763642"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 19k</span></span> Yann LeCun’s new startup <a class="link" href="https://amilabs.xyz/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Advanced Machine Intelligence</a> (AMI) has secured $1.03 billion in one of the largest seed rounds in history to develop AI systems capable of advanced reasoning, persistent memory, and world-model understanding. These models are designed to understand the physical world while featuring persistent memory and the ability to reason, plan, and operate safely.</p></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW Distillation Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;"><b>We just added a new chapter on Distillation</b> <b>too!</b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have an early bird offer, where you would get 40% off on the yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="open-claw-rl-train-any-agent-simply">OpenClaw-RL: Train Any Agent Simply by Talking</h2><p class="paragraph" style="text-align:left;"><i>Wang et al. [Princeton University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 674 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Every time we interact with an AI, it receives immediate feedback: a follow-up question, a software error, or a screen transition. Existing AI systems throw this valuable experience away. They treat our replies merely as context for their very next move, missing a massive opportunity to learn. Researchers wanted to solve this by capturing these everyday reactions, what they call next-state signals, and turning them into a live, continuous learning stream.</p><p class="paragraph" style="text-align:left;">We can build a system where an AI naturally improves simply by being used, turning ordinary conversations and software tasks into a seamless training loop without needing to pause for offline updates. The researchers built a unified framework called OpenClaw-RL.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ccc5932b-f549-4e87-9c66-3688f60fcce5/framework.png?t=1773760687"/></div><p class="paragraph" style="text-align:left;">The AI receives two powerful forms of information. First, there are evaluative signals, which act like a simple score indicating whether an action succeeded or frustrated the user. Second, there are directive signals. When a user corrects an AI by explaining how it should have responded, or when a software tool outputs a detailed error, it provides a clear map for improvement.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ffc13183-eb06-465b-8f18-17dcb068f6fb/rlserver.png?t=1773760715"/></div><p class="paragraph" style="text-align:left;">The framework extracts these specific textual hints to give the AI rich, word-by-word guidance. Because the system is built asynchronously, the AI can chat with a user, a background judge can evaluate its performance, and a training engine can update the AI&#39;s core behavior all at the exact same time, without ever interrupting the workflow.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/99ad0a95-e5a4-4b24-b4fe-9b7b6521cbf9/openclawrl1performance.png?t=1773760703"/></div><p class="paragraph" style="text-align:left;">By combining basic scoring with these rich textual hints, researchers found that personal assistants rapidly adapt their tone, becoming much more natural after just a handful of conversations.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.10165?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="neural-thickets-diverse-task-expert">Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights</h2><p class="paragraph" style="text-align:left;"><i>Gan and Isola [MIT CSAIL]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 818 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Pretraining </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching an AI a new skill feels like searching for a microscopic needle in a haystack. Researchers have long believed that adapting a massive, billion-parameter model required highly complex, meticulous, step-by-step mathematical adjustments just to find a version of the system that performed a specific task well. Blindly guessing the right settings was considered mathematically impossible and entirely out of the question.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f6a5b673-9ba8-466d-855d-b147850e6379/CleanShot_2026-03-17_at_20.52.39_2x.png?t=1773760972"/></div><p class="paragraph" style="text-align:left;">However, this study discovered that as models grow larger and undergo extensive initial training, adapting them to specific tasks like mathematical reasoning or coding no longer requires such rigid, painstaking effort. The heavy lifting of learning has already been done, signaling an exciting era where customizing powerful technology is becoming surprisingly natural, fast, and accessible.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/da09c942-e16a-427c-b4ff-cbec467d428e/image.png?t=1773760932"/></div><p class="paragraph" style="text-align:left;">The researchers found that scaling up these models fundamentally transforms their underlying structure. Instead of a desolate landscape where a good solution is a lone needle, massive models are surrounded by a dense, flourishing &quot;thicket&quot; of specialized solutions.</p><p class="paragraph" style="text-align:left;">By making random, tiny adjustments to the model&#39;s underlying numbers, the researchers uncovered an abundance of hidden specialists waiting nearby. One random tweak might produce an expert in creative writing, while another creates a brilliant chemist, each specializing in one area while forgetting others.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b55e95a0-a834-4c9e-b907-ba6cf42e88e3/image.png?t=1773760946"/></div><p class="paragraph" style="text-align:left;">To harness this rich diversity, the team tested a beautifully simple approach: they generated thousands of random tweaks simultaneously, kept the top performers for a specific task, and had them vote on the final answer. This parallel guess-and-check strategy matched the accuracy of today’s most advanced training methods but operated in a <b>fraction of the time</b> because it avoided slow, sequential updates. </p><div class="embed"><a class="embed__url" href="https://thickets.mit.edu/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank"><div class="embed__content"><p class="embed__title"> Neural Thickets · MIT </p><p class="embed__description"> Diverse Task Experts Are Dense Around Pretrained Weights </p><p class="embed__link"> thickets.mit.edu </p></div><img class="embed__image embed__image--right" src=""/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.12228?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="index-cache-accelerating-sparse-att">IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse</h2><p class="paragraph" style="text-align:left;"><i>Bai et al. [Tsinghua University, Z.ai]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 540 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI models use an attention mechanism to connect different pieces of information, but as the text gets longer, this demands a staggering amount of computing power. Developers recently introduced a clever shortcut called sparse attention, which uses a specialized &quot;indexer&quot; tool to scan the text and pick only the most relevant words for the AI to focus on at each step. It is a brilliant fix, but researchers ran into a new bottleneck. This indexer operates independently at every single layer of the AI network. As the text grows, just running this indexer consumes a massive chunk of the system&#39;s processing time, slowing everything down.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0faeefe4-a5f3-4892-9ada-6e795d560984/CleanShot_2026-03-17_at_20.59.42_2x.png?t=1773761393"/><div class="image__source"><span class="image__source_text"><p>Side-by-side comparison of inference loops.</p></span></div></div><p class="paragraph" style="text-align:left;">Looking closely at how these models process data, researchers noticed an incredible inefficiency. They discovered that consecutive layers of the AI were repeatedly selecting almost the exact same important words, often with a near-perfect overlap. To solve this, the researchers developed an elegant, hopeful solution called IndexCache. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5c9dd9f4-553b-4d7b-8983-fc61ed557c7a/CleanShot_2026-03-17_at_20.59.59_2x.png?t=1773761411"/></div><p class="paragraph" style="text-align:left;">Instead of forcing every layer to do the heavy lifting of scanning and selecting information, IndexCache designates a few specific layers as &quot;Full&quot; layers to run the indexer. The remaining layers become &quot;Shared&quot; layers, which simply borrow the selected words from the nearest Full layer. The team created two ways to apply this: a training-free method that calculates the absolute best pattern of Full and Shared layers for existing models, and a training-aware method that actively teaches the AI to share this data.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e5e8d358-d8b0-4118-b1ca-7bb548b9a795/CleanShot_2026-03-17_at_21.00.30_2x.png?t=1773761439"/><div class="image__source"><span class="image__source_text"><p>Training-free IndexCache at 1/2, 1/4, and 1/8 indexer retention. ‘Long’ and ‘G&R’ aggregate benchmark scores.</p></span></div></div><p class="paragraph" style="text-align:left;">By simply reusing this cached information, researchers eliminated 75 percent of the heavy indexer computations with negligible drops in the model&#39;s reasoning quality. When tested on massive systems, this straightforward change nearly doubled the speed at which the AI reads information and significantly accelerated its ability to generate answers.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.12201?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">Training Language Models via Neural Cellular Automata</h2><p class="paragraph" style="text-align:left;"><i>Lee et al. [MIT, Improbable AI Lab]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.5k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Sampling </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">LLMs rely on massive amounts of human-written text to learn how to reason and communicate. However, this approach faces a looming wall: high-quality human data is finite, often riddled with biases, and mixes actual reasoning with messy, subjective language. They investigated whether models could learn the fundamental mechanics of reasoning by training on purely synthetic, non-linguistic data before ever seeing a human sentence.</p><p class="paragraph" style="text-align:left;">To test this, researchers turned to neural cellular automata (NCA), algorithmic systems that generate complex, ever-changing grid patterns using simple, local rules. Unlike static text, these patterns can be generated cheaply and in infinite supply, allowing for precise control over the &quot;complexity&quot; of the training data.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5524ce76-66dd-4891-be54-98935b454324/image.png?t=1773761613"/></div><p class="paragraph" style="text-align:left;">The team pre-trained models on these synthetic grid trajectories, then followed up with standard training on natural language. The results show that models that practiced on just 164 million synthetic tokens learned faster and performed better than those trained solely on significantly larger amounts of traditional internet text.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7e2ad5b0-e9c5-4e39-8137-0c2051ab7c38/image.png?t=1773761635"/></div><p class="paragraph" style="text-align:left;">While traditional text training can cause a model to rely on human biases or semantic shortcuts, the synthetic grids force the model to focus purely on tracking long-range patterns and inferring underlying rules. Furthermore, the researchers found they could tune the complexity of this synthetic data to match specific domains; for instance, code benefited from simpler rules, while math and web text thrived on higher complexity.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/50dd5649-a6bb-49af-8857-54207f5f196e/CleanShot_2026-03-17_at_21.04.21_2x.png?t=1773761678"/></div><div class="embed"><a class="embed__url" href="https://hanseungwook.github.io/blog/nca-pre-pre-training/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank"><div class="embed__content"><p class="embed__title"> Training Language Models via Neural Cellular Automata </p><p class="embed__link"> hanseungwook.github.io/blog/nca-pre-pre-training </p></div><img class="embed__image embed__image--right" src=""/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.10055?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="glmocr-technical-report">GLM-OCR Technical Report</h2><p class="paragraph" style="text-align:left;"><i>Duan et al. [Zhipu AI, Tsinghua University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Modern information systems rely heavily on extracting knowledge from complex, visually dense documents like financial reports, invoices, and scientific papers. While recent multimodal AI models have improved how we read these documents, they often suffer from a major drawback: their massive size makes them slow, memory-intensive, and difficult to deploy in practical, real-world settings.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/115f675f-7e46-41af-b090-3147108a1f17/CleanShot_2026-03-17_at_21.08.05_2x.png?t=1773761902"/><div class="image__source"><span class="image__source_text"><p>Architecture and workflow of the GLM-OCR framework</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers developed GLM-OCR, a lightweight, highly optimized framework that packs significant power into a compact <b>0.9-billion-parameter</b> design. The system merges a specialized visual encoder with a streamlined language decoder. What makes it particularly clever is the shift away from standard “one-token-at-a-time” generation, which is notoriously slow for structured documents.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/33a6a399-2319-4610-90c4-3c1092f69e02/docparse.png?t=1773761864"/></div><p class="paragraph" style="text-align:left;">Instead, the researchers implemented a Multi-Token Prediction mechanism that allows the model to predict several tokens simultaneously. By using a shared-parameter scheme to keep memory usage low, the model effectively boosts throughput, the speed at which it processes data, by roughly 50% without sacrificing accuracy.</p><p class="paragraph" style="text-align:left;">To handle real-world complexities, the system employs a two-stage pipeline. First, it uses an analysis module to detect the layout of a document, breaking a complex page into manageable regions. These regions are then processed in parallel, allowing for faster and more robust recognition of everything from handwritten text to complicated table structures.</p><div class="embed"><a class="embed__url" href="https://huggingface.co/zai-org/GLM-OCR?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it" target="_blank"><div class="embed__content"><p class="embed__title"> zai-org/GLM-OCR · Hugging Face </p><p class="embed__description"> We’re on a journey to advance and democratize artificial intelligence through open source and open science. </p><p class="embed__link"> huggingface.co/zai-org/GLM-OCR </p></div><img class="embed__image embed__image--right" src="https://cdn-thumbnails.huggingface.co/social-thumbnails/models/zai-org/GLM-OCR.png"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.10910?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=you-can-train-openclaw-just-by-talking-to-it"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/qznFV59f3Uk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=18f3814c-56e8-4b68-8e85-f4475dd94f24&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Flash Attention 4 is nuts</title>
  <description>and more about Speculative Speculative Decoding, SWE-CI, and Beyond Language Modeling</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/242099e8-0ada-4e3c-8919-7c8f74e9104f/issue_98.jpg" length="421025" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/flash-attention-4-is-nuts</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/flash-attention-4-is-nuts</guid>
  <pubDate>Tue, 10 Mar 2026 19:30:00 +0000</pubDate>
  <atom:published>2026-03-10T19:30:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Mar 3rd ~ Mar 10th</i><br><i>#98 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 6.8k</span></span> Sarvam AI has announced the <a class="link" href="https://www.sarvam.ai/blogs/sarvam-30b-105b?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" target="_blank" rel="noopener noreferrer nofollow">open-source release of its 30B and 105B parameter models</a>, which were developed entirely in-house to target both global benchmarks and Indian language tasks. The model weights are now available on Hugging Face and AIKosh, featuring day-zero support for SGLang with vLLM compatibility expected soon. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cf4452b8-d497-484a-a7df-31b778e5e372/image.png?t=1773165675"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 23k</span></span> OpenAI has <a class="link" href="https://openai.com/index/introducing-gpt-5-4/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" target="_blank" rel="noopener noreferrer nofollow">launched GPT-5.4 Thinking and GPT-5.4 Pro models</a> across ChatGPT, the API, and Codex. These new models are designed to enhance reasoning, coding, and agentic workflows. It also includes features such as advanced deep web research capabilities and the ability for users to interrupt the model mid-process to provide real-time instructions or course corrections.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0483e7fc-d71d-437f-a91e-adc9d49bb8d5/SWE-Bench_Pro__public_.png?t=1773165809"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 15k</span></span> Andrej Karpathy has developed an <a class="link" href="https://github.com/karpathy/autoresearch?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" target="_blank" rel="noopener noreferrer nofollow">&quot;autoresearch&quot; agentic workflow</a> designed to autonomously optimize neural network training through iterative experimentation. When applied to his nanochat project, the tool identified 20 additive changes that reduced the &quot;Time to GPT-2&quot; training benchmark from 2.02 hours to 1.80 hours. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cc49f92d-ccdd-4a13-b192-b33e98b2076d/image.png?t=1773165985"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW MoE Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;"><b>We just added a new chapter on MoE</b>, that goes through the history, the key techniques, and the current state of MoE that frontier model uses. With over 10,000 words written!</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e379f8f4-ca30-4804-a3fd-bfb8df42f3d8/image.png?t=1773169810"/></div><p class="paragraph" style="text-align:left;">We currently have a early bird offer, where you would get 40% off yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="speculative-speculative-decoding">Speculative Speculative Decoding</h2><p class="paragraph" style="text-align:left;"><i>Kumar et al. [Stanford University, Princeton University, Together AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 22k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Decoding </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">The primary hurdle in making LLMs feel instantaneous is a phenomenon known as the sequential bottleneck. Standard AI models generate text one word, or &quot;token&quot;, at a time, which fails to fully use the massive parallel computing power of modern hardware. Researchers previously introduced &quot;speculative decoding&quot;, a method where a small, fast model drafts a few guesses for a larger model to verify. The drafting model has to wait for the larger model to finish checking its work before it can start guessing the next set of words. This creates a persistent lag that limits how fast even the most advanced systems can communicate.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6c894daa-cc17-47b9-99da-8c9fb68f1991/CleanShot_2026-03-10_at_22.56.12_2x.png?t=1773163585"/><div class="image__source"><span class="image__source_text"><p>Ordinary speculative decoding (SD) requires the verifier to wait idly for the draft to speculate.</p></span></div></div><p class="paragraph" style="text-align:left;">To eliminate this idle time, researchers have developed Speculative Speculative Decoding (SSD) via an optimized algorithm called Saguaro. The biggest change is decoupling the &quot;guesser&quot; from the &quot;checker&quot; entirely. While the large target model is busy verifying a current batch of text, Saguaro’s draft model looks ahead and predicts several possible outcomes of that verification. It prepares a menu of &quot;potential futures.&quot; If the larger model confirms one of these predicted outcomes, the system can immediately provide the next set of words without any drafting delay. This parallel approach effectively hides the time spent guessing, transforming a sequential process into a streamlined, continuous flow.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8209abcd-3b89-425b-b631-38eb8d009ab5/CleanShot_2026-03-10_at_22.57.11_2x.png?t=1773163646"/></div><p class="paragraph" style="text-align:left;">By using a clever &quot;geometric fan-out&quot; strategy, Saguaro focuses its computational effort on the most likely verification results, ensuring the &quot;speculation cache&quot; is highly accurate even at higher creative temperatures. This method is entirely lossless, meaning it achieves the exact same high-quality output as the original model but at much higher speeds.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7e6b4af7-9d1e-4ca6-b2e2-b94ddd748d6d/CleanShot_2026-03-10_at_22.58.11_2x.png?t=1773163702"/><div class="image__source"><span class="image__source_text"><p>Advantage of geometric fan out strategy increases at higher temperatures, improving both speculation cache hit rate (right) and thus end-to-end speed (left).</p></span></div></div><p class="paragraph" style="text-align:left;">Initial results show that Saguaro can deliver text up to five times faster than traditional generation methods and twice as fast as previous state-of-the-art speculative techniques.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.03251?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="beyond-language-modeling-an-explora">Beyond Language Modeling: An Exploration of Multimodal Pretraining</h2><p class="paragraph" style="text-align:left;"><i>Tong et al. [FAIR, New York University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 424 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Multimodal LLMs </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI is getting good at manipulating language, but it lacks a fundamental grasp of the physical world. Researchers describe this limitation using the &quot;allegory of the cave&quot;: current models have mastered the description of shadows on a wall, text, without ever seeing the actual objects casting those shadows. Because text is a human abstraction, it is a &quot;lossy&quot; version of reality that misses the raw physics, geometry, and causality of our environment. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b9e32d82-c9ea-43b4-ad90-ee7563414f84/CleanShot_2026-03-10_at_23.03.51_2x.png?t=1773164041"/><div class="image__source"><span class="image__source_text"><p>Overview of this study.</p></span></div></div><p class="paragraph" style="text-align:left;">By training models to &quot;see&quot; and &quot;read&quot; simultaneously from birth, we can build a more grounded intelligence that understands the world’s dynamics directly.</p><p class="paragraph" style="text-align:left;">The researchers developed a unified model using a framework called Transfusion, which trains a single &quot;brain&quot; to perform two different tasks at once: predicting the next word in a sequence and reconstructing visual frames through a process called diffusion.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5e68c62c-b95a-47d1-8559-e1402a06236c/CleanShot_2026-03-10_at_23.04.21_2x.png?t=1773164081"/><div class="image__source"><span class="image__source_text"><p>Examples of training data.</p></span></div></div><p class="paragraph" style="text-align:left;">They discovered that the most effective way to do this is by using a single, high-quality visual representation known as a Representation Autoencoder. This contradicts the traditional belief that you need different &quot;eyes&quot; for understanding an image versus creating one; instead, a unified representation excels at both while keeping the model’s language skills sharp.</p><p class="paragraph" style="text-align:left;">The study also showed that learning from diverse visual data, like video and image-text pairs, actually improves the model’s performance on downstream tasks like reasoning and planning. To manage this complexity, they utilized a &quot;Mixture-of-Experts&quot; architecture.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/753b0252-9988-4ad6-a43d-8c9c3251ae9a/CleanShot_2026-03-10_at_23.05.00_2x.png?t=1773164109"/><div class="image__source"><span class="image__source_text"><p>Multimodal co-training exceeds unimodal performance.</p></span></div></div><p class="paragraph" style="text-align:left;">This design allows the model to naturally evolve specialized internal &quot;experts&quot; for different tasks. It learned to dedicate more capacity to language while efficiently processing the massive amounts of data required by vision. Most impressively, this unified training allowed &quot;world modeling&quot; capabilities to emerge.</p><p class="paragraph" style="text-align:left;">The model could predict the physical outcome of actions, like navigating a robot through a room, using simple text commands, proving that a truly multimodal foundation can bridge the gap between human language and physical reality.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b040cd9f-0240-4e01-b5a9-4265be7ae36d/dense_scaling.png?t=1773164142"/><div class="image__source"><span class="image__source_text"><p>Scaling laws for unified dense models. </p></span></div></div><div class="embed"><a class="embed__url" href="https://beyond-llms.github.io/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts" target="_blank"><div class="embed__content"><p class="embed__title"> Beyond Language Modeling: An Exploration of Multimodal Pretraining </p><p class="embed__description"> Empirical insights on representation, data, architecture, and scaling for native multimodal pretraining. </p><p class="embed__link"> beyond-llms.github.io </p></div><img class="embed__image embed__image--right" src="https://beyond-llms.github.io/assets/figures/progression.png"/></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.03276?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="flash-attention-4-algorithm-and-ker">FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling</h2><p class="paragraph" style="text-align:left;"><i>Zadouri et al. [Princeton University, Meta, Colfax Research, NVIDIA, Georgia Tech, Together AI]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 430 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Latest AI hardware is significantly faster at basic matrix multiplication, but other components, like the units responsible for specialized math or moving data around, haven&#39;t kept the same pace. This creates a digital traffic jam where the fastest parts of a processor are frequently left idling, waiting for slower sections to finish their work.</p><p class="paragraph" style="text-align:left;">FlashAttention-4 was designed to bridge this gap, by offering a clever software design that can overcome the physical limitations of even the most advanced hardware.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4410b593-2aea-4683-8a58-7cfd1edeef45/CleanShot_2026-03-10_at_23.09.29_2x.png?t=1773164404"/><div class="image__source"><span class="image__source_text"><p>FlashAttention-4 forward pipeline.</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers taught the software to find creative shortcuts around these hardware bottlenecks. Because the chip’s dedicated unit for calculating exponentials is often the slowest link in the chain, they developed a way to emulate its functions using more plentiful, general-purpose math units.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/416000d7-9f8d-4de6-84f4-b3375a1e6f4e/CleanShot_2026-03-10_at_23.10.29_2x.png?t=1773164446"/><div class="image__source"><span class="image__source_text"><p>FlashAttention-4 backward computation graph (5 MMA operations + 2 elementwise operations), showing the 1-CTA MMA mode software pipeline order across the prologue, main loop, and tail.</p></span></div></div><p class="paragraph" style="text-align:left;">By using a mathematical technique called polynomial approximation, the software effectively mimics the specialized unit, achieving nearly identical accuracy at much higher speeds. Additionally, they introduced a &quot;conditional rescaling&quot; method that intelligently skips redundant calculations unless they are truly necessary for precision.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/84be4ffc-797a-4dee-95e7-3fbabf052516/CleanShot_2026-03-10_at_23.11.11_2x.png?t=1773164482"/><div class="image__source"><span class="image__source_text"><p>Backward pass TFLOPS on B200 (FP16/BF16) with head dimension 128.</p></span></div></div><p class="paragraph" style="text-align:left;">By keeping more data in high-speed &quot;tensor memory&quot; and coordinating tasks so that different parts of the chip work in perfect sync, FlashAttention-4 can reach incredible speeds, hitting up to 1613 trillion operations per second on the latest Blackwell GPUs. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.05451?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration</h2><p class="paragraph" style="text-align:left;"><i>Chen et al. [Sun Yat-sen University, Alibaba Group]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 855 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Evaluation </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Software engineering is rarely about writing a perfect piece of code in one go; it is an ongoing marathon of maintenance, updates, and evolving requirements. While current AI models have become quite good at solving isolated, one-shot coding tasks, these models are often tested on their ability to provide a quick fix rather than their capacity to sustain a healthy codebase over time.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6a9b38f2-3c2f-42bc-bce1-3794987fc8c1/CleanShot_2026-03-10_at_23.17.38_2x.png?t=1773164868"/><div class="image__source"><span class="image__source_text"><p>Unlike previous benchmarks, SWE-CI proposes an evolution-based evaluation.</p></span></div></div><p class="paragraph" style="text-align:left;">To address this, a new benchmark called SWE-CI shifts the focus from simple functional correctness to long-term maintainability, providing a much-needed lens into how AI agents handle the messy, continuous reality of software development.</p><p class="paragraph" style="text-align:left;">The researchers built this benchmark using actual evolutionary histories from real-world software repositories, with tasks spanning an average of seven months and dozens of consecutive updates. To mirror a professional environment, they employed a dual-agent system where one AI acts as an Architect to identify gaps and set requirements, while a second Programmer agent implements the changes. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3fd18bb6-d095-4d4a-b4b1-2319b1542084/CleanShot_2026-03-10_at_23.18.07_2x.png?t=1773164901"/><div class="image__source"><span class="image__source_text"><p>Data curation process of SWE-CI.</p></span></div></div><p class="paragraph" style="text-align:left;">Success is measured by a novel metric called EvoScore, which specifically rewards agents that make decisions facilitating future growth rather than just immediate fixes. </p><p class="paragraph" style="text-align:left;">The results show that the newest AI models are improving at an accelerating pace, but most still struggle to prevent regressions, the frustrating phenomenon where adding a new feature accidentally breaks an old one. This discovery highlights that the next great frontier for AI developers isn&#39;t just writing code that works, but writing code that lasts.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e3bf44d7-f49b-4a1b-ba2a-d8471677be5f/CleanShot_2026-03-10_at_23.18.41_2x.png?t=1773164952"/><div class="image__source"><span class="image__source_text"><p>SWE-CI uses an architect-programmer dual-agent workflow to model the continuous integration cycle of professional software teams in the real world.</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.03823?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h1 class="heading" style="text-align:left;" id="real-money-fake-models-deceptive-mo"><b>Real Money, Fake Models: Deceptive Model Claims in Shadow APIs</b></h1><p class="paragraph" style="text-align:left;"><i>Zhang et al. [CISPA Helmholtz Center for Information Security]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 805 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Access to frontier models is restricted by high costs, complex payment barriers, or geographical limitations. To bridge this gap, many researchers and developers have turned to &quot;shadow APIs&quot;, third-party services that promise the same power as official models like GPT-5 or Gemini through unofficial, often cheaper channels. While these services appear to democratize access, they operate in a digital gray market with almost no transparency.</p><p class="paragraph" style="text-align:left;">Researchers recently set out to investigate whether these shadow services are truly delivering what they advertise or if they are quietly undermining the integrity of the scientific work built upon them. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0e303853-0a65-491c-84ce-90e8cc3e488b/CleanShot_2026-03-10_at_23.21.12_2x.png?t=1773165088"/><div class="image__source"><span class="image__source_text"><p>Landscape of the shadow APIs.</p></span></div></div><p class="paragraph" style="text-align:left;">The investigation revealed a concerning gap between marketing promises and technical reality. By auditing these services through &quot;model fingerprinting&quot;, a technique that identifies an AI by analyzing the unique statistical patterns and &quot;signatures&quot; in its responses, researchers discovered that nearly half of the shadow services were not using the models they claimed.</p><p class="paragraph" style="text-align:left;">In a classic &quot;<b>bait-and-switch</b>&quot;, premium proprietary models were frequently swapped for cheaper, open-source alternatives behind the scenes. This deception led to significant performance collapses; in high-stakes fields like medicine and law, accuracy dropped by nearly half when compared to official versions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60bea12e-9481-4b55-b0f6-9aa9422f0d04/CleanShot_2026-03-10_at_23.26.05_2x.png?t=1773165377"/><div class="image__source"><span class="image__source_text"><p>Fingerprinting results via LLMmap matched model and mean cosine distance D with standard.</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers found that while these shadow services might handle simple tasks well, they often fail during complex reasoning and display unpredictable safety behaviors. In addition to identifying the mismatch, the study used statistical testing to prove that the outputs from these services were fundamentally different from the official sources.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2603.01919?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=flash-attention-4-is-nuts"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/aQr_FWJETOk" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=1fddbe5c-9018-46d7-9b07-7ec128ff5df5&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Compress Context... Into a LoRA!?</title>
  <description>plus more on Learning Without Training and The Geometry of Noise</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/786806a7-849f-46b5-9ae8-211df95da43c/issue_97.jpg" length="138646" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/compress-context-into-a-lora</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/compress-context-into-a-lora</guid>
  <pubDate>Wed, 04 Mar 2026 20:00:00 +0000</pubDate>
  <atom:published>2026-03-04T20:00:00Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Feb 24th ~ Mar 4th</i><br><i>#97 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 10k</span></span> Google has begun rolling out Nano Banana 2, its latest image generation model. The updated model uses real-time web search data to improve real-world accuracy and introduces the ability to render clear, multilingual text for designs like posters and logos. Additionally, Nano Banana 2 brings faster generation speeds alongside enhancements to lighting, textures, and overall image detail. Try it today via the <a class="link" href="https://gemini.google.com/app?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Gemini app</a> and <a class="link" href="https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-flash-image-preview&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">web interface</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/50859a00-0a5a-49b2-8549-48636a46a2b0/image.png?t=1772642049"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 3.7k</span></span> <a class="link" href="https://openai.com/index/our-agreement-with-the-department-of-war/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">OpenAI has signed a classified deployment contract with the Department of War</a>, insisting that a &quot;cloud-only&quot; architecture, an internal safety stack, and cleared engineers will somehow strictly prevent the military from using their models for <b>autonomous lethal weapons</b> or mass NSA surveillance. Entrusting a tech corporation to independently self-police lethal and intelligence applications sounds straight out of a black mirror episode.</p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 20k</span></span> Alibaba&#39;s Qwen team has launched the Qwen 3.5 Small Model Series, a family of native multimodal models ranging from a highly compact <b>0.8B</b> to a remarkably capable <b>9B</b> designed for edge devices and lightweight agents. Both the base and instruct models are now available on <a class="link" href="https://huggingface.co/collections/Qwen/qwen35?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Hugging Face</a> and <a class="link" href="https://modelscope.cn/collections/Qwen/Qwen35?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">ModelScope</a>. Additionally, the entire suite is already optimized for local deployment via <a class="link" href="https://ollama.com/library/qwen3.5?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Ollama</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/48c9eb2c-c884-45a8-8601-e59eaa5828de/image.png?t=1772642591"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 8k</span></span> After small series, Alibaba has also introduced the <a class="link" href="https://chat.qwen.ai/?models=qwen3.5-flash&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Qwen 3.5 Medium Model</a> Series, which includes the <b>27B</b>, <b>35B-A3B</b>, and <b>122B-A10B</b> models designed to bridge the gap between mid-sized and frontier AI capabilities. Highlighting a shift toward architectural efficiency over parameter size, the new <b>35B-A3B</b> model notably outperforms the previous generation&#39;s massive 235B model thanks to improved data quality and reinforcement learning. <a class="link" href="https://chat.qwen.ai/?models=qwen3.5-122b-a10b&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Try it in browser</a>.</p></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 8.2k</span></span> Google has announced <a class="link" href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-lite/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Gemini 3.1 Flash-Lite</a>, its fastest and most cost-effective Gemini 3 model to date. It has dynamic &quot;thinking levels&quot; and the model instantly processes high-volume queries while scaling its reasoning for complex edge cases, delivering a 2.5X faster time-to-first-token than its 2.5 Flash predecessor. <a class="link" href="https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-flash-lite-preview&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Try it in browser</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8df63b5f-70b3-4fbe-b844-00acfd49ddfa/gemini-3.1-flash-lite-table_1.gif?t=1772643155"/></div></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Intuitive AI Academy - NEW MoE Chapter!</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><p class="paragraph" style="text-align:left;"><b>We just added a new chapter on MoE</b>, that goes through the history, the key techniques, and the current state of MoE that frontier model uses. With over 10,000 words written.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a0011805-b53d-4cc5-9d88-d2cc98e1c565/image.png?t=1772650337"/></div><p class="paragraph" style="text-align:left;">We currently have a early bird offer, where you would get 40% off yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h1 class="heading" style="text-align:left;" id="learning-without-training"><b>Learning Without Training</b></h1><p class="paragraph" style="text-align:left;"><i>Ryan O’Dowd [Claremont Graduate University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 720 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Training </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Engineers usually build machine learning models by guessing a structure and running exhaustive optimization processes to train them. Researchers wanted to know if there was a more elegant way to overcome the hurdles of high-dimensional, noisy data without relying on these brute-force training methods.</p><p class="paragraph" style="text-align:left;">The current approach assumes a model will eventually learn the underlying patterns, but it lacks constructive mathematical guarantees. By rooting their approach in classical approximation theory, we can save immense computational power while tackling complex problems like tracking brain diseases or analyzing hyperspectral images.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/14240e47-718d-40f1-a2b0-780d02534b82/CleanShot_2026-03-04_at_21.02.01_2x.png?t=1772638331"/><div class="image__source"><span class="image__source_text"><p>Normalized histogram of the density of interest (left), paired with our density estimation by σ128 based on 3900 samples (right).</p></span></div></div><p class="paragraph" style="text-align:left;">This paper discovered how to mathematically construct highly accurate models directly on unknown, complex data surfaces, known as manifolds. This method bypasses the need to map out the entire geometry of the dataset first; it only requires knowing the dimension of the data. </p><p class="paragraph" style="text-align:left;">The team unlocked a breakthrough in transfer learning by figuring out how to successfully lift learned information from just a localized portion of one data space and apply it to a completely different domain. As a result, adapting a massive model to a new problem no longer requires processing the entire original dataset.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/545991e8-e24f-4603-ba38-67eec3f1b16f/CleanShot_2026-03-04_at_21.02.32_2x.png?t=1772638363"/></div><p class="paragraph" style="text-align:left;">Finally, the researchers reimagined data classification by treating it like a signal separation problem. By mathematically estimating the underlying sources of these signals, their new algorithm quickly zeroes in on the most informative data points. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.17985?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="docto-lo-ra-learning-to-instantly-i">Doc-to-LoRA: Learning to Instantly Internalize Contexts</h2><p class="paragraph" style="text-align:left;"><i>Charakorn et al. [Sakana AI, Minerva University]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1.4k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LoRA </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Let’s assume you asked an AI to analyze a massive technical manual. Currently, every time you ask a follow-up question, the system has to re-read the entire document. This repetitive reading eats up massive amounts of computing power, memory, and time.</p><p class="paragraph" style="text-align:left;">While researchers can technically train the AI to memorize the document permanently, that traditional training process is painfully slow, expensive, and completely impractical for quick updates.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bbb89628-6d3d-4f78-9864-1885cb320838/CleanShot_2026-03-04_at_21.12.40_2x.png?t=1772638972"/></div><p class="paragraph" style="text-align:left;">To solve this, researchers developed a brilliant workaround called Doc-to-LoRA, or D2L. Instead of forcing the main AI to constantly re-read text or undergo grueling training, they built a specialized, lightweight helper system called a hypernetwork.</p><p class="paragraph" style="text-align:left;">This helper reads the document exactly once and instantly generates a tiny, customized plug-in. Think of it like instantly downloading a new skill directly into the AI&#39;s brain. Once this plug-in is attached, the main AI can answer subsequent queries fluidly without ever needing the original text in its prompt. It performs this complex mental compression in just a single step, completely bypassing the steep costs of standard training.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a5557701-506c-43e2-b515-718c6c17edd8/CleanShot_2026-03-04_at_21.13.00_2x.png?t=1772638992"/><div class="image__source"><span class="image__source_text"><p>QA performance on SQuAD compared to the used context length ratio (left), update latency (middle), and additional memory needed for model updates (right).</p></span></div></div><p class="paragraph" style="text-align:left;">The results are incredibly promising. In testing, D2L successfully hunted down specific facts hidden inside massive walls of text, achieving near-perfect accuracy on documents over four times larger than the AI’s normal limits.</p><p class="paragraph" style="text-align:left;">It works drastically faster and uses far less memory than previous memorization methods. It can even translate visual information from image-based models into these text plug-ins, allowing a text-only AI to classify images.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d1e8c8ec-6143-43e9-8c20-2e131d5ae1d5/CleanShot_2026-03-04_at_21.13.32_2x.png?t=1772639027"/><div class="image__source"><span class="image__source_text"><p>Long document QA performance. LLMLingua-2 compresses the input with [20%, 40%, 60%, 80%, 90%] compression rates from right to left (gray dots).</p></span></div></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.15902?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="the-geometry-of-noise-why-diffusion">The Geometry of Noise: Why Diffusion Models Don&#39;t Need Noise Conditioning</h2><p class="paragraph" style="text-align:left;"><i>Sahraee-Ardakan et al. [Google]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 382 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Diffusion Models </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Can you restore a severely damaged painting without knowing how much damage was originally done? Standard diffusion models avoid this problem by relying on a strict timer that tells them exactly how much &quot;noise&quot; or corruption they are dealing with at any given step.</p><p class="paragraph" style="text-align:left;">Recently, researchers have been incredibly hopeful about &quot;autonomous&quot; models that strip away this timer, learning a single rule to handle everything from pure static to nearly perfect data. However, this creates a profound mathematical paradox. </p><p class="paragraph" style="text-align:left;">As these blind models approach the clean data, the underlying mathematical landscape forms an infinitely deep pit. The directional signals diverge completely, creating a severe geometric singularity. By all conventional logic, these models should become hopelessly unstable and crash.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c0930cfa-02fa-4142-bee5-dd014a509eb1/CleanShot_2026-03-04_at_21.14.23_2x.png?t=1772639072"/><div class="image__source"><span class="image__source_text"><p>The Singular Geometry of the Marginal Energy Landscape.</p></span></div></div><p class="paragraph" style="text-align:left;">Yet, researchers have beautifully resolved this mystery. They discovered that these autonomous systems are actually charting a course across a unified map called &quot;Marginal Energy.&quot;</p><p class="paragraph" style="text-align:left;">More importantly, the scientists proved that the models naturally develop a hidden geometric shock absorber. As the AI approaches that infinitely deep mathematical pit, this built-in feature perfectly counteracts the extreme steepness. It transforms a catastrophic plunge into a smooth, stable descent known as a Riemannian gradient flow.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28dbfec3-260b-496d-9501-8ecacb8f0570/CleanShot_2026-03-04_at_21.14.50_2x.png?t=1772639104"/><div class="image__source"><span class="image__source_text"><p>Generative performance on Fashion MNIST.</p></span></div></div><p class="paragraph" style="text-align:left;">Models attempting to directly predict the noise act like faulty amplifiers, magnifying tiny errors until the system catastrophically breaks down. Conversely, models based on predicting &quot;velocity&quot; inherently absorb that uncertainty into a smooth, stable drift. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.18428?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">dLLM: Simple Diffusion Language Modeling</h2><p class="paragraph" style="text-align:left;"><i>Zhou et al. [UC Berkeley, UIUC]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 120 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Diffusion LLM </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Language models traditionally generate text strictly left to right. But recently, researchers have found a promising alternative: diffusion language models. These systems can generate words in any order and iteratively refine their answers, unlocking highly flexible AI.</p><div class="embed"><a class="embed__url" href="https://github.com/ZHZisZZ/dllm?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank"><div class="embed__content"><p class="embed__title"> GitHub - ZHZisZZ/dllm: dLLM: Simple Diffusion Language Modeling </p><p class="embed__description"> dLLM: Simple Diffusion Language Modeling. Contribute to ZHZisZZ/dllm development by creating an account on GitHub. </p></div><img class="embed__image embed__image--right" src="https://opengraph.githubassets.com/cf5bb5f1995e865bce2db24a0df4ad74a8833f65b1be06aa1997556ee01b71c2/ZHZisZZ/dllm"/></a></div><p class="paragraph" style="text-align:left;">However, the underlying code was scattered across complex, isolated research repositories, making it incredibly difficult for developers to reproduce results or build upon each other’s work. To solve this, researchers created dLLM, a unified open-source framework that elegantly standardizes the development pipeline so the community can innovate together.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6f00097b-7faf-4531-85f6-5f94d3c16f32/CleanShot_2026-03-04_at_21.15.35_2x.png?t=1772639149"/><div class="image__source"><span class="image__source_text"><p>Inference pipeline: sampler swap from vanilla to FastdLLM MDLM sampler.</p></span></div></div><p class="paragraph" style="text-align:left;">The framework seamlessly connects the core pillars of AI development: training, generation, and testing. Using a highly modular design, dLLM allows developers to snap different components together effortlessly. A researcher can easily swap out a training method or plug in a high-speed generation algorithm without rewriting the model&#39;s core architecture.</p><p class="paragraph" style="text-align:left;">By standardizing how these models are evaluated, researchers also uncovered a hidden quirk: diffusion models are intensely sensitive to tiny adjustments in generation settings. A single tweaked parameter can drastically alter a model&#39;s performance, highlighting exactly why a transparent, shared testing environment is vital for meaningful progress.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/594a3fcc-0f64-435a-991e-e6dfc3ce600b/CleanShot_2026-03-04_at_21.16.01_2x.png?t=1772639171"/><div class="image__source"><span class="image__source_text"><p>Terminal Visualizer showing transition from masked to decoded tokens.</p></span></div></div><p class="paragraph" style="text-align:left;">Perhaps the most hopeful breakthrough is how this framework democratizes AI research. The team proved that building these dynamic models does not require massive supercomputers. Using their new recipes, they successfully transformed standard, off-the-shelf systems (including traditional discriminative architectures like BERT) into functional diffusion chatbots. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cb8c5c67-b464-432e-8919-46e3ca1f9f61/CleanShot_2026-03-04_at_21.16.23_2x.png?t=1772639200"/><div class="image__source"><span class="image__source_text"><p>Sensitivity to decoding hyperparameters.</p></span></div></div><p class="paragraph" style="text-align:left;">They achieved this with minimal computing power and simple fine-tuning, requiring no structural changes to the original models.</p><div class="embed"><a class="embed__url" href="https://huggingface.co/dllm-hub?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora" target="_blank"><img class="embed__image embed__image--left" src="https://cdn-thumbnails.huggingface.co/social-thumbnails/dllm-hub.png"/><div class="embed__content"><p class="embed__title"> dllm-hub (dLLM) </p><p class="embed__description"> Org profile for dLLM on Hugging Face, the AI community building the future. </p></div></a></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.22661?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=compress-context-into-a-lora"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/qFttD0060QA" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=7c762817-d958-44d2-a40b-cb9a8160ccdf&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Google Presents A Brand New Way To Train Latents</title>
  <description>plus more about Experiential RL, GLM-5 Report, and Attention Matching</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/12aeafde-3816-4305-8cf8-c2cfb5de4a4c/issue_96.jpg" length="167350" type="image/jpeg"/>
  <link>https://mail.bycloud.ai/p/google-presents-a-brand-new-way-to-train-latents</link>
  <guid isPermaLink="true">https://mail.bycloud.ai/p/google-presents-a-brand-new-way-to-train-latents</guid>
  <pubDate>Tue, 24 Feb 2026 19:15:22 +0000</pubDate>
  <atom:published>2026-02-24T19:15:22Z</atom:published>
    <dc:creator>by cloud</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><h6 class="heading" style="text-align:left;" id="nov-18-th-nov-24-th-33-latest-ai-re"><i>Feb 19th ~ Feb 24th</i><br><i>#96 Latest AI Research Explained Simply</i></h6><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="industry-news-in-1-line">🗞️ Industry News in 1 Line</h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 8.2k</span></span> <a class="link" href="https://x.com/Replit/status/2024578806208745637?s=20&utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" target="_blank" rel="noopener noreferrer nofollow">Replit has introduced Replit Animation</a>, a new tool that lets users create animated videos in minutes using conversational prompts. It is powered by Gemini 3.1 Pro and makes it easier to produce polished, shareable content without traditional editing software.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/63b38e00-c1a8-45c6-8d31-40180c9b4379/CleanShot_2026-02-24_at_21.31.48_2x.png?t=1771948927"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 22k</span></span> Anthropic has released Claude Sonnet 4.6, which comes with stronger performance in complex spreadsheet tasks, multi-step web forms, and expanded integrations through Excel MCP connectors. The Claude API now also supports <a class="link" href="https://claude.com/blog/improved-web-search-with-dynamic-filtering?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" target="_blank" rel="noopener noreferrer nofollow">more accurate web search</a>, dynamic filtering, and general availability of code execution, memory, and programmatic tool use.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/88ebd7bb-7fd3-4c53-8ec0-5f79625645e8/1206645ef5a618dabce8587b472b21c67a30a0db-3840x1948.webp?t=1771949348"/></div></li><li><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;">♥ 49k</span></span> <a class="link" href="https://www.anthropic.com/news/claude-code-security?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" target="_blank" rel="noopener noreferrer nofollow">Anthropic has announced Claude Code Security</a> to scan codebases for vulnerabilities and suggest targeted patches for human review. This release triggered a sharp selloff in cybersecurity stocks, with companies like JFrog, CrowdStrike, Okta, and Cloudflare all recording declines. The system uses reasoning to trace data flows and flag subtle errors, while enforcing a human-in-the-loop safeguard to ensure developer oversight. If you want to try it, then <a class="link" href="https://claude.com/solutions/claude-code-security?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" target="_blank" rel="noopener noreferrer nofollow">join the waitlist here</a>.</p></li></ol><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:transparent;border-color:#2C81E5;border-style:solid;border-width:5px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;">Learn LLMs Intuitively - Intuitive AI Academy</h2><div class="image"><a class="image__link" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/734c79dc-aaa6-46ce-ac7d-41a5f4d84381/image.png?t=1769003669"/></a></div><p class="paragraph" style="text-align:left;">Want to learn about LLMs, but never have a good place to start?</p><p class="paragraph" style="text-align:left;">My latest project: Intuitive AI Academy has the perfect starting point for you! We focus on<b> building your intuition to understand LLMs</b>, from transformer components, to post-training logic. All in one place.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/377a8b3f-2557-4a32-8115-242eb6a2146e/image.png?t=1769004014"/><div class="image__source"><span class="image__source_text"><p>content overview</p></span></div></div><p class="paragraph" style="text-align:left;">We currently have a early bird offer, where you would get 40% off yearly plan for our early users. </p><p class="paragraph" style="text-align:left;">Use code: <b>TIMELINE</b></p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://www.intuitiveai.academy/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents"><span class="button__text" style=""> Check Out Intuitive AI Academy </span></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://theaitimeline.carrd.co/?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents" target="_blank" rel="noopener noreferrer nofollow">Advertise with The AI Timeline! </a></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="unified-latents-ul-how-to-train-you">Unified Latents (UL): How to train your latents</h2><p class="paragraph" style="text-align:left;"><i>Heek [Google DeepMind Amsterdam]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 2.1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> ALDiffusion </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">To create high-quality images and videos efficiently, models usually compress data into a &quot;latent&quot; space (a digital shorthand that represents the original image in a smaller, more manageable package). If you compress the data too much, the AI loses fine details like textures and sharp edges; if you don&#39;t compress it enough, the AI becomes incredibly slow and expensive to train.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d95f0533-b5ca-4584-9088-db02e704fa9c/CleanShot_2026-02-23_at_17.10.28_2x.png?t=1771846840"/><div class="image__source"><span class="image__source_text"><p>Schematic overview of our model, include the Encoder, the prior latent diffusion model, and the diffusion decoder model.</p></span></div></div><p class="paragraph" style="text-align:left;">Traditional methods often rely on manual tuning to strike this balance, which is often more of an art than a science. Researchers have recently introduced a framework called Unified Latents (UL) to turn this guesswork into a systematic, more efficient process. By co-training the compression and generation steps together, they have found a way to maintain stunningly high-quality details while actually lowering the computational cost required for training.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/87043c06-a8f6-4960-97b8-5cd4ff3956ef/CleanShot_2026-02-23_at_17.11.12_2x.png?t=1771846881"/></div><p class="paragraph" style="text-align:left;">Instead of treating the compressed representation as a static container, Unified Latents uses a &quot;diffusion prior&quot; to monitor and regularize the information flow. By linking the noise level produced during the encoding process directly to the precision of the diffusion model, researchers could create a mathematically tight way to control the latent bitrate.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cb8077f4-0de0-4ad7-9919-3391d102c993/CleanShot_2026-02-23_at_17.11.31_2x.png?t=1771846903"/><div class="image__source"><span class="image__source_text"><p>A selection of samples from a text-to-image trained with Unified Latents</p></span></div></div><p class="paragraph" style="text-align:left;">They also paired this with a diffusion-based decoder, which is remarkably better at reconstructing high-frequency details than previous methods. This unified approach allows the system to navigate the trade-off between compression and quality with much greater precision.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.17270?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="fast-kv-compaction-via-attention-ma">Fast KV Compaction via Attention Matching</h2><p class="paragraph" style="text-align:left;"><i>Zweiger et al. [Massachusetts Institute of Technology]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 196 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> Attention </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> bycloud’s pick </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">AI models can take on more complex tasks like long-form coding and multi-day conversations, but they face a significant memory hurdle known as the key-value (KV) cache bottleneck. Every word the model processes adds to a digital &quot;short-term memory&quot; that can quickly balloon into several gigabytes of data. </p><p class="paragraph" style="text-align:left;">Until now, researchers have managed this by either summarizing the text (which often strips away vital nuances) or by using expensive optimization techniques that take hours of computing time to compress a single document.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7c67a3bb-49f1-4d53-8ac7-bbec52deb6a5/CleanShot_2026-02-23_at_17.16.12_2x.png?t=1771847181"/></div><p class="paragraph" style="text-align:left;">This paper has introduced a technique called Attention Matching that changes how we think about memory compaction. Instead of relying on slow, iterative training to condense information, this approach treats memory like a mathematical puzzle that can be solved directly in &quot;latent space&quot;.</p><p class="paragraph" style="text-align:left;">By focusing on how a model &quot;pays attention&quot; to specific pieces of information, the researchers found they could create a compact version of the memory that mimics the original&#39;s behavior. They discovered that this problem can be broken down into smaller sub-problems with efficient, closed-form solutions, allowing them to bypass the slow trial-and-error process of traditional machine learning.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ab8f5664-b736-4ff6-881f-5072e0016409/CleanShot_2026-02-23_at_17.16.30_2x.png?t=1771847205"/><div class="image__source"><span class="image__source_text"><p>Accuracy vs. compaction ratio across methods.</p></span></div></div><p class="paragraph" style="text-align:left;">This new framework can shrink a model&#39;s memory by up to 50 times in just a matter of seconds, rather than hours, with almost no impact on the quality of the model&#39;s output. </p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.16284?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="experiential-reinforcement-learning">Experiential Reinforcement Learning</h2><p class="paragraph" style="text-align:left;"><i>Shi et al. [KRAFTON, University of Wisconsin–Madison, UC Berkeley, Microsoft Research]</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 1k </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM RL </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Teaching artificial intelligence to navigate complex tasks is often a game of high-stakes guessing. In standard reinforcement learning, a model typically receives a single &quot;reward&quot; signal. This makes it incredibly difficult for the AI to pinpoint exactly where it tripped up or how to adjust its behavior for the next try.</p><p class="paragraph" style="text-align:left;">Researchers recognized that this &quot;blind&quot; trial-and-error is far less efficient than how humans naturally learn. When we fail at a task, we don&#39;t just try again at random; we stop, reflect on what went wrong, and form a mental plan to do better. To solve this, researchers introduced Experiential Reinforcement Learning (ERL), a new approach designed to turn silent failures into structured, durable lessons.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aeff92cc-0799-4f0f-9a31-c3d7ff12a915/CleanShot_2026-02-23_at_17.26.15_2x.png?t=1771847784"/><div class="image__source"><span class="image__source_text"><p>In Experiential Reinforcement Learning (ERL), instead of learning from feedback or outcome directly</p></span></div></div><p class="paragraph" style="text-align:left;">The researchers developed a clever &quot;experience-reflection-consolidation&quot; loop that embeds human-like reasoning directly into the training process. Instead of moving on immediately after an attempt, the model receives feedback from its environment and is prompted to generate a verbal reflection on its own performance.</p><p class="paragraph" style="text-align:left;">It looks at its errors and produces a self-critique that guides a refined second attempt at the same task. If this second try succeeds, the model &quot;internalizes&quot; the successful correction. This is the breakthrough moment: by training the model to reproduce the improved behavior from the original task alone, the AI eventually learns to skip the reflection step entirely.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5aa552dc-9e18-4470-a671-d245bb33690c/CleanShot_2026-02-23_at_17.26.47_2x.png?t=1771847818"/><div class="image__source"><span class="image__source_text"><p>Conceptual comparison of learning dynamics in RLVR and Experiential Reinforcement Learning (ERL)</p></span></div></div><p class="paragraph" style="text-align:left;">This ensures that the final model is both smarter and faster, maintaining high performance at deployment without any extra computational cost.</p><p class="paragraph" style="text-align:left;">In complex multi-step tasks like Sokoban, which require deep planning and spatial reasoning, ERL improved performance by a staggering 81% over standard methods. It also showed reliable gains in agentic reasoning tasks that involve using external tools to answer questions. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4da17c0b-eb2c-4ba5-895a-f81c534edd59/CleanShot_2026-02-23_at_17.27.15_2x.png?t=1771847847"/><div class="image__source"><span class="image__source_text"><p>Overview of Experiential Reinforcement Learning (ERL).</p></span></div></div><p class="paragraph" style="text-align:left;">By allowing the AI to accumulate &quot;corrective knowledge&quot; in a persistent memory, researchers have created a way for models to build on their past successes rather than repeating the same mistakes.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e57f4395-edd3-42e9-83bb-02865e4ff28c/CleanShot_2026-02-23_at_17.27.40_2x.png?t=1771847869"/></div><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.13949?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><h2 class="heading" style="text-align:left;" id="not-all-bits-are-equal-scale-depend">GLM-5: from Vibe Coding to Agentic Engineering</h2><p class="paragraph" style="text-align:left;"><i>Zhipu AI & Tsinghua University</i></p><p class="paragraph" style="text-align:left;"><span style="background-color:#e0e0e0;"><span style="color:rgb(255, 58, 58);font-size:0.6rem;"> ♥ 290 </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span><span style="background-color:#e0e0e0;"><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> LLM Agents </span></span><span style="color:rgb(44, 129, 229);font-size:0.6rem;"> </span></p><p class="paragraph" style="text-align:left;">Researchers are working to shift our relationship with AI from &quot;vibe coding&quot; (where humans provide constant prompts) toward a more autonomous era of &quot;agentic engineering.&quot; Traditional models often struggle with the sheer computational cost of keeping track of long conversations, or they lose their way during tasks that require hours of planning and execution.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/34475712-6dcc-4869-bb82-35137a00e8d3/bench.png?t=1771848560"/></div><p class="paragraph" style="text-align:left;">GLM-5 was designed to overcome these bottlenecks by creating a more independent assistant that can plan, implement, and iterate on technical challenges with minimal human intervention.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/49ee0681-4b06-4479-9225-9cab867c60a6/realworld_bench.png?t=1771848577"/></div><p class="paragraph" style="text-align:left;">To make this possible, the researchers introduced a dynamic mechanism called DeepSeek Sparse Attention. Rather than the model straining to analyze every single word in a massive document with equal intensity (which is incredibly expensive and slow), it now identifies which tokens are truly important for the task at hand.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ab27461c-7c68-47be-827e-776bd9b3448d/CleanShot_2026-02-23_at_17.40.29_2x.png?t=1771848644"/></div><p class="paragraph" style="text-align:left;">This allows the model to manage massive amounts of data, such as entire codebases or long-term business simulations, while significantly reducing the hardware power required. The model also features an &quot;interleaved thinking&quot; process where it pauses to reason before every action it takes.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/deec37e5-87f7-48b7-91fb-649be1ed9748/vending_bench.png?t=1771848588"/></div><p class="paragraph" style="text-align:left;">It can &quot;preserve&quot; these thoughts across a long conversation, ensuring it doesn&#39;t lose its train of thought when moving between different stages of a project. By using a new asynchronous training infrastructure that allows the model to learn from complex, it was able to reach an <b>unprecedented ability</b> to solve end-to-end software engineering challenges.</p><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="" href="https://arxiv.org/abs/2602.15763?utm_source=mail.bycloud.ai&utm_medium=newsletter&utm_campaign=google-presents-a-brand-new-way-to-train-latents"><span class="button__text" style=""> Read Full Paper </span></a></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><p class="paragraph" style="text-align:left;"></p></div><iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="true" class="youtube_embed" frameborder="0" height="100%" src="https://youtube.com/embed/httnhdpu_W4" width="100%"></iframe></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=81de0e5a-b3cd-49dd-9898-aaa969440917&utm_medium=post_rss&utm_source=the_ai_timeline">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
