<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>LLMs Research</title>
    <description>Daily newsletter categorizing &amp; easily explaining LLMs research papers as they published.</description>
    
    <link>https://llm.beehiiv.com/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/Apn7agm48K.xml" rel="self"/>
    
    <lastBuildDate>Sun, 13 Sep 2026 11:33:39 +0000</lastBuildDate>
    <pubDate>Sat, 22 Mar 2025 13:00:00 +0000</pubDate>
    <atom:published>2025-03-22T13:00:00Z</atom:published>
    <atom:updated>2026-09-13T11:33:39Z</atom:updated>
    
      <category>Data Science</category>
      <category>Machine Learning</category>
      <category>Artificial Intelligence</category>
    <copyright>Copyright 2026, LLMs Research</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/a15a9679-18ec-43e6-8e45-6c2b064f95d6/new_logo.png</url>
      <title>LLMs Research</title>
      <link>https://llm.beehiiv.com/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>LLM Research Highlights: March 1-15, 2025 [ Part 2/2 ]</title>
  <description>Exploring Innovations in reasoning and context length improvement for Large Language Models (LLMs)</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4f5a3b53-e4cc-4f55-b412-19d0e1690cff/March_1_to_15_part2.png" length="380044" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llm-research-highlights-march-1-15-2025-part-2-2</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llm-research-highlights-march-1-15-2025-part-2-2</guid>
  <pubDate>Sat, 22 Mar 2025 13:00:00 +0000</pubDate>
  <atom:published>2025-03-22T13:00:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ul><li><p class="paragraph" style="text-align:left;">Chain of Draft (CoD) cuts LLM token usage by up to 92.4% while retaining 91% accuracy on GSM8k. </p></li><li><p class="paragraph" style="text-align:left;">LADDER boosts a 3B LLM’s math accuracy from 1% to 82% via self-generated problem variants. </p></li><li><p class="paragraph" style="text-align:left;">LMM-R1 enhances 3B multimodal LLMs, improving MathVerse scores by 4.83% with rule-based RL. </p></li><li><p class="paragraph" style="text-align:left;">SoRFT-Qwen-7B resolves 21.4% of software issues, outperforming larger models on SWE-Bench. </p></li><li><p class="paragraph" style="text-align:left;">TOKENSWIFT speeds up 100K-token generation by 3×, reducing LLaMA3.1-8B time from 5 hours to 90 minutes.</p></li></ul><div class="recommendation" id="a95b1f62-6a5b-4696-b999-0c62a751f6e9"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/4f5a3b53-e4cc-4f55-b412-19d0e1690cff/March_1_to_15_part2.png?t=1742609157"/></figure><h3 class="recommendation__title"> Reasoning and context length improvement of large language models </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOlwiI2RlY2RjZFwiLFwiYmFja2dyb3VuZFRoZW1lXCI6XCJsaWdodFwiLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2L2E5NWIxZjYyLTZhNWItNDY5Ni1iOTk5LTBjNjJhNzUxZjZlOS9QYXJ0JTIwMl8lMjBNYXJjaCUyMDFzdCUyMHRvJTIwMTV0aC53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDBaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9ZDhmODhmMDlhNmE3MDA5OGYxMTY4NTZjMWJjNzUxODY4YzFiNGUxOWI4YzJjOGRlMDQxNGZlZmI0YTM2NWVjN1wiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlLzRmNWEzYjUzLWU0Y2MtNGY1NS1iNDEyLTE5ZDBlMTY5MGNmZi9NYXJjaF8xX3RvXzE1X3BhcnQyLnBuZz90PTE3NDI2MDkxNTdcIixcInRpdGxlXCI6XCJSZWFzb25pbmcgYW5kIGNvbnRleHQgbGVuZ3RoIGltcHJvdmVtZW50IG9mIGxhcmdlIGxhbmd1YWdlIG1vZGVsc1wifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h1 class="heading" style="text-align:center;"><b>Reasoning</b></h1><h2 class="heading" style="text-align:left;"><b>Chain of Draft: Thinking Faster by Writing Less</b></h2><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2502.18600?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2502.18600</a><br><b>Authors:</b> Silei Xu et al. (Zoom Communications)<br><b>Focus:</b> Reducing verbosity in LLM reasoning while maintaining accuracy<br><b>Code:</b> <a class="link" href="https://github.com/sileix/chain-of-draft?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://github.com/sileix/chain-of-draft</a></p><p class="paragraph" style="text-align:left;">This research paper introduces a novel prompting strategy, Chain of Draft (CoD), designed to make LLMs reason more efficiently. Unlike the verbose Chain-of-Thought (CoT) prompting, CoD mimics human shorthand by encouraging concise, essential intermediate reasoning steps, cutting down on token usage and latency.</p><p class="paragraph" style="text-align:left;"><b>Key Innovations:</b> </p><p class="paragraph" style="text-align:left;"><b>Minimalistic Reasoning:</b> CoD prompts LLMs to produce brief, information-dense drafts at each step, reducing unnecessary elaboration. </p><p class="paragraph" style="text-align:left;"><b>Efficiency Optimization:</b> It achieves CoT-level accuracy with significantly fewer tokens, enhancing inference speed and lowering computational costs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Evaluated on benchmarks like GSM8k, date understanding, and coin flipping, CoD reduces token usage by up to 92.4% compared to CoT. For instance, on GSM8k, it maintains over 91% accuracy while cutting output tokens by 80% (e.g., from 205.1 to 43.9 for GPT-4o). Latency drops by 76.2% for GPT-4o and 48.4% for Claude 3.5 Sonnet, making CoD ideal for real-time applications without compromising reasoning quality.</p><hr class="content_break"><h2 class="heading" style="text-align:left;"><b>LADDER: Self-Improving LLMs Through Recursive Problem Decomposition</b></h2><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2503.00735?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.00735</a><br><b>Authors:</b> Toby Simonds et al. (Tufa Labs)<br><b>Focus:</b> Enabling LLMs to autonomously enhance reasoning via self-generated problem variants </p><p class="paragraph" style="text-align:left;">This paper presents a framework called LADDER, that allows LLMs to self-improve by recursively breaking down complex problems into simpler variants. This autonomous learning process, guided by reinforcement learning (RL) and numerical verification, eliminates the need for human supervision or curated datasets.</p><p class="paragraph" style="text-align:left;"><b>Key Innovations:</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8baf6bba-8edf-4f3a-b803-9a789542c018/image.png?t=1742607016"/></div><p class="paragraph" style="text-align:left;"><b>Recursive Variant Generation:</b> LLMs create a tree of progressively simpler problem versions, forming a difficulty gradient for incremental learning. </p><p class="paragraph" style="text-align:left;"><b>Test-Time Reinforcement Learning (TTRL):</b> At inference, TTRL refines solutions by applying RL to problem-specific variants, boosting performance dynamically.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> On mathematical integration tasks, LADDER improves a Llama 3.2 3B model’s accuracy from 1% to 82% on undergraduate-level problems. A Qwen2.5 7B Deepseek-R1 Distilled model achieves 73% on the 2025 MIT Integration Bee qualifying exam, surpassing GPT-4o (42%). With TTRL, accuracy rises to 90%, outpacing OpenAI’s o1 (80%), demonstrating the power of self-directed learning for complex reasoning.</p><hr class="content_break"><h2 class="heading" style="text-align:left;"><b>LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL</b></h2><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2503.07536?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.07536</a><br><b>Authors:</b> Yingzhe Peng et al. (Southeast University, Ant Group, et al.)<br><b>Focus:</b> Enhancing reasoning in compact 3B-parameter multimodal LLMs<br><b>Code:</b> <a class="link" href="https://github.com/TideDra/lmm-r1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://github.com/TideDra/lmm-r1</a></p><p class="paragraph" style="text-align:left;">&quot;LMM-R1&quot; proposes a two-stage rule-based RL framework to boost reasoning in 3B-parameter large multimodal models (LMMs), overcoming their limited capacity and the scarcity of high-quality multimodal reasoning data. It first strengthens foundational reasoning with text-only data, then generalizes it to multimodal contexts.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/56b17c91-31dd-439d-8fec-1350822017aa/image.png?t=1742607100"/></div><p class="paragraph" style="text-align:left;"><b>How:</b> &quot;LMM-R1&quot; proposes two stage RL which is a foundational reasoning enhancement (FRE) uses text-only data to build reasoning skills, followed by multimodal generalization training (MGT) for multimodal application. It also proposes a rule based rewards which combines format and accuracy rewards to guide learning efficiently, avoiding extensive human annotation.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Applied to Qwen2.5-VL-Instruct-3B, LMM-R1 achieves average gains of 4.83% on multimodal benchmarks (e.g., MathVerse: 34.64% to 41.55%) and 4.5% on text-only benchmarks (e.g., MATH500: 63.4% to 65.8%). On the complex Football Game task, it improves by 3.63% (15.36 to 18.99), outperforming larger models like GPT-4o (21.20). This approach proves text-based reasoning can effectively transfer to multimodal domains, offering a cost-efficient training strategy.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching</b></p><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2503.05179?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.05179</a><br><b>Authors:</b> Not specified (PDF unavailable)<br><b>Focus:</b> Improving LLM reasoning efficiency with cognitive-inspired sketching </p><p class="paragraph" style="text-align:left;">This paper introduces a method, Sketch-of-Thought, to enhance reasoning efficiency in LLMs. Drawing from cognitive processes, it uses adaptive sketching to streamline problem-solving, though detailed insights are limited due to the unavailable PDF.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f41731f8-a468-44b4-95a9-48641d906cdd/image.png?t=1742607161"/></div><p class="paragraph" style="text-align:left;"><b>How:</b> This paper proposes adaptive sketching. It likely employs concise, cognitive-inspired sketches to guide reasoning, reducing computational overhead. Also, it aims to maintain accuracy while speeding up inference, though specifics are unclear without the full paper.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Without access to the paper, specific results cannot be detailed. Based on the abstract, it likely demonstrates improved reasoning efficiency across various tasks, aligning with the trend of optimizing LLM performance for practical use.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h1 class="heading" style="text-align:center;"><b>Context length improvement</b></h1><p class="paragraph" style="text-align:left;"><b>SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning</b></p><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2502.20127?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2502.20127</a><br><b>Authors:</b> Zexiong Ma et al. (Peking University & ByteDance)<br><b>Focus:</b> Enhancing context handling for issue resolving in software development</p><p class="paragraph" style="text-align:left;">While not primarily focused on context length extension, <b>SoRFT</b> indirectly improves LLMs’ ability to manage complex, long-context tasks like software issue resolving. Targeting open-source models, it addresses cost and privacy concerns of commercial APIs by decomposing issue resolution into manageable subtasks: file localization, function localization, line localization, and code edit generation.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/74726982-f53a-45c4-a9d8-60e21c11e4a7/image.png?t=1742607268"/></div><p class="paragraph" style="text-align:left;"><b>Two-Stage Training:</b></p><ul><li><p class="paragraph" style="text-align:left;"><b>Rejection-Sampled Supervised Fine-Tuning (SFT):</b> Filters Chain-of-Thought (CoT) data using ground-truth from pull requests, teaching the model structured reasoning.</p></li><li><p class="paragraph" style="text-align:left;"><b>Rule-Based Reinforcement Learning (RL):</b> Employs Proximal Policy Optimization (PPO) with rewards based on β scores (prioritizing recall), refining the model’s precision across subtasks.</p></li><li><p class="paragraph" style="text-align:left;"><b>Context Relevance:</b> By breaking down repository-scale tasks, SoRFT enables LLMs to process extensive codebases effectively, leveraging long-context data from open-source projects.</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results:</b> On SWE-Bench Verified, SoRFT-Qwen-7B resolves 21.4% of issues, outperforming larger models like SWE-Gym-Qwen-32B (20.6%). It also boosts general code tasks (e.g., 90% on RepoQA vs. 85% baseline), showcasing enhanced long-context comprehension and generalization.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://arxiv.org/abs/2502.18890?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation up to 100K Tokens</a></b></p><p class="paragraph" style="text-align:left;"><b>Authors:</b> Tong Wu et al. (Shanghai Jiao Tong University & BIGAI)<br><b>Focus:</b> Accelerating ultra-long sequence generation<br><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2502.18890?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2502.18890</a></p><p class="paragraph" style="text-align:left;"><b>TOKENSWIFT</b>, introduced in this paper, revolutionizes the generation of ultra-long sequences (up to 100K tokens) by slashing processing time without compromising quality. Traditional autoregressive (AR) methods take hours (e.g., 5 hours for LLaMA3.1-8B), a bottleneck for applications like creative writing or reasoning traces.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5335af98-3761-4953-be3e-d381480ce487/image.png?t=1742607343"/></div><p class="paragraph" style="text-align:left;">The authors identify three hurdles:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Frequent Model Reloading:</b> Slows generation due to I/O bottlenecks.</p></li><li><p class="paragraph" style="text-align:left;"><b>Prolonged KV Cache Growth:</b> Increases complexity as sequences lengthen.</p></li><li><p class="paragraph" style="text-align:left;"><b>Repetitive Content:</b> Degrades quality in long outputs.</p></li></ol><p class="paragraph" style="text-align:left;">TOKENSWIFT counters these with:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Multi-Token Generation & Token Reutilization:</b> Generates multiple tokens per forward pass and reuses frequent n-grams, reducing reloads.</p></li><li><p class="paragraph" style="text-align:left;"><b>Dynamic KV Cache Management:</b> Updates partial caches iteratively, maintaining efficiency.</p></li><li><p class="paragraph" style="text-align:left;"><b>Contextual Penalty:</b> Mitigates repetition, enhancing diversity (Distinct-n scores improve).</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results:</b> TOKENSWIFT achieves over 3× speedup (e.g., 90 minutes vs. 5 hours for 100K tokens on LLaMA3.1-8B), with lossless accuracy across models (1.5B to 14B). It saves up to 5.54 hours on a 14B model, making ultra-long generation practical.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h1 class="heading" style="text-align:center;"><b>Explore Past Highlights</b></h1><p class="paragraph" style="text-align:left;">If you liked what you read and interested in reading more than here are our achieved newsletters we sent in past.</p><ul><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.llmsresearch.com/p/llm-research-highlights-march-1-15-2025-part1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">LLM Research Highlights: March 1–15, 2025 [Part 1/2]</a></b></p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://www.llmsresearch.com/p/breakthrough-papers-improving-llms-performance?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow">Breakthrough papers improving LLMs performance (February 16–28, 2025)</a></b></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/llms-application-february-15-28-2025?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow"><b>Creative way of using LLMs (February 15–28, 2025)</b></a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-1-3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow"><b>Research papers improving performance of LLMs [1/3] (January 16–February 15, 2025)</b></a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/exploring-new-llm-applications-from-january-1-15-2025-research-papers?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025-part-2-2" target="_blank" rel="noopener noreferrer nofollow"><b>Exploring New LLM Applications: Insights from January 1–15, 2025</b></a></p></li></ul></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=b8a355c3-2e1e-4140-af96-38ddeb66b7e3&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLM Research Highlights: March 1-15, 2025</title>
  <description>Exploring Innovations in Performance, Instruction Tuning, Cache Management, Quantization, and Unlearning for Large Language Models</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3280ae7-546e-49f8-a49f-399ae9cc2364/march_1_to_15__1_2_.png" length="405377" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llm-research-highlights-march-1-15-2025-part1</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llm-research-highlights-march-1-15-2025-part1</guid>
  <pubDate>Fri, 21 Mar 2025 10:30:00 +0000</pubDate>
  <atom:published>2025-03-21T10:30:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ul><li><p class="paragraph" style="text-align:left;"><span style="color:black;font-family:sans-serif;font-size:inherit;"><b>Performance Boosts</b></span><span style="color:black;font-family:sans-serif;font-size:inherit;">: Forgetting Transformer, Multi-Attempt RL, and R1-Searcher improve efficiency, math accuracy, and search with selective memory, feedback, and RL.</span></p></li><li><p class="paragraph" style="text-align:left;"><span style="color:black;font-family:sans-serif;font-size:inherit;"><b>Simplified Design</b></span><span style="color:black;font-family:sans-serif;font-size:inherit;">: Normalization-Free Transformers speed up training and inference using Dynamic Tanh in a streamlined architecture.</span></p></li><li><p class="paragraph" style="text-align:left;"><span style="color:black;font-family:sans-serif;font-size:inherit;"><b>Data Optimization</b></span><span style="color:black;font-family:sans-serif;font-size:inherit;">: RDS+ enhances instruction tuning, achieving top performance with only 6% of the data pool.</span></p></li><li><p class="paragraph" style="text-align:left;"><span style="color:black;font-family:sans-serif;font-size:inherit;"><b>Memory Efficiency</b></span><span style="color:black;font-family:sans-serif;font-size:inherit;">: Q-Filters and RSQ optimize long-context handling and quantization by compressing KV Cache and prioritizing key tokens.</span></p></li><li><p class="paragraph" style="text-align:left;"><span style="color:black;font-family:sans-serif;font-size:inherit;"><b>Compression & Fairness</b></span><span style="color:black;font-family:sans-serif;font-size:inherit;">: TinyR1-32B-Preview and Group-Robust Unlearning deliver high accuracy and equitable data removal via distillation and unlearning techniques.</span></p></li></ul><div class="recommendation" id="2bd64a30-83b2-4ddc-b647-5f92209f5fe6"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/f3280ae7-546e-49f8-a49f-399ae9cc2364/march_1_to_15__1_2_.png?t=1742543185"/></figure><h3 class="recommendation__title"> LLM Research Highlights: March 1-15, 2025 </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzJiZDY0YTMwLTgzYjItNGRkYy1iNjQ3LTVmOTIyMDlmNWZlNi9QYXJ0JTIwMV8lMjBNYXJjaCUyMDFzdCUyMC0lMjAxNXRoJTJDJTIwMjAyNS53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDBaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9MzRkOGJiNmJhNDZmM2U5NTJkYzQyM2Q5YzYzNjIzYWU5Y2U1MDU1YmE1MGQ0YTNlN2QyYzg3ZWRhMjVlZmE4N1wiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlL2YzMjgwYWU3LTU0NmUtNDlmOC1hNDlmLTM5OWFlOWNjMjM2NC9tYXJjaF8xX3RvXzE1X18xXzJfLnBuZz90PTE3NDI1NDMxODVcIixcInRpdGxlXCI6XCJMTE0gUmVzZWFyY2ggSGlnaGxpZ2h0czogTWFyY2ggMS0xNSwgMjAyNVwifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h1 class="heading" style="text-align:center;"><b>Core research</b></h1><p class="paragraph" style="text-align:left;"><b>Forgetting Transformer: Softmax Attention with a Forget Gate</b></p><p class="paragraph" style="text-align:left;"><b>Paper: </b><a class="link" href="https://arxiv.org/abs/2503.02130?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.02130</a><br><b>Authors:</b> Zhixuan Lin et al. (Mila & Université de Montréal)<br><b>Focus:</b> Enhancing Transformer performance with a data-dependent forgetting mechanism<br><b>Code: </b><a class="link" href="https://github.com/zhixuan-lin/forgetting-transformer?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/zhixuan-lin/forgetting-transformer</a></p><p class="paragraph" style="text-align:left;">The Forgetting Transformer (FoX) addresses a key limitation in standard Transformers: their lack of an explicit mechanism to selectively forget past information, a feature common in recurrent sequence models via forget gates. While Transformers excel at long-context tasks, they often retain irrelevant details, impacting efficiency and performance on both short and long sequences. FoX introduces a novel approach by integrating a forget gate into softmax attention.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/07da7a2c-316c-403a-a5e8-02328157388e/image.png?t=1742540453"/></div><p class="paragraph" style="text-align:left;"><b>Key contribution:</b></p><p class="paragraph" style="text-align:left;">Forgetting Attention: A scalar forget gate, computed as <span style="background-color:#d6c1c1;">f</span><span style="background-color:#d6c1c1;"><sub>t</sub></span><span style="background-color:#d6c1c1;">​=σ(w</span><span style="background-color:#d6c1c1;"><sub>f</sub></span><span style="background-color:#d6c1c1;"><sup>⊤</sup></span><span style="background-color:#d6c1c1;">​x</span><span style="background-color:#d6c1c1;"><sub>t</sub></span><span style="background-color:#d6c1c1;">​+b</span><span style="background-color:#d6c1c1;"><sub>f​</sub></span><span style="background-color:#d6c1c1;">)</span>, down-weights unnormalized attention scores in a data-dependent manner. This is applied as <span style="background-color:#d6c1c1;">F</span><span style="background-color:#d6c1c1;"><sub>ij</sub></span><span style="background-color:#d6c1c1;">​=∏</span><span style="background-color:#d6c1c1;"><sup>i</sup></span><span style="background-color:#d6c1c1;"><sub>l=j+1​</sub></span><span style="background-color:#d6c1c1;"><span style="font-family:&quot;Times New Roman&quot;,Baskerville,Georgia,serif;">f</span></span><span style="background-color:#d6c1c1;"><sub>l</sub></span>, allowing the model to prioritize relevant context dynamically.</p><p class="paragraph" style="text-align:left;"><b>Pro Block Design:</b> An enhanced architecture incorporating recurrent-inspired components like output gates and token shifts, boosting performance across tasks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Evaluated on LongCrawl64 with a 16K-token context, FoX outperforms the baseline Transformer on long-context language modeling (e.g., lower per-token loss), length extrapolation, and short-context downstream tasks (e.g., 50.85% avg. accuracy on LM-eval-harness vs. 50.79% for Transformer-Pro). It matches Transformer performance on long-context downstream tasks (LongBench) while requiring no positional embeddings and remaining compatible with FlashAttention. FoX’s ability to retain long-context retrieval (near-perfect needle-in-the-haystack scores) while improving efficiency marks a significant step forward for LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Learning from Failures in Multi-Attempt Reinforcement Learning</b></p><p class="paragraph" style="text-align:left;"><b>Paper: </b><a class="link" href="https://arxiv.org/abs/2503.04808?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.04808</a><br><b>Authors:</b> Stephen Chung et al. (DualityRL & Shanghai AI Lab)<br><b>Focus:</b> Boosting LLM reasoning through multi-attempt training with feedback </p><p class="paragraph" style="text-align:left;">This paper extends reinforcement learning (RL) for LLMs by shifting from single-turn question-answering to a multi-attempt framework, where models refine responses based on feedback after incorrect attempts. Inspired by DeepSeek R1, the authors argue that allowing multiple tries enhances reasoning by encouraging self-refinement, a capability often absent in single-turn trained models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1e73b629-80a4-4d0c-8965-6ac8fe6f1d3f/image.png?t=1742540534"/></div><p class="paragraph" style="text-align:left;"><b>Key Innovations:</b></p><p class="paragraph" style="text-align:left;"><b>Multi-Attempt Task:</b> The model gets <b>N</b> attempts (sampled from 1 to 5), with a transition function terminating the dialogue on a correct answer or exhausted attempts. Feedback prompts refinement after errors.</p><p class="paragraph" style="text-align:left;"><b>Reward Design:</b> +1 for a correct answer, -0.5 for a wrong answer in correct format, and -1 otherwise, incentivizing exploration and correction without penalizing attempt count.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Fine-tuning Qwen 2.5 Math 1.5B on 8K math questions, the multi-attempt LLM improves from 45.6% accuracy (1 attempt) to 52.5% (2 attempts) across five math benchmarks (e.g., AIME 2024, MATH 500). The single-turn baseline, in contrast, rises only from 42.3% to 43.2%. Even in single-attempt evaluations, the multi-attempt model edges out the baseline (45.4% vs. 43.5%), showcasing its superior adaptability and reasoning refinement—crucial for real-world LLM applications.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning</b></p><p class="paragraph" style="text-align:left;"><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2503.05592?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.05592</a><br><b>Authors:</b> Huatong Song et al. (Renmin University of China)<br><b>Focus:</b> Enhancing LLMs with autonomous external search capabilities<br><b>Code:</b> <a class="link" href="https://github.com/RUCAIBox/R1-Searcher?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/RUCAIBox/R1-Searcher</a></p><p class="paragraph" style="text-align:left;">R1-Searcher tackles a persistent LLM weakness: reliance on internal knowledge, which falters on time-sensitive or knowledge-intensive tasks, leading to hallucinations. By integrating retrieval-augmented generation (RAG) with a two-stage RL approach, it trains LLMs to invoke external search systems effectively, improving reasoning without supervised fine-tuning (SFT).</p><p class="paragraph" style="text-align:left;"><b>Contributions:</b></p><p class="paragraph" style="text-align:left;"><b>Two-Stage RL:</b> Stage 1 uses a retrieval reward (0.5 if search is invoked) to teach correct query formatting; Stage 2 adds an answer reward (F1 score-based) to optimize problem-solving with retrieved data.</p><p class="paragraph" style="text-align:left;"><b>RAG-based Rollout:</b> Special tags (e.g., &lt;begin_of_query&gt;) pause generation for retrieval, integrating results seamlessly into reasoning.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Trained on HotpotQA and 2WikiMultiHopQA, R1-Searcher (Qwen-2.5-7B-Base) outperforms GPT-40-mini-based ReARTeR by 48.2% on HotpotQA and 21.7% on 2Wiki (LLM-as-Judge scores). On the unseen Bamboogle dataset with online search, it achieves an 11.4% gain over Search-o1 (32B). This pure RL approach enhances generalization and inference efficiency, making LLMs more robust for complex, real-world queries.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Transformers without Normalization: Weight Space Analysis and Improved Training Dynamics</b></p><p class="paragraph" style="text-align:left;"><b>Paper:</b><a class="link" href="https://arxiv.org/abs/2503.10622?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.10622</a><br><b>Authors:</b> Jiachen Zhu et al. (Meta, NYU, MIT, Princeton)<br><b>Focus:</b> Eliminating normalization layers in Transformers with minimal performance trade-offs<br><b>Code: </b><a class="link" href="https://jiachenzhu.github.io/DyT?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://jiachenzhu.github.io/DyT</a></p><p class="paragraph" style="text-align:left;">Normalization layers like Layer Norm (LN) and RMSNorm are cornerstones of Transformer architectures, widely assumed to be critical for stable training and strong performance. &quot;Transformers without Normalization&quot; challenges this assumption, introducing a surprisingly simple alternative—<b>Dynamic Tanh (DyT)</b>—to replace normalization entirely. The authors observe that LN often produces tanh-like, S-shaped mappings, squashing extreme values while scaling inputs. DyT leverages this insight with an element-wise operation, <span style="background-color:#d6c1c1;">DyT(</span><span style="background-color:#d6c1c1;"><i>x</i></span><span style="background-color:#d6c1c1;">)=tanh(</span><span style="background-color:#d6c1c1;"><i>αx</i></span><span style="background-color:#d6c1c1;">)</span>, where <i>α</i> is a learnable parameter, bypassing the need for statistical computations.</p><p class="paragraph" style="text-align:left;"><b>Contribution:</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b440e4ee-67b8-49aa-95f1-68e6b63f3bf3/image.png?t=1742540849"/></div><p class="paragraph" style="text-align:left;"><b>Dynamic Tanh (DyT):</b> A lightweight replacement for normalization, DyT uses <span style="background-color:#d6c1c1;">tanh(</span><span style="background-color:#d6c1c1;"><i>αx</i></span><span style="background-color:#d6c1c1;">)</span> to mimic LN’s squashing and scaling effects, with <span style="background-color:#d6c1c1;"><i>α</i></span><i> </i>dynamically adjusting to input ranges. </p><p class="paragraph" style="text-align:left;"><b>Drop-in Simplicity:</b> DyT integrates seamlessly into existing Transformer designs, replacing LN or RMSNorm without altering other components or requiring extensive hyperparameter tweaks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b></p><p class="paragraph" style="text-align:left;">Tested across diverse domains—vision (ViT, ConvNeXt), language (LLaMA), speech (wav2vec 2.0), and DNA modeling (HyenaDNA)—DyT matches or exceeds the performance of normalized models. For LLaMA (7B to 70B), DyT achieves equivalent training loss and zero-shot accuracy on 15 tasks, while cutting training time by 8.2% and inference time by 7.8% on a 7B model. Its efficiency stems from avoiding mean/variance calculations, making it a leaner option. Unlike alternatives like Fixup or SkipInit, DyT maintains stability and performance without complex initialization tricks, using just a default <i>α = 0.5 </i>for most tasks (though tuned for LLMs). This approach redefines Transformer design, offering a faster, simpler alternative to normalization while preserving—or even enhancing—LLM capability, especially in resource-sensitive settings.</p><hr class="content_break"><p class="paragraph" style="text-align:left;">Discover more about recent research papers that enhance the performance of LLMs, published in the latter half of February 2025.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-1-3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-1-3</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-2-3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-2-3</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-3-3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://www.llmsresearch.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-3-3</a></p></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Instruction Tuning</b></h2><p class="paragraph" style="text-align:left;"><b>Large-Scale Data Selection for Instruction Tuning</b></p><p class="paragraph" style="text-align:left;"><b>Paper</b>: <a class="link" href="https://arxiv.org/abs/2503.01807?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.01807</a><br><b>Code</b>: <a class="link" href="https://github.com/hanishivi/automated-instruction-selection?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/hanishivi/automated-instruction-selection</a><br><b>Authors</b>: Hamish Ivison et al. (University of Washington, Allen Institute for AI, University of Southern California)<br><b>Focus</b>: Optimizing data selection for instruction-tuning at scale </p><p class="paragraph" style="text-align:left;">Instruction-tuning drives language model performance, but scaling data selection to millions of samples remains tricky. This paper evaluates nine automated methods on pools up to 5.8 million samples, introducing RDS+, a simple embedding approach that outperforms complex techniques across single and multi-task settings. This paper offers RDS+ method, which uses weighted mean pooling of LM hidden states, consistently beating advanced methods like LESS and IFD. This method thrives with larger datasets, outperforming human-curated mixtures and improving multi-task generalization.</p><p class="paragraph" style="text-align:left;"><b>Results</b>: On TULU 2, RDS+ selects 326,000 samples (6% of 5.8 million) to exceed the TULU 2 mixture and match full-pool performance, with 2-point gains over baselines and random selection. This efficient method shines at scale, enhancing large-scale instruction-tuning.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Cache management</b></h2><p class="paragraph" style="text-align:left;"><b>Q-Filters: Leveraging Query-Key Geometry for Efficient Key-Value Cache Compression</b></p><p class="paragraph" style="text-align:left;"><b>Paper</b>: <a class="link" href="https://arxiv.org/abs/2503.02812?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.02812</a><br><b>Code</b>: <a class="link" href="https://github.com/NathanGodey/qfilters?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/NathanGodey/qfilters</a><br><b>Authors</b>: Nathan Godey et al. (Sorbonne Université, Inria, Sapienza University of Rome, University of Edinburgh, Miniml.AI)<br><b>Focus</b>: Compressing the KV Cache for LLMs using Query-Key geometry </p><p class="paragraph" style="text-align:left;">The KV Cache’s memory demands hinder long-context LLMs. Q-Filters leverages Query-Key geometry to estimate attention scores and prune less critical KV pairs without attention weights, ensuring compatibility with FlashAttention for efficient compression.</p><p class="paragraph" style="text-align:left;"><b>Key contribution</b>: </p><ul><li><p class="paragraph" style="text-align:left;"><b>Geometric Insight</b>: Projects Keys onto Query eigenvectors to assess KV importance. </p></li><li><p class="paragraph" style="text-align:left;"><b>Training-Free & FlashAttention-Compatible</b>: Computes filters once, integrating with memory-efficient attention.</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results</b>: On Llama-3.1-8B, achieves 99% accuracy with 32x compression in needle-in-a-haystack tests and cuts perplexity drop by 65% over Streaming-LLM on the Pile with a 512-item cache.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h1 class="heading" style="text-align:center;"><b>Quantization</b></h1><h2 class="heading" style="text-align:left;"><b>RSQ: Learning from Important Tokens Leads to Better Quantized LLMs</b></h2><p class="paragraph" style="text-align:left;"><b>Paper</b>: <a class="link" href="https://arxiv.org/abs/2503.01820?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.01820</a><br><b>Authors</b>: Yi-Lin Sung et al. (University of North Carolina at Chapel Hill)<br><b>Focus</b>: Enhancing post-training quantization (PTQ) of LLMs by prioritizing important tokens<br><b>Code</b>: <a class="link" href="https://github.com/ylsung/rsq?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/ylsung/rsq</a></p><p class="paragraph" style="text-align:left;">RSQ introduces a layer-wise quantization method that boosts LLM compression by focusing on high-importance tokens (e.g., those with large attention scores). It uses a three-step process (Rotate, Scale, Quantize) and token importance strategies to maintain key information, cutting computational costs.</p><p class="paragraph" style="text-align:left;"><b>Key Innovations:</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/80447ec4-1191-4059-9097-60dbbda30948/image.png?t=1742542357"/></div><ul><li><p class="paragraph" style="text-align:left;"><b>Three-Step Process</b>: Rotate (reduces outliers), Scale (adjusts features by importance), Quantize (uses scaled Hessian). Rotate applies orthogonal transformations to mitigate weight outliers, reducing quantization errors. While scale adjusts token features based on their importance, using a modified loss function: <span style="background-color:#d6c1c1;">‖(WX - ŴX)R‖₂²</span>, where R scales tokens dynamically. and, Quantize uses the GPTQ framework with a scaled Hessian matrix (<span style="background-color:#d6c1c1;">H</span><span style="background-color:#d6c1c1;"><sub>RSQ</sub></span><span style="background-color:#d6c1c1;"> = 2XR²Xᵀ</span>) for efficient quantization.</p></li><li><p class="paragraph" style="text-align:left;"><b>Token Importance</b>: Heuristic (e.g., First-N) and dynamic (AttnCon) strategies improve accuracy.</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results</b>: On LLaMA3-8B-Instruct, RSQ gains 1.6% accuracy at 3-bit precision, with up to 3.0% improvement on long-context tasks over QuaRot, shining in extreme compression.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation</b></p><p class="paragraph" style="text-align:left;"><b>Paper</b>: <a class="link" href="https://arxiv.org/abs/2503.04872?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.04872</a><br><b>Authors</b>: Lin Sun et al. (Qiyuan Tech, Peking University)<br><b>Focus</b>: Compressing LLMs via distillation with enhanced accuracy </p><p class="paragraph" style="text-align:left;">TinyR1-32B-Preview uses Branch-Merge distillation to compress a 671B model (DeepSeek-R1) into a 32B student, improving accuracy in math, coding, and science while enabling efficient, quantization-ready deployment. It combines models using KL divergence for parameter selection.</p><p class="paragraph" style="text-align:left;"><b>Results</b>: Outperforms baselines by +5.5 (Math), +4.4 (Coding), +2.9 (Science), nearly matching the 671B model and surpassing a 70B variant.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Unlearning</b></h2><p class="paragraph" style="text-align:left;"><b>Group-robust Machine Unlearning</b></p><p class="paragraph" style="text-align:left;"><br><b>Paper:</b> <a class="link" href="https://arxiv.org/abs/2503.09330?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://arxiv.org/abs/2503.09330</a><br><b>Authors:</b> Thomas De Min et al. (University of Trento)<br><b>Focus: </b>Ensuring fairness in machine unlearning for LLMs<br><b>Code: </b><a class="link" href="https://github.com/tdemin16/group-robust_machine_unlearning?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llm-research-highlights-march-1-15-2025" target="_blank" rel="noopener noreferrer nofollow">https://github.com/tdemin16/group-robust_machine_unlearning</a></p><p class="paragraph" style="text-align:left;">Standard unlearning can harm model fairness when the forget set over represents certain groups, degrading performance unevenly. This paper proposes group-robust unlearning to maintain accuracy across groups in large language models (LLMs) handling diverse data.</p><p class="paragraph" style="text-align:left;"><b>Key Innovations:</b></p><ul><li><p class="paragraph" style="text-align:left;"><b>REWEIGHT:</b> Reweights sampling in retraining for fair exact unlearning.</p></li><li><p class="paragraph" style="text-align:left;"><b>MIU:</b> Approximate unlearning that reduces group-specific bias by minimizing mutual information, aligning with the original model.</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results:</b> On CelebA (0.5 unlearning ratio), MIU with REWEIGHT hits 69.0% group accuracy, beating SCRUB’s 62.9%, and keeps equalized odds delta at 0.6 vs. SCRUB’s 3.2. This method ensures equitable unlearning, enhancing LLM fairness for data removal tasks.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=87079946-dd75-4e66-98f8-2ce4803356db&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Breakthrough papers improving LLMs performance</title>
  <description>Pivotal advancements from February 16th–28th, 2025, redefining performance and efficiency in large language models</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f2953384-9fd4-4e9b-b276-89dcda1f1f05/core_feb_16_to_28.png" length="312024" type="image/png"/>
  <link>https://llm.beehiiv.com/p/breakthrough-papers-improving-llms-performance</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/breakthrough-papers-improving-llms-performance</guid>
  <pubDate>Thu, 06 Mar 2025 04:21:51 +0000</pubDate>
  <atom:published>2025-03-06T04:21:51Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Highlights</b></h2><ul><li><p class="paragraph" style="text-align:left;"><b>Enhanced Attention Mechanisms:</b> Native Sparse Attention and SpargeAttn propose hardware-aligned and universal sparse frameworks that accelerate long-context inference and training without compromising performance.</p></li><li><p class="paragraph" style="text-align:left;"><b>Knowledge Integration & Continual Learning:</b> &quot;How Do LLMs Acquire New Knowledge?&quot; and &quot;How Much Knowledge Can You Pack into a LoRA Adapter&quot; explore strategies to embed new information into LLMs while preserving established capabilities.</p></li><li><p class="paragraph" style="text-align:left;"><b>Temporal & Diffusion Innovations:</b> The Continuous Diffusion Model and Temporal Heads work address the dynamic nature of language by refining outputs iteratively and specializing in time-specific information recall.</p></li><li><p class="paragraph" style="text-align:left;"><b>Advanced Fine-Tuning Strategies:</b> Make LoRA Great Again, PAFT, and Thinking Preference Optimization introduce novel adaptation methods to improve efficiency and enhance chain-of-thought reasoning during fine-tuning.</p></li><li><p class="paragraph" style="text-align:left;"><b>Architectural Enhancements:</b> You Do Not Fully Utilize Transformer&#39;s Representation Capacity and MUDDFormer offer solutions to overcome residual bottlenecks by improving cross-layer information flow and expanding representational power.</p></li></ul><div class="recommendation" id="59dad3c7-38da-468a-9bba-f11ba0ba407a"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/f2953384-9fd4-4e9b-b276-89dcda1f1f05/core_feb_16_to_28.png?t=1741234086"/></figure><h3 class="recommendation__title"> Papers imprving performance of LLMs published between February 16th-28th 2025 </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzU5ZGFkM2M3LTM4ZGEtNDY4YS05YmJhLWYxMWJhMGJhNDA3YS9jb3JlJTIwRmViJTIwMTZ0aC0yOHRoJTI3MjAyNS53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDFaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9MmZkNjQwMWU5MmY2MTY4NWEwZTNhYTk1ZWY3NjU3NDg5ZTQ4YWFkZmIyOWJlZDc3OGNkYWJlNjZjMGRhYzljMVwiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlL2YyOTUzMzg0LTlmZDQtNGU5Yi1iMjc2LTg5ZGNkYTFmMWYwNS9jb3JlX2ZlYl8xNl90b18yOC5wbmc_dD0xNzQxMjM0MDg2XCIsXCJ0aXRsZVwiOlwiUGFwZXJzIGltcHJ2aW5nIHBlcmZvcm1hbmNlIG9mIExMTXMgcHVibGlzaGVkIGJldHdlZW4gRmVicnVhcnkgMTZ0aC0yOHRoIDIwMjVcIn0i." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Survey paper</b></h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.14776?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">SurveyX: Academic Survey Automation via Large Language Model</a></p></li><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.10708?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey</a></p></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h3 class="heading" style="text-align:left;" id="optimize-global-it-operations-with-">Optimize global IT operations with our World at Work Guide</h3><div class="image"><a class="image__link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_abcd3a7a-6705-404e-b746-a27ffb8db053_c39077bb&bhcl_id=6b90bf2d-8f7e-4cec-9ea7-9e83ae516941_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/72e1a843-4cc5-47d1-a733-66cd584a6b88/1200x600.png?t=1740413803"/></a></div><p class="paragraph" style="text-align:left;">Explore this <a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_abcd3a7a-6705-404e-b746-a27ffb8db053_c39077bb&bhcl_id=6b90bf2d-8f7e-4cec-9ea7-9e83ae516941_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">ready-to-go guide</a> to support your IT operations in 130+ countries. Discover how:</p><ul><li><p class="paragraph" style="text-align:left;">Standardizing global IT operations enhances efficiency and reduces overhead</p></li><li><p class="paragraph" style="text-align:left;">Ensuring compliance with local IT legislation to safeguard your operations</p></li><li><p class="paragraph" style="text-align:left;">Integrating Deel IT with EOR, global payroll, and contractor management optimizes your tech stack</p></li></ul><p class="paragraph" style="text-align:left;">Leverage <a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_abcd3a7a-6705-404e-b746-a27ffb8db053_c39077bb&bhcl_id=6b90bf2d-8f7e-4cec-9ea7-9e83ae516941_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Deel IT</a> to manage your global operations with ease.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_abcd3a7a-6705-404e-b746-a27ffb8db053_c39077bb&bhcl_id=6b90bf2d-8f7e-4cec-9ea7-9e83ae516941_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Download free guide</a></p><p class="paragraph" style="text-align:left;"></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">LLMs architecture improvement </h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.11089?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention</a> <br><br>Standard attention mechanisms are notoriously computationally expensive for long-context modeling. This paper fills the gap by proposing a sparse attention mechanism that not only reduces computation but is also natively trainable on modern hardware.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bdae9080-6714-479d-8645-34ce8c38a21f/image.png?t=1741232254"/></div><p class="paragraph" style="text-align:left;">The authors introduce NSA - a dynamic hierarchical sparse strategy. In simple terms, NSA compresses tokens coarsely while selectively preserving crucial tokens at a fine-grained level. This two-tiered approach balances global context awareness with local precision. The design is carefully optimized for modern hardware by balancing arithmetic intensity, enabling end-to-end training without the pre-training overhead typical of full attention models.</p><p class="paragraph" style="text-align:left;"><b>Results: </b>Experiments reveal that models pre-trained with NSA not only match or exceed full-attention baselines on general benchmarks and long-context tasks but also achieve substantial speedups on 64k-length sequences during decoding, forward, and backward passes.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.11196?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training</a> | <a class="link" href="https://github.com/zjunlp/DynamicKnowledgeCircuits?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">GitHub</a><br><br>While LLMs excel at knowledge-intensive tasks, the internal mechanisms that assimilate new knowledge remain poorly understood. This paper explores how neural circuits evolve to incorporate new facts during continual pre-training.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ba4fb4df-34fa-4f06-bee2-cdb49ad0915b/image.png?t=1741232430"/></div><p class="paragraph" style="text-align:left;">Using a “knowledge circuits” lens, the study analyzes the formation and optimization of computational subgraphs responsible for knowledge storage. The researchers show that new information is more effectively integrated when it relates to pre-existing knowledge, and they identify a distinct deep-to-shallow evolution pattern in these circuits.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.11564?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Continuous Diffusion Model for Language Modeling</a> | <a class="link" href="https://github.com/harryjo97/RDLM?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">GitHub</a><br><br>Traditional diffusion models for discrete data lose valuable iterative refinement signals. This study addresses the limitations of discrete diffusion in language modeling by proposing a continuous alternative.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b><br>The paper bridges discrete diffusion and continuous flow on a statistical manifold. By incorporating the geometry of the underlying categorical distribution, the authors design a diffusion process that generalizes previous discrete models. A simulation-free training framework based on radial symmetry further alleviates the high-dimensionality challenge.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.16894?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment</a><br><br>Although Low-Rank Adaptation (LoRA) enables efficient fine-tuning of LLMs, its performance traditionally lags behind full fine-tuning. This paper seeks to close that gap.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/845f7b53-ba53-4d7f-8fa9-1c39ec4eb86f/image.png?t=1741232619"/></div><p class="paragraph" style="text-align:left;">The authors propose GOAT, a framework that dynamically integrates SVD-based priors with a Mixture-of-Experts (MoE) architecture. By deriving a theoretical scaling factor, GOAT realigns optimization dynamics to better harness pre-trained knowledge without modifying existing architecture or training routines.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.18137?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference</a> | <a class="link" href="https://github.com/thu-ml/SpargeAttn?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">GitHub</a><br><br>Even though many models exhibit sparse attention maps, most optimizations have been model-specific. There is a need for a universal sparse attention method that accelerates inference across diverse architectures.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6ac7bd1a-4011-410c-bbc2-9ba28ddadd4e/image.png?t=1741232708"/></div><p class="paragraph" style="text-align:left;">SpargeAttn employs a two-stage online filtering mechanism. Initially, it predicts the attention map to decide which matrix multiplications can be skipped. Then, an online softmax-aware filter further eliminates unnecessary computations without extra overhead.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.14502?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?</a><br><br>Integrating new facts into LLMs via LoRA fine-tuning can sometimes lead to a degradation in general performance, particularly in question-answering benchmarks.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b><br>This study systematically varies the proportion of new versus known facts in the training data when fine-tuning a Llama-3.1-8B-instruct model. The experiments highlight the delicate balance required to incorporate additional knowledge without biasing the model’s outputs toward overrepresented entities.</p><p class="paragraph" style="text-align:left;"><b>Results: </b>The best performance is observed when the training data is a balanced mix of known and new facts. However, when biased, the model shows a decline in benchmark performance and a tendency to overcommit to few dominant answers—emphasizing the importance of careful data composition during LoRA-based updates.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.14258?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information</a></p><p class="paragraph" style="text-align:left;">While LLMs excel in fact recall, their ability to handle temporally dynamic information is less understood. This work investigates the specific mechanisms that store time-related knowledge.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/52221545-7af9-4c9a-beec-60a8032bf710/image.png?t=1741232792"/></div><p class="paragraph" style="text-align:left;">The researchers identify “Temporal Heads”—specialized attention heads that process time-specific information. By disabling these heads and observing performance drops in temporal recall (without affecting time-invariant knowledge), they demonstrate the critical role these heads play.</p><p class="paragraph" style="text-align:left;"><b>Results: </b>The findings confirm that temporal heads are activated by both numeric cues (e.g., “in 2004”) and textual phrases (e.g., “in the year…”). Their study lays the groundwork for targeted editing of temporal information within LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.09245?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">You Do Not Fully Utilize Transformer&#39;s Representation Capacity</a> | <a class="link" href="https://github.com/corl-team/lime?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;">Standard Transformers limit themselves by relying solely on the immediately preceding layer’s output, leading to suboptimal representation capacity.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b><br>The authors introduce Layer-Integrated Memory (LIMe), which aggregates hidden states from earlier layers without increasing the overall memory footprint. This strategy counters representation collapse by enriching the model’s contextual information.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.12859?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">PAFT: Prompt-Agnostic Fine-Tuning</a><br><br>Fine-tuned LLMs often overfit to specific prompt formulations, resulting in fragile performance when the prompts vary. PAFT addresses this by aiming for prompt-agnostic robustness.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e42f0f2f-8bb0-459a-b7e2-f3a8f31de93e/image.png?t=1741232926"/></div><p class="paragraph" style="text-align:left;">PAFT operates in two stages. First, it generates a diverse set of synthetic candidate prompts. Then, during fine-tuning, it randomly samples from this set so that the model learns to rely on underlying task principles rather than fixed prompt cues.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.12170?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections</a> | <a class="link" href="https://github.com/Caiyun-AI/MUDDFormer?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><p class="paragraph" style="text-align:left;">Residual connections in Transformers can act as bottlenecks, limiting cross-layer information flow and thus hindering overall performance.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ed882a3e-2ac6-4d2d-a401-ca39e9092236/image.png?t=1741233004"/></div><p class="paragraph" style="text-align:left;">MUDDFormer introduces Multiway Dynamic Dense (MUDD) connections, which generate connection weights dynamically based on the hidden states for each input stream (query, key, value, or residual). This approach effectively integrates information from earlier layers, breaking through the limitations of traditional residual architectures.</p><p class="paragraph" style="text-align:left;"><b>Results: </b>The paper reports striking improvements MUDDFormer achieves the performance of Transformers trained with 1.8×–2.4× compute. For example, the MUDDPythia-2.8B model matches the pretraining perplexity of Pythia-6.9B and rivals Pythia-12B in five-shot tasks—all while adding only 0.23% more parameters and 0.4% extra computation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.13173?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=breakthrough-papers-improving-llms-performance" target="_blank" rel="noopener noreferrer nofollow">Thinking Preference Optimization</a><br><br>Supervised Fine-Tuning (SFT) for chain-of-thought reasoning often leads to performance plateaus or declines with repeated training, and acquiring high-quality long responses is costly.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b><br>Thinking Preference Optimization (ThinkPO) is a post-SFT method that leverages readily available short chain-of-thought responses as “rejected” answers and long responses as “chosen” ones. By applying direct preference optimization, the model is nudged to generate more comprehensive reasoning without the need for new, expensive data collection.</p><p class="paragraph" style="text-align:left;"><b>Results: </b>ThinkPO boosts performance significantly—math reasoning accuracy increases by 8.6% and output length by 25.9%. Notably, when applied to the DeepSeek-R1-Distill-Qwen-7B model, performance on the MATH500 benchmark improves from 87.4% to 91.2%.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=d911211d-c25e-4ce6-8b6c-16dd3ef4ee34&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Creative way of using LLMs</title>
  <description>Exploring interesting applications of LLMs from research papers published between February 15th-28th, 2025</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6f76ef92-ee5f-447f-9bea-f91e29e91cac/feb15-28-application.png" length="374126" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-application-february-15-28-2025</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-application-february-15-28-2025</guid>
  <pubDate>Tue, 04 Mar 2025 15:15:00 +0000</pubDate>
  <atom:published>2025-03-04T15:15:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Highlights</h2><ul><li><p class="paragraph" style="text-align:left;"><b>LEDD</b>: <b>Find your data lake treasures</b> with LLM-powered semantic search!</p></li><li><p class="paragraph" style="text-align:left;"><b>LLMs in Mobile Apps</b>: Discover the <b>secrets & struggles of adding LLMs to your Android apps</b>!</p></li><li><p class="paragraph" style="text-align:left;"><b>BayesGenie</b>: <b>Edit like a pro</b> with AI that combines LLMs & Bayesian optimization!</p></li><li><p class="paragraph" style="text-align:left;"><b>Talking Like P&IDs</b>: <b>Chat with your P&IDs</b> using natural language and AI-powered knowledge graphs!</p></li><li><p class="paragraph" style="text-align:left;"><b>ArtInsight</b>: <b>Bridge vision gaps</b> with AI that describes children&#39;s art for BLV families!</p></li><li><p class="paragraph" style="text-align:left;"><b>Learning Code-Edit Embedding</b>: <b>Level up your debugging</b> with AI that learns from your coding edits!</p></li></ul><div class="recommendation" id="738c529e-6148-4858-82c5-4357e7c91ffd"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/6effa6bf-bf13-41fa-935d-ace7d7ef9ada/feb15-28-application.png?t=1741091983"/></figure><h3 class="recommendation__title"> (Feb 15-28 2025) Innovative applications using LLMs </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzczOGM1MjllLTYxNDgtNDg1OC04MmM1LTQzNTdlN2M5MWZmZC9BcHBsaWNhdGlvbl8lMjBGZWIlMjAxNXRoLTI4dGglMkMlMjAyMDI1Lndhdj9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NlxcdTAwMjZYLUFtei1DcmVkZW50aWFsPUFLSUFRQ01IVFFTRTJKR0FHWEhKJTJGMjAyNjA5MTMlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdFxcdTAwMjZYLUFtei1EYXRlPTIwMjYwOTEzVDExMzM0MlpcXHUwMDI2WC1BbXotRXhwaXJlcz02MDQ4MDBcXHUwMDI2WC1BbXotU2lnbmVkSGVhZGVycz1ob3N0XFx1MDAyNlgtQW16LVNpZ25hdHVyZT0zYmJjNWIxZTQ3ZGJkZGY0ZjFiODhlNjhmODIxY2UxMzE2MjZhNzMzMjYwNjhjMTQyOTRmZjg3YTIwYjcxZjA2XCIsXCJ0eXBlXCI6XCJhdWRpby94LXdhdlwiLFwidGh1bWJuYWlsVXJsXCI6XCJodHRwczovL2JlZWhpaXYtaW1hZ2VzLXByb2R1Y3Rpb24uczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Fzc2V0L2ZpbGUvNmVmZmE2YmYtYmYxMy00MWZhLTkzNWQtYWNlN2Q3ZWY5YWRhL2ZlYjE1LTI4LWFwcGxpY2F0aW9uLnBuZz90PTE3NDEwOTE5ODNcIixcInRpdGxlXCI6XCIoRmViIDE1LTI4IDIwMjUpIElubm92YXRpdmUgYXBwbGljYXRpb25zIHVzaW5nIExMTXNcIn0i." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Creative ways to use LLMs!!</b></h2><h4 class="heading" style="text-align:left;">The AI That Called 📞 2,739 Strangers (And Got Them Talking)</h4><p class="paragraph" style="text-align:left;">Paper: <a class="link" href="http://arxiv.org/abs/2502.20140v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b2b05101-e206-475a-8810-4a2195f4dc61/image.png?t=1741091533"/></div><p class="paragraph" style="text-align:left;">Traditional telephone surveys require substantial human resources for recruiting and training interviewers. This research introduces an autonomous AI system that dramatically scales survey deployment without compromising data quality.</p><p class="paragraph" style="text-align:left;"><span style="color:#000000;font-size:medium;">In a pilot study with </span><span style="color:#000000;font-size:medium;"><b>75 calls</b></span><span style="color:#000000;font-size:medium;"> in the </span><span style="color:#000000;font-size:medium;"><b>United States</b></span><span style="color:#000000;font-size:medium;"> and a larger study with </span><span style="color:#000000;font-size:medium;"><b>2,739 calls</b></span><span style="color:#000000;font-size:medium;"> in </span><span style="color:#000000;font-size:medium;"><b>Peru</b></span><span style="color:#000000;font-size:medium;">, the system was </span>deployed an STT-LLM-TTS pipeline that achieved <b>89% response parity</b> with human interviews across Peru. But here&#39;s the twist: When respondents hesitated, the AI used &quot;cognitive mirroring&quot; - subtly matching regional speech patterns, reducing hang-ups by <b>22%</b> compared to rigid IVR systems.</p><p class="paragraph" style="text-align:left;"> <b>At $0.83 per completed survey vs. $25+ for human teams</b>, this could democratize global policy research. The AI even detected vocal stress patterns correlating with response accuracy (r=0.67, p&lt;0.01)</p><hr class="content_break"><h4 class="heading" style="text-align:left;">Stanford&#39;s code-editing model analyzed <b>1.4M student </b>submissions</h4><p class="paragraph" style="text-align:left;">Paper: <a class="link" href="http://arxiv.org/abs/2502.19407v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Learning Code-Edit Embedding to Model Student Debugging Behavior</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8be06709-5603-4db3-ae5d-68aa238487bd/image.png?t=1741091579"/></div><p class="paragraph" style="text-align:left;">Providing timely, personalized feedback in programming education is challenging. By modeling how students debug their code, educators can design targeted support tools that improve learning outcomes.</p><p class="paragraph" style="text-align:left;">This paper presents an encoder-decoder model that learns &quot;code-edit embeddings&quot; from consecutive student code submissions. By analyzing the patterns in how students edit their code after receiving test case feedback, the system can generate personalized next-step coding suggestions, identify common debugging patterns through clustering, and better understand the learning process. Educators can provide more targeted assistance based on individual problem-solving approaches.</p><hr class="content_break"><h3 class="heading" style="text-align:left;" id="optimize-global-it-operations-with-">Optimize global IT operations with our World at Work Guide</h3><div class="image"><a class="image__link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_77b03252-f598-4e0f-a2ec-9f6f81dcb916_c39077bb&bhcl_id=151cf66c-b68b-4a1c-a024-326955ef4120_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/72e1a843-4cc5-47d1-a733-66cd584a6b88/1200x600.png?t=1740413803"/></a></div><p class="paragraph" style="text-align:left;">Explore this <a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_77b03252-f598-4e0f-a2ec-9f6f81dcb916_c39077bb&bhcl_id=151cf66c-b68b-4a1c-a024-326955ef4120_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">ready-to-go guide</a> to support your IT operations in 130+ countries. Discover how:</p><ul><li><p class="paragraph" style="text-align:left;">Standardizing global IT operations enhances efficiency and reduces overhead</p></li><li><p class="paragraph" style="text-align:left;">Ensuring compliance with local IT legislation to safeguard your operations</p></li><li><p class="paragraph" style="text-align:left;">Integrating Deel IT with EOR, global payroll, and contractor management optimizes your tech stack</p></li></ul><p class="paragraph" style="text-align:left;">Leverage <a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_77b03252-f598-4e0f-a2ec-9f6f81dcb916_c39077bb&bhcl_id=151cf66c-b68b-4a1c-a024-326955ef4120_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Deel IT</a> to manage your global operations with ease.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.deel.com/resources/deel-guide-to-the-world-of-work-in-2024-with-deel-it/?utm_medium=sponsored-newsletter&utm_source=beehiiv&utm_term={{publication_alphanumeric_id}}&utm_campaign=ww_engage_download_beehiiv_sponnewsletter_it-theworldatwork-feb25_it_all&utm_content=engage_it_sponnewsletter_theworldatwork-sponnews400-it_en&_bhiiv=opp_77b03252-f598-4e0f-a2ec-9f6f81dcb916_c39077bb&bhcl_id=151cf66c-b68b-4a1c-a024-326955ef4120_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Download free guide</a></p><p class="paragraph" style="text-align:left;"></p><hr class="content_break"><h4 class="heading" style="text-align:left;">ArtInsight: Making Children&#39;s Art Accessible to All</h4><p class="paragraph" style="text-align:left;">Paper:<b> </b><a class="link" href="http://arxiv.org/abs/2502.19263v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">ArtInsight: Enabling AI-Powered Artwork Engagement for Mixed Visual-Ability Families</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9c8be90c-f823-4eae-b2dd-0ec4c63f11d6/image.png?t=1741091652"/></div><p class="paragraph" style="text-align:left;">Art is meant to be shared and appreciated, but how do blind or low-vision family members engage with children&#39;s artwork? ArtInsight offers a touching solution to this challenge using LLMs.</p><p class="paragraph" style="text-align:left;"><b>The Innovation</b>: The ArtInsight system uses large language models to generate rich, descriptive content about children&#39;s artwork. It goes beyond basic visual descriptions to create detailed artistic interpretations, audio recordings of children explaining their art, and AI-generated questions to prompt meaningful discussions</p><p class="paragraph" style="text-align:left;">This paper shows how LLMs can bridge experiential gaps and create more inclusive family interactions around visual creativity.</p><hr class="content_break"><h4 class="heading" style="text-align:left;">Engineering Diagrams That Speak Your Language</h4><p class="paragraph" style="text-align:left;">Paper: <a class="link" href="http://arxiv.org/abs/2502.18928v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Talking like Piping and Instrumentation Diagrams (P&IDs)</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4952b29e-4d4a-4f0f-bf76-32c1be9c1413/image.png?t=1741091687"/></div><p class="paragraph" style="text-align:left;">Process engineers are familiar with the complex schematics known as Piping and Instrumentation Diagrams (P&IDs). Now, researchers have enabled natural language communication with these diagrams through a three-component system: transforming P&IDs into graph representations via the DEXPI data model, creating P&ID knowledge graphs, and integrating these with LLMs through graph-based retrieval augmented generation.</p><p class="paragraph" style="text-align:left;">This innovation allows engineers to query complex diagrams using plain language, potentially reducing errors and improving efficiency in industrial settings where these diagrams are essential.</p><hr class="content_break"><h4 class="heading" style="text-align:left;">BayesGenie: Precision Image Editing Through Language</h4><p class="paragraph" style="text-align:left;">Paper: <a class="link" href="http://arxiv.org/abs/2502.18116v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Bayesian Optimization for Controlled Image Editing via LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1ba3ae74-0010-4520-b096-e233327b007a/image.png?t=1741091727"/></div><p class="paragraph" style="text-align:left;">Image editing requires precise control and semantic accuracy—critical for both creative professionals and everyday users. This research introduces a novel method that marries LLMs with Bayesian optimization to refine image edits using natural language instructions.</p><p class="paragraph" style="text-align:left;">The model, dubbed BayesGenie, uses LLM-generated edits and applies Bayesian optimization to fine-tune parameters, thereby ensuring that the edited image retains both accuracy and semantic consistency. Users can specify desired changes in natural language without manually selecting image areas.</p><hr class="content_break"><h4 class="heading" style="text-align:left;">Discovering Hidden Data in Vast Data Lakes</h4><p class="paragraph" style="text-align:left;">Paper:<b> </b><a class="link" href="https://arxiv.org/abs/2502.15182?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow">LEDD: Large Language Model-Empowered Data Discovery in Data Lakes</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cecfb620-b118-4f6c-87c4-94eb5310242d/image.png?t=1741091848"/></div><p class="paragraph" style="text-align:left;">Many organizations store vast amounts of unstructured data in data lakes, and finding useful information within them can be overwhelming. </p><p class="paragraph" style="text-align:left;">The LEDD system uses an LLM to automatically generate hierarchical catalogs that organize data semantically. It also supports natural language queries, which make it easier to search for specific information across millions of data points. Though exact numerical improvements in search speed or accuracy are not detailed in the abstract, the paper highlights that LEDD can significantly reduce the time and effort needed to locate valuable data.</p><hr class="content_break"><h4 class="heading" style="text-align:left;">A detailed survey paper on LLMs in mobile apps</h4><p class="paragraph" style="text-align:left;">Paper: <a class="link" href="http://arxiv.org/abs/2502.15908v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=creative-way-of-using-llms" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LLMs in Mobile Apps: Practices, Challenges, and Opportunities</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is significant because it explores how LLMs can transform mobile application development, making apps more intelligent and personalized while addressing the unique challenges developers face.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research methodology involved constructing a dataset of 149 LLM-enabled Android apps and conducting an exploratory analysis to examine the deployment and usage of LLMs within these applications, focusing on integration strategies and challenges.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The analysis revealed key characteristics of the dataset, prevalent integration strategies, and common challenges faced by developers in the context of mobile app development.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=e7c0aa06-a41b-4154-b7bd-9cc93fd96d58&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Research papers improving performance of LLMs [3/3]</title>
  <description>Research papers published from January 16th to February 15th, 2025 proposing context length and architectural changes in LLMs </description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2ed86333-be53-40ac-a924-1e079a5e1b06/3rd_edition_collage.png" length="320600" type="image/png"/>
  <link>https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-3-3</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-3-3</guid>
  <pubDate>Sat, 22 Feb 2025 16:47:00 +0000</pubDate>
  <atom:published>2025-02-22T16:47:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:center;"><b>In partnership with</b></p><div class="image"><a class="image__link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42cc7ce3-bec6-4a63-94bc-a31e2bdaec9b/Prolific_Logo_Blue__1_.png?t=1740074994"/></a></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">What’s in it today?</h2><ul><li><p class="paragraph" style="text-align:left;">SCONE: who needs a bigger vocabulary when you can just <b>contextualize</b> the heck out of your n-grams?</p></li><li><p class="paragraph" style="text-align:left;">DAAs: Making LLMs <b>agree with humans</b>, one preference at a time (and sometimes only needing 5% of the data to do it!)</p></li><li><p class="paragraph" style="text-align:left;">CT-KL: Ignoring the KL penalty and focusing on <b>critical tokens</b> to boost LLMs, because sometimes you just need to be a rebel</p></li><li><p class="paragraph" style="text-align:left;">LIMO: <b>Less is More</b> for Reasoning - Proving you don&#39;t need a mountain of data, just a tiny, well-curated molehill!</p></li></ul><p class="paragraph" style="text-align:left;">Don’t have much time to read entire newsletter? No problem? Listen to this fun and engaging podcast covering these research papers in detail.</p><div class="recommendation" id="345ebbdd-90cb-4fa9-af23-1e5164763534"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/7877ce4a-e10a-48fa-8022-a4c4ff638f4d/image.png?t=1740241839"/></figure><h3 class="recommendation__title"> Papers improving performance of LLMs published in 2025 </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzM0NWViYmRkLTkwY2ItNGZhOS1hZjIzLTFlNTE2NDc2MzUzNC8lNUIzJTVEJTIwSmFuJTIwMTZ0aC1GZWIlMjAxNXRoJTIwMjAyNS53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDNaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9YTkzZjRmZGU4MWMyNGEzMjRjYmRkMzY5NWUwYmRiMzNhY2ZiYWZmNWExZjMwYjk1NzQ5MjlkZTMwNWQ2OTJmM1wiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlLzc4NzdjZTRhLWUxMGEtNDhmYS04MDIyLWE0YzRmZjYzOGY0ZC9pbWFnZS5wbmc_dD0xNzQwMjQxODM5XCIsXCJ0aXRsZVwiOlwiUGFwZXJzIGltcHJvdmluZyBwZXJmb3JtYW5jZSBvZiBMTE1zIHB1Ymxpc2hlZCBpbiAyMDI1XCJ9Ig." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">Scaling Embedding Layers in Language Models</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.01637?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> </p><p class="paragraph" style="text-align:left;">Embedding layers in language models maps discrete tokens to continuous vector representations. Scaling embedding layers enhances model performance but simply increasing vocabulary size has limitations due to the tight coupling between input and output embedding layers, especially the increasing computational cost of the output layer&#39;s logits calculation. The benefits of simply scaling the vocabulary also diminish due to the proliferation of low-frequency &quot;tail tokens&quot; which are rarely updated during training, resulting in lower quality embeddings. </p><p class="paragraph" style="text-align:left;">To address these issues, the paper introduces SCONE (Scalable, Contextualized, Offloaded, N-gram Embedding), a novel approach designed to disentangle the input and output embeddings, enabling effective input embedding scaling with minimal additional inference cost.</p><p class="paragraph" style="text-align:left;"><b>Proposed approach</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8547e528-c13e-4650-b9d1-7676cbcc18a0/image.png?t=1740235721"/></div><p class="paragraph" style="text-align:left;">The core of SCONE lies in augmenting the existing token vocabulary with contextualized variants derived from frequent n-grams (f-grams). Instead of directly increasing the vocabulary size, SCONE uses a set of frequently occurring n-grams to provide contextualized representations for each input token. These contextualized tokens are used only for input embedding computation, allowing for a massive augmented input embedding table without increasing the output layer&#39;s computational burden. The embeddings for these contextualized tokens are generated by a separate embedding transformer model, referred to as the f-gram model, which is jointly trained with the main language model. This allows for rich contextualized representations without the sparsity issues associated with simply increasing the vocabulary.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/47276ee2-9916-45a8-8bc4-68e78d45bda1/image.png?t=1740237171"/></div><p class="paragraph" style="text-align:left;">The paper proposes an efficient method, analogous to continuing a Byte Pair Encoding (BPE) tokenizer&#39;s training, to identify frequent n-grams (f-grams) in the training corpus. This process involves scanning the corpus to count the occurrences of n-grams of length up to a specified maximum (n). To reduce memory usage, a minimum frequency threshold is applied during the counting process. The SCONE method maps a sequence of tokens to a sequence of embedding vectors. During training, the method utilizes an f-gram transformer model (Af-gram) to generate embeddings for contextualized tokens. At inference time, a precomputed f-gram embedding layer (F) is used to map the f-grams to embedding vectors, allowing for efficient retrieval of contextualized embeddings. After that, The embeddings generated by the SCONE method are passed to a standard transformer model (Amain), referred to as the main model, followed by a prediction head (D). This combination enables next-word prediction with SCONE. The f-gram model and the main model are trained jointly, allowing the f-gram model to learn contextualized representations that are aligned with the main model&#39;s objective.</p><p class="paragraph" style="text-align:left;"><b>Results</b></p><p class="paragraph" style="text-align:left;">The paper shows that scaling both the number of cached f-gram embeddings and the size of the f-gram model allows SCONE to outperform a 1.9B parameter baseline model across diverse corpora. For example, with 10M f-grams and a 1.8B f-gram model, a 1.3B main model matches the perplexity of the 1.9B baseline. With 1B f-grams and a 1.8B f-gram model, a 1B main model surpasses the 1.9B baseline. Also, SCONE enables improved performance while using only half the inference-time FLOPS of the baseline model. </p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h3 class="heading" style="text-align:left;">Sponsored: Special offer for you claim $50 free credits!</h3><div class="image"><a class="image__link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/842ecda9-55ef-4274-978a-0eba2a662568/Neuron_Newsletter_image_2025.png?t=1740066909"/></a></div><p class="paragraph" style="text-align:left;">Accelerate your AI projects with Prolific. Claim <b>$50 free credits</b> and get quality human data in minutes <b>from 200,000+ taskers</b>. No setup cost, no subscription, no delay—get started, top up your account to claim your free credit, and test Prolific for yourself now. <b>Use code:</b> LLM-RESEARCH-50 <br></p><p class="paragraph" style="text-align:center;"><b><a class="link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Claim free credit</a></b><b><a class="link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow"> now</a></b></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">The Differences Between Direct Alignment Algorithms are a Blur</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.01237?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> </p><p class="paragraph" style="text-align:left;">Aligning Large Language Models (LLMs) with human values is a challenging task. Generally methods like Supervised Fine-Tuning (SFT), Reward Modeling (RM), and Reinforcement Learning from Human Feedback (RLHF) is used for that. Recenlty Direct Alignment Algorithms (DAAs)  is gained a traction. It directly optimize the policy based on human preferences, without the explicit reward modeling step. However, the paper notes the confusing landscape of DAAs, differing in their theoretical underpinnings (pairwise vs. pointwise objectives), implementation details (e.g., using a reference policy or an odds ratio), and whether they require a separate SFT phase (one-stage vs. two-stage). The paper aims to clarify the relationships between these algorithms, determine the importance of the SFT stage, and identify the key factors influencing their alignment quality.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><p class="paragraph" style="text-align:left;">The core of the paper&#39;s methodology is a systematic empirical comparison of different Direct Alignment Algorithms. The paper investigates the role of the SFT stage, the influence of a tempering factor (β), and the impact of pairwise vs. pointwise preference optimization. To achieve this, the paper focuses primarily on two single-stage DAA methods, ORPO (Odds Ratio Preference Optimization) and ASFT (Aligned Supervised Fine-Tuning), and modifies them in several ways.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Introducing an Explicit SFT Stage</b></span><b>:</b> The paper explores whether adding a separate SFT phase <i>before</i> applying ORPO and ASFT improves their performance. This involves training a model using standard supervised fine-tuning on a dataset of instructions and responses, and then using ORPO or ASFT to align the model with human preferences.</p></li><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Adding a Tempering Factor (β)</b></span><b>:</b> The paper introduces a scaling parameter β into the ORPO and ASFT loss functions. This parameter, inspired by similar parameters used in other DAAs like DPO, controls the strength of the preference optimization. By varying the value of β, the paper investigates its impact on the alignment quality of ORPO and ASFT.</p></li><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Theoretical Analysis and Equivalence Proofs</b></span><b>:</b> The paper presents theoretical analyses demonstrating relationships between different DAA loss functions. These analyses include theorems proving the equivalence of ASFT to a binary cross-entropy loss, establishing an upper bound of ASFT on ORPO, and demonstrating collinearity of gradients under certain conditions. These theoretical results are used to support and interpret the empirical findings.</p></li><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Controlled Experiments</b></span><b>:</b> The paper conducts a series of carefully controlled experiments to evaluate the performance of different DAA configurations. These experiments involve training and evaluating models on benchmark datasets, using metrics designed to assess alignment quality.</p></li><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Varying SFT Data Size</b></span><b>:</b> The paper examines how the final alignment quality depends on the amount of data used in the SFT stage. This is done to determine whether it&#39;s necessary to use the full dataset for SFT, or if a smaller subset is sufficient.</p></li></ol><p class="paragraph" style="text-align:left;"><b>Result</b></p><p class="paragraph" style="text-align:left;">The paper shows that incorporating an explicit SFT stage significantly improves the performance of both ORPO and ASFT. The paper demonstrates that introducing a tempering factor (β) enhances the alignment quality of ASFT and ORPO. The authors found that tuning β was essential for achieving good performance.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.06533?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/jvasso/llm-rl-arithmetic?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Implementation</a> </p><p class="paragraph" style="text-align:left;">The paper presents a novel approach to improving the exploration capabilities of LLMs during reinforcement learning (RL) fine-tuning. The core challenge in RL fine-tuning is balancing exploration and stability. Traditional methods often use a Kullback-Leibler (KL) penalty to prevent the model from deviating too far from its pre-trained behavior, ensuring stability but potentially limiting exploration.</p><p class="paragraph" style="text-align:left;">The paper introduces the <b>Critical Token KL (CT-KL)</b> method, which dynamically adjusts the KL penalty during training based on critical tokens identified as crucial decision points. These critical tokens are determined using attention patterns and gradient information, indicating where targeted exploration can significantly impact performance. By reducing or eliminating the KL penalty for these tokens, CT-KL encourages more extensive exploration at these key junctures while maintaining stability elsewhere.</p><p class="paragraph" style="text-align:left;">This selective approach contrasts with traditional methods that apply uniform penalties across all tokens or eliminate penalties entirely, which can lead to instability or inefficient learning. The CT-KL method modifies the standard transformer architecture by adjusting its training objective without altering its base structure.</p><p class="paragraph" style="text-align:left;"><b>Implementation Details</b></p><p class="paragraph" style="text-align:left;">This method integrates CT-KL into existing RL frameworks for LLMs. The first step is identifying critical tokens within input sequences using attention weights and gradients. After that, the training objective is modified to selectively reduce or remove KL penalties for these identified critical tokens. Models are than trained with this modified objective on various tasks requiring complex reasoning and problem-solving skills.</p><p class="paragraph" style="text-align:left;"><b>Results</b></p><p class="paragraph" style="text-align:left;">The paper demonstrates significant improvements in performance when using CT-KL compared to standard approaches: On mathematical reasoning tasks, models trained with CT-KL showed an improvement of up to 8.2% compared to those using traditional uniform penalties.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">LIMO: Less is More for Reasoning</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.03387?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/GAIR-NLP/LIMO?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Implementation</a> | <a class="link" href="https://huggingface.co/datasets/GAIR/LIMO?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Dataset</a> | <a class="link" href="https://huggingface.co/GAIR/LIMO?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-3-3" target="_blank" rel="noopener noreferrer nofollow">Model</a></p><p class="paragraph" style="text-align:left;">This paper challenges the conventional wisdom that complex reasoning in LLMs requires extensive training data. This hypothesis tests that if a model has sufficient pre-trained knowledge and is presented with effective &quot;cognitive templates&quot; demonstrating problem-solving processes, it can achieve strong reasoning capabilities with minimal fine-tuning data.</p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><p class="paragraph" style="text-align:left;">Paper carefully curated a small training dataset to act as &quot;cognitive templates&quot; for reasoning. The key is not the <i>quantity</i> of the data, but its <i>quality</i> and how well it demonstrates the reasoning process. While the specific details of how these training samples are orchestrated is limited from the paper, they must have done so effectively to get such positive outcomes.</p><p class="paragraph" style="text-align:left;">The general idea for the methodology is:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Leveraging Pre-trained Knowledge:</b> The paper builds upon modern foundation models that have used significant mathematical knowledge during pre-training. This pre-trained knowledge serves as the foundation for reasoning.</p></li><li><p class="paragraph" style="text-align:left;"><b>Minimal Exemplars as Cognitive Templates:</b> The core approach involves creating a small set of training examples that demonstrate how to utilize the pre-trained knowledge to solve complex reasoning problems.</p></li><li><p class="paragraph" style="text-align:left;"><b>Emphasis on Inference-Time Computation:</b> The approach is designed to encourage extended deliberation and systematic application of pre-trained knowledge during inference. This highlights the importance of having sufficient computational resources at inference time.</p></li></ol><p class="paragraph" style="text-align:left;"><b>Results</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7877ce4a-e10a-48fa-8022-a4c4ff638f4d/image.png?t=1740241839"/></div><p class="paragraph" style="text-align:left;">The paper reports impressive results on mathematical reasoning benchmarks:</p><ul><li><p class="paragraph" style="text-align:left;"><b>AIME Benchmark:</b> LIMO achieves 57.1% accuracy on the AIME benchmark, a significant improvement over previous SFT-based models which achieved only 6.5% accuracy. This demonstrates the ability to achieve strong performance on a challenging reasoning task.</p></li><li><p class="paragraph" style="text-align:left;"><b>MATH Benchmark:</b> LIMO achieves 94.8% accuracy on the MATH benchmark, also a substantial improvement over previous models (59.2%).</p></li><li><p class="paragraph" style="text-align:left;"><b>Out-of-Distribution Generalization:</b> Most remarkably, LIMO demonstrates exceptional out-of-distribution generalization, achieving a 40.5% absolute improvement across 10 diverse benchmarks, outperforming models trained on 100x more data. This result challenges the common belief that SFT inherently leads to memorization rather than generalization.</p></li></ul></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=578fa359-e959-4f30-bae6-cbd9bb11025d&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Research papers improving performance of LLMs [2/3]</title>
  <description>Research papers published from January 16th to February 15th, 2025 proposing context length and architectural changes in LLMs </description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/04053706-0356-44c3-ae9c-ee5d6dbad7b2/part2.png" length="283592" type="image/png"/>
  <link>https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-2-3</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-2-3</guid>
  <pubDate>Fri, 21 Feb 2025 21:00:14 +0000</pubDate>
  <atom:published>2025-02-21T21:00:14Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:center;"><b>In partnership with</b></p><div class="image"><a class="image__link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42cc7ce3-bec6-4a63-94bc-a31e2bdaec9b/Prolific_Logo_Blue__1_.png?t=1740074994"/></a></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">What’s in it today?</h2><ul><li><p class="paragraph" style="text-align:left;"><b>ReLearn </b>makes LLMs forget unwanted knowledge and remember how to speak good</p></li><li><p class="paragraph" style="text-align:left;"><b>Coupled Adam</b> fixes Adam so language model embeddings aren&#39;t too &quot;extra&quot;</p></li><li><p class="paragraph" style="text-align:left;"><b>TransMLA </b>converts GQA models to MLA ones for better LLM expression, because apparently size does matter</p></li><li><p class="paragraph" style="text-align:left;"><b>LASP-2 </b>makes linear attention training zoom by decluttering communication</p></li></ul><p class="paragraph" style="text-align:left;">Don’t have much time to read entire newsletter? Well, listen to this fun and engaging podcast covering these research papers in detail.</p><div class="recommendation" id="f2638165-6f97-4d50-97ac-95911b3561a4"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/04053706-0356-44c3-ae9c-ee5d6dbad7b2/part2.png?t=1740171526"/></figure><h3 class="recommendation__title"> [2] Jan 16th-Feb 15th 2025 </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2L2YyNjM4MTY1LTZmOTctNGQ1MC05N2FjLTk1OTExYjM1NjFhNC8lNUIyJTVEJTIwSmFuJTIwMTZ0aC1GZWIlMjAxNXRoJTIwMjAyNS53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDRaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9ZGMxOTUyYzczNzJjOTZkMmJiNTljZjZkMzk3MGQ4M2UyMzViMjA4MWU1M2MyMzEyOGVjYjg3OGQxNmFjZWM4MlwiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlLzA0MDUzNzA2LTAzNTYtNDRjMy1hZTljLWVlNWQ2ZGJhZDdiMi9wYXJ0Mi5wbmc_dD0xNzQwMTcxNTI2XCIsXCJ0aXRsZVwiOlwiWzJdIEphbiAxNnRoLUZlYiAxNXRoIDIwMjVcIn0i." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">ReLearn: Unlearning via Learning for Large Language Models</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.11190?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/zjunlp/unlearn?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Paper implementation</a> </p><p class="paragraph" style="text-align:left;"><b>Why this research is important? </b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ff555564-ee45-4d86-945b-b17b7bda9d3e/image.png?t=1739983256"/></div><p class="paragraph" style="text-align:justify;">LLMs are getting trained on more and more data and its consumption is increasing drastically. LLM providers use large-scale AI training datasets that often contain unauthorized private and copyrighted information. Recent legal actions, like the New York Times lawsuit against OpenAI, underscore the urgency of addressing these issues. The core problem is that current unlearning methods often rely on reverse optimization, which reduces the probabilities of target tokens but degrades the model&#39;s fundamental language generation capabilities, resulting in repetitive or incoherent outputs. The paper argues that current evaluation metrics are also inadequate, as they primarily focus on contextual forgetting while failing to capture broader limitations in fluency and relevance. Therefore, this research is crucial for developing more effective and reliable unlearning techniques that comply with privacy and copyright regulations without compromising the model&#39;s linguistic coherence.</p><p class="paragraph" style="text-align:left;"><b>Approach:</b></p><p class="paragraph" style="text-align:justify;">Paper proposes unlearning pipeline based on data augmentation and positive optimization. Instead of suppressing token probabilities (reverse optimization), ReLearn overwrites sensitive information with new, authorized knowledge by training the model on augmented data. This process preserves the model&#39;s linguistic ability while forgetting target knowledge, akin to human memory updating. To better understand I divide ReLearn pipeline in following steps:</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42bb003f-6c73-47c9-80ec-f7f44050d300/image.png?t=1739983135"/></div><ol start="1"><li><p class="paragraph" style="text-align:justify;">Unlearning data synthesis: In this step ReLearn pipeline synthesize non-sensitive training data by augmenting the forget set (data to be unlearned) with diverse variations, ensuring comprehensive coverage of the knowledge to be forgotten. This process is entirely performed by an LLM using specific prompts. It performs question and answer augmentation. For question augment it iterates on each question-answer pair in the forget set, the method synthesizes four types of question variations: (1) Simple Variants to prevent overfitting to specific phrasings, (2) Contextual Variants to ensure forgetting across contexts, (3) Noise Variants to enhance robustness to noisy inputs, and (4) Logical Variants to adapt to different knowledge forms by altering the logic of the questions. and, for answer augmentation it iterates on each augmented question, the method synthesizes new pairs with relevant, deliberately vague answers. These answers must be (1) Unlearned, containing no original sensitive content; (2) Relevant, aligning with the question context; and (3) No-risk, avoiding the introduction of new sensitive content.</p></li><li><p class="paragraph" style="text-align:justify;">Content verification: To ensure the safety of the augmented data, the method uses a Content Verification process for the synthesized answers. This process utilizes LLMs to conduct Chain-of-Thought (COT) analysis on each augmented answer, evaluating it against predefined safety criteria. If verification fails, indicating a potential risk in the augmented data, the process returns to the step of &quot;Answer Augmentation.&quot;</p></li><li><p class="paragraph" style="text-align:justify;">Data diversification: To prevent QA format overfitting and catastrophic forgetting, the method uses two main strategies sentence completion and generic dataset. In sentence completion, the augmented data is augmented with sentence completion pairs, split from each answer. For example, &quot;Isabella Marquez can be reached through conventional electronic communication channels.&quot; is split into &quot;Isabella Marquez can be reached through&quot; and the label &quot;conventional electronic communication channels.&quot; and in generic dataset, ReLearn applies generic data by randomly sampling questions from WikiQA and Chatbot Instruction datasets.</p></li><li><p class="paragraph" style="text-align:justify;">Unlearning via learning: The unlearning objective is formulated using three datasets: the augmented forget set, the sentence completion set, and the generic dataset. The vanilla model is then fine-tuned on a combination of these datasets. The loss function combines generative loss, masked language model loss, and knowledge loss, which is then used to perform parameter update and fine-tuning of the vanilla model.</p></li></ol><p class="paragraph" style="text-align:left;"><b>Results:</b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0ec17d62-ddb6-4352-9df3-848230d13211/image.png?t=1739983176"/></div><p class="paragraph" style="text-align:justify;">Paper identified limitations in existing unlearning metrics and proposes three metrics: Knowledge Forgetting Rate (KFR), Knowledge Retention Rate (KRR), and Linguistic Score (LS). KFR measures the extent of knowledge forgetting, and KRR measures the extent of knowledge retention, while LS evaluates the linguistic quality of the unlearned model, capturing linguistic degradation patterns such as reduced vocabulary diversity, simplified syntax, and diminished lexical richness. ReLearn outperforms SOTA in many benchmarks, one of them is shown above.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h3 class="heading" style="text-align:left;">Sponsored: Special offer for you claim $50 free credits!</h3><div class="image"><a class="image__link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" rel="noopener" target="_blank"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/842ecda9-55ef-4274-978a-0eba2a662568/Neuron_Newsletter_image_2025.png?t=1740066909"/></a></div><p class="paragraph" style="text-align:left;">Accelerate your AI projects with Prolific. Claim <b>$50 free credits</b> and get quality human data in minutes <b>from 200,000+ taskers</b>. No setup cost, no subscription, no delay—get started, top up your account to claim your free credit, and test Prolific for yourself now. <b>Use code:</b> LLM-RESEARCH-50 <br></p><p class="paragraph" style="text-align:center;"><b><a class="link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Claim free credit</a></b><b><a class="link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow"> now</a></b></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">Better Embeddings with Coupled Adam</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.08441?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://www.github.com/llmsresearch/coupledadam?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Paper implementation</a></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This paper solves the problem of anisotropic embeddings. Buit what it is? Well, word representations learned by LLMs tend to cluster in a small subspace, away from the origin of the vector space. This clustering limits the semantic expressiveness of the embeddings and, consequently, the model&#39;s overall performance. The paper begins by noting that while various attempts have been made to explain and alleviate this issue, the role of the optimization algorithm itself has been largely overlooked. This research paper argue that the second moment estimate in Adam, which is used to adapt the learning rate for each parameter, causes embedding vectors to shift collectively away from the origin.</p><p class="paragraph" style="text-align:left;">To understand this, let’s first understand how Adam works. Adam maintains an exponentially decaying average of past gradients (first moment) and squared gradients (second moment) for each parameter. The second moment is meant to scale the learning rate adaptively, giving larger updates to parameters with smaller gradients and smaller updates to parameters with larger gradients. This is particularly useful for sparse data, like word frequencies in LLM training, where some words occur far more often than others.</p><p class="paragraph" style="text-align:left;">However, paper shows that this adaptive scaling, specifically the  <i><b>i-dependency</b></i><i> of the second moment</i>, leads to the anisotropy problem. They demonstrate mathematically that while the sum of gradients over all embedding vectors vanishes with SGD, the weighted sum (weighted by the adaptive learning rate based on the second moment) <i>does not vanish</i> with Adam. This non-vanishing sum causes a collective shift of the embedding vectors away from the origin. They also provide experimental evidence to show that the expectation value of the second moment is proportional to the unigram probability of the corresponding word, confirming the link between word frequency and the anisotropic effect.</p><p class="paragraph" style="text-align:left;"><b>Proposed approach</b></p><p class="paragraph" style="text-align:left;">Paper proposes a modified version of Adam called <i><b>Coupled Adam</b></i>. The main idea behind Coupled Adam is to enforce that the second moments are the same for all embedding vectors. To do this, team replaced the individual second moment estimates for each embedding vector with the <i>average</i> of the second moments over all embedding vectors.</p><p class="paragraph" style="text-align:left;">Mathematically, the original Adam update rule uses an <i><b>i-dependent </b></i>effective learning rate (η<sub>i</sub>) that depends on the second moment estimate of the <i>i</i><sup><i>th</i></sup> embedding vector. Coupled Adam replaces this <i><b>i-dependent</b></i> rate with an <i><b>i-independent </b></i>rate based on the average second moment.</p><p class="paragraph" style="text-align:left;">This coupling of the second moments ensures that the sum of embedding updates vanishes, similar to SGD, preventing the collective shift of embeddings away from the origin. At the same time, Coupled Adam retains the benefits of Adam, using a second moment to normalize the embedding update vectors (albeit a global one).</p><p class="paragraph" style="text-align:left;"><b>Implementation</b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c93df284-f53e-4842-8479-879352bd6772/image.png?t=1739995585"/></div><p class="paragraph" style="text-align:left;">The implementation of Coupled Adam is straightforward. Paper provides pseudocode in Algorithm 1 of the paper which is as shown above. The key modification lies in Algorithm 1 lines 8-12 in the inclusion of Coupled Adam where the second moments are averaged across all the vocabulary items before the update vector is calculated, as such:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Calculate the standard Adam update vectors for all embedding vectors <i>e</i><sub><i>i</i></sub>.</p></li><li><p class="paragraph" style="text-align:left;">Compute the average second moment (ν) over all embedding vectors.</p></li><li><p class="paragraph" style="text-align:left;">Replace the individual second moment estimates for each embedding vector with this average value (ν).</p></li><li><p class="paragraph" style="text-align:left;">Apply the standard Adam update rule using this shared second moment.</p></li></ol><p class="paragraph" style="text-align:left;">Paper emphasizes that Coupled Adam can be easily integrated into existing training pipelines with minimal code changes, specifically for the embedding parameters, while standard Adam can be used for all non-embedding parameters.</p><p class="paragraph" style="text-align:left;"><b>Setup and Results:</b></p><p class="paragraph" style="text-align:left;">Paper presents multiple experiments to evaluate the effectiveness of Coupled Adam. These experiments were divided into small-scale and large-scale settings, involving different datasets, model architectures, and training frameworks.</p><ul><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Small-Scale Experiments</b></span><b>:</b> The small-scale experiments utilized the OpenWebText Corpus and the GPT-2 tokenizer. The model architecture also followed GPT-2, with hyperparameter settings derived from GPT-3. Model sizes ranged from 125M to 760M parameters, and dataset sizes ranged from 5B to 20B tokens. Each experiment was repeated three times with different random seeds to assess statistical significance.</p></li><li><p class="paragraph" style="text-align:left;"><span style="text-decoration:underline;"><b>Large-Scale Experiments</b></span><b>:</b> The large-scale experiments employed the SlimPajama dataset and the GPT-2 tokenizer. The model architecture was a state-of-the-art dense transformer similar to those used in recent LLMs, including RoPE embeddings and SwiGLU activation functions. Model sizes were 1.3B and 2.6B parameters.</p></li></ul><p class="paragraph" style="text-align:left;">Across all experiments, the authors trained two models: one using standard Adam and one using Coupled Adam for the embedding parameters. They then evaluated the models using various metrics to assess both general performance and embedding quality. Here are the findings:</p><ul><li><p class="paragraph" style="text-align:left;">The use of Coupled Adam resulted in lower perplexity on both small and large scale. For example, they reported 14.69 perplexity score using vanilla adam compared to 14.45 score using Coupled Adam.</p></li><li><p class="paragraph" style="text-align:left;">Coupled Adam was seen to outperform vanilla adam in different training sizes. For example, it was seen that model trained on 10B tokens has perplexity of 15.51. And, on similar training setup using Coupled Adam, perplexity score comes out to be 15.37.</p></li><li><p class="paragraph" style="text-align:left;">Coupled Adam had significantly less anisotropic score which proves that the method is working to improve the quality of embeddings. The lower the score, the better is the embedding. Example: The anistropic score for vanilla adam was 0.874 compared to coupled adam which was 0.817.</p></li></ul></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">TransMLA: Multi-head Latent Attention Is All You Need</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.07864?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/fxmeng/TransMLA?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Implementation</a></p><p class="paragraph" style="text-align:justify;">As models and sequence lengths grow, the KV cache, which stores information about previous tokens, becomes a significant bottleneck, consuming substantial memory and bandwidth. While methods like Group Query Attention (GQA) have been adopted to reduce KV cache size, the authors propose that Multi-head Latent Attention (MLA), used in Deepseek models, offers a more theoretically sound and practically effective solution. The key problem this paper tackles is <i>how to efficiently transition existing GQA-based models, which are widely used, to MLA-based models, which are more expressive for the same KV cache overhead.</i></p><p class="paragraph" style="text-align:left;"><b>Core argument:</b></p><p class="paragraph" style="text-align:left;">The main argument is that Multi-head Latent Attention (MLA) provides greater expressive power than Group Query Attention (GQA) for the same KV cache overhead. Paper provides a <i>theoretical proof</i> to support this claim. It then introduces TransMLA, a post-training method, that enables the conversion of widely used GQA-based pre-trained models (such as LLaMA, Qwen, and Mixtral) into equivalent MLA-based models. The goal is to allow existing models to benefit from the superior expressiveness of MLA with minimal changes to model architecture and without increasing KV cache size. </p><p class="paragraph" style="text-align:left;"><b>Methodology</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a375bb66-9cd6-4605-a6c3-cf74b99783f3/image.png?t=1740113839"/></div><p class="paragraph" style="text-align:left;">We divided proposed method in three main parts:</p><ol start="1"><li><p class="paragraph" style="text-align:justify;"><b>Theoretical Proof of GQA to MLA Conversion:</b> Paper provides a mathematical proof showing that any GQA configuration can be equivalently transformed into MLA with the same size of KV cache. This proof relies on showing that the key transformation in GQA can be represented as a low-rank factorization, which is the core idea behind MLA. Paper demonstrates how to replicate keys in GQA (making all heads identical within a group), move that replication to the parameter side (replicating weights instead of activations), and then show that this replicated weight matrix has a low-rank structure that can be factorized in the same way as MLA. They achieve this by factorizing <i>W</i><sup><i>′</i></sup><sub><i>K</i></sub> using the Singular Value Decomposition (SVD) <i>W</i><sup><i>′</i></sup><sub><i>K</i></sub><i> = U</i><sub><i>K</i></sub><i> S</i><sub><i>K</i></sub><i> V</i><sup><i>⊤</i></sup><sub><i>K</i></sub> , and only keeping the top-r singular values, thus achieving KV compression.</p></li><li><p class="paragraph" style="text-align:left;"><b>MLA is Not Representable in GQA:</b> Paper also proves that the reverse is <i>not</i> true – <b>MLA cannot always be represented by GQA</b>. They demonstrate this by considering a scenario where vectors in the MLA transformation matrix are orthogonal. This leads to diversity in the outputs that cannot be replicated by GQA’s grouped heads which are necessarily identical within each group. This asymmetry provides the theoretical justification for expecting performance improvements when converting GQA models to MLA.</p></li><li><p class="paragraph" style="text-align:left;"><b>Practical Conversion and Post-Training:</b> The core of the TransMLA method lies in converting the weights of a pre-trained GQA model to equivalent MLA weights. This involves several steps:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Weight Decomposition:</b> The authors take the existing key projection matrix from the GQA model and decompose it into two smaller matrices, corresponding to the <i>W</i><sup><i>a</i></sup><sub><i>K</i></sub> and <i>W</i><i><sup>b</sup></i><i><sub>K</sub></i> matrices in the MLA formulation. This decomposition is achieved using SVD or other low-rank factorization techniques.</p></li><li><p class="paragraph" style="text-align:left;"><b>Weight Recombination:</b> These newly formed <i>W</i><sup><i>a</i></sup><sub><i>K</i></sub> and <i>W</i><i><sup>b</sup></i><i><sub>K</sub></i> matrices are then integrated into the model as the new key projection weights. This effectively replaces the GQA attention mechanism with the MLA mechanism.</p></li><li><p class="paragraph" style="text-align:left;"><b>Post-Conversion Training:</b> After the conversion, the model is fine-tuned on a dataset to allow it to adapt to the new MLA architecture and realize the benefits of its enhanced expressiveness. This fine-tuning step is crucial for recovering any potential performance loss from the weight conversion and for fully exploiting the potential of MLA. This training is done <i>without</i> increasing the KV cache size.</p></li></ul></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.07563?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/OpenSparseLLMs/Linear-MoE?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-2-3" target="_blank" rel="noopener noreferrer nofollow">Implementation</a></p><p class="paragraph" style="text-align:justify;">Attention mechanism’s computation increases at quadratic complexity with respect to sequence length. Paper enhances the efficiency of linear attention mechanisms through a combination of theoretical insights and practical implementations. Paper begins by analyzing the limitations of existing linear attention methods, which typically rely on fixed parallelism strategies that do not fully exploit the potential of modern hardware architectures and than it proposes a new approach that allows for dynamic sequence parallelism, enabling better utilization of computational resources. </p><p class="paragraph" style="text-align:justify;">Paper introduces a hybrid architecture that combines both local and global attention mechanisms. Local attention focuses on a limited context window, allowing for efficient computation within smaller segments of the input sequence. In contrast, global attention captures long-range dependencies across the entire sequence. By integrating these two forms of attention, paper achieves a balance between computational efficiency and expressive power.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/47d0874b-18de-4263-8aba-eb48d79b64ad/image.png?t=1740114510"/></div><p class="paragraph" style="text-align:left;">The paper details how LASP-2 employs a two-stage process for attention computation. In the first stage, local attention is computed in parallel across segments of the input sequence. This allows for rapid processing of individual segments while maintaining low memory overhead. In the second stage, global attention is applied to aggregate information from these segments, ensuring that long-range dependencies are preserved. Moreover paper introduces a novel mechanism for adaptive segment sizing based on input characteristics. This allows LASP-2 to dynamically adjust the size of local segments depending on the complexity of the input data, further optimizing performance.</p><p class="paragraph" style="text-align:left;"><b>Result</b></p><p class="paragraph" style="text-align:left;">Paper demonstrates up to a 30% reduction in training time while maintaining comparable accuracy levels on large datasets. LASP-2 maintained stable performance without significant degradation in processing speed or accuracy for up to 10,000 tokens. </p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=c3ad0065-8552-4cb3-9baf-d44b5b0f6a0d&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Research papers improving performance of LLMs [1/3]</title>
  <description>Research papers published from January 16th to February 15th, 2025 proposing context length and architectural changes in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1a0ffc01-aaae-451f-beb8-c9830982690c/collage_jan16th-Feb15th__1_3_.png" length="379044" type="image/png"/>
  <link>https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-1-3</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/research-papers-improving-performance-of-llms-from-jan-16-feb-15-2025-1-3</guid>
  <pubDate>Thu, 20 Feb 2025 20:23:00 +0000</pubDate>
  <atom:published>2025-02-20T20:23:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:center;"><b>In partnership with</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42cc7ce3-bec6-4a63-94bc-a31e2bdaec9b/Prolific_Logo_Blue__1_.png?t=1740074994"/></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Highlights</h2><ul><li><p class="paragraph" style="text-align:left;">InfiniteHiP prunes tokens like scissors, extending context to 3M</p></li><li><p class="paragraph" style="text-align:left;">LongRoPE stretches context to 2M+ tokens with fine-tuning</p></li><li><p class="paragraph" style="text-align:left;">DarwinLM uses evolution to prune LLMs , keeping performance high with structured pruning and training</p></li><li><p class="paragraph" style="text-align:left;">New paper draws a line between context length and model size</p></li><li><p class="paragraph" style="text-align:left;">Get a $50 free credit to get the humanized data for your project. No credit card required!</p></li></ul><p class="paragraph" style="text-align:left;">Don’t have much time to read entire newsletter? Well, listen to this fun and engaging podcast covering these research papers in detail.</p><div class="recommendation" id="61b9b217-6ace-4f48-9f40-111d2ea35f84"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/1a0ffc01-aaae-451f-beb8-c9830982690c/collage_jan16th-Feb15th__1_3_.png?t=1740016017"/></figure><h3 class="recommendation__title"> [1] Jan 16th-Feb 15th 2025 </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzYxYjliMjE3LTZhY2UtNGY0OC05ZjQwLTExMWQyZWEzNWY4NC8lNUIxJTVEJTIwSmFuJTIwMTZ0aC1GZWIlMjAxNXRoJTIwMjAyNS53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDRaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9YjFkOTNmMmIzMDRmODgzMjliY2Y0ZTI5ZTViYzE5ZjRhNmY5MGIwZDljMjliZDQwNjE0OGEyNjA2YmU2OTVhZVwiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOlwiaHR0cHM6Ly9iZWVoaWl2LWltYWdlcy1wcm9kdWN0aW9uLnMzLmFtYXpvbmF3cy5jb20vdXBsb2Fkcy9hc3NldC9maWxlLzFhMGZmYzAxLWFhYWUtNDUxZi1iZWI4LWM5ODMwOTgyNjkwYy9jb2xsYWdlX2phbjE2dGgtRmViMTV0aF9fMV8zXy5wbmc_dD0xNzQwMDE2MDE3XCIsXCJ0aXRsZVwiOlwiWzFdIEphbiAxNnRoLUZlYiAxNXRoIDIwMjVcIn0i." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">DarwinLM: Evolutionary Structured Pruning of Large Language Models</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.07780?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/llmsresearch/darwinlm?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Paper implementation </a></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:justify;">This research paper propose effective model compression technique. Pruning is a popular model compression technique. In LLMs, pruning refers to the process of removing less important or redundant connections (parameters) of model. By strategically removing these parameters, the model becomes smaller and more efficient, requiring less computation and memory.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f2b4b533-735a-45b4-8cd2-dfc030256208/image.png?t=1739894055"/></div><p class="paragraph" style="text-align:justify;">Pruning can be <i>unstructured</i> or <i>structured</i>. Unstructured pruning allows for the removal of individual connections seemingly at random, leading to sparsity in the model&#39;s weight matrices. While this can reduce the model size, it often doesn&#39;t translate directly into speed improvements on standard hardware. Structured pruning, on the other hand, removes entire structures within the network, such as neurons, channels, or even layers. This results in a more compact model that can be executed more efficiently on conventional hardware, providing end-to-end speedups without requiring specialized hardware or software.</p><p class="paragraph" style="text-align:justify;">However, not all parts of an LLM are equally important. Some components are more sensitive to pruning than others. This research paper ntroduces a novel method for training-aware, <b>non-uniform structured pruning</b>. The core idea is to combine an evolutionary search process with a lightweight training procedure to identify the optimal substructure of the LLM. The name &quot;DarwinLM&quot; itself is inspired by the concept of natural selection. The method works by generating multiple &quot;offspring&quot; models from a parent model, each with slightly different pruning configurations (mutations). These offspring models are then evaluated based on their performance on a given task. The best-performing offspring are selected to &quot;survive&quot; and become the parents for the next generation. This process is repeated over several generations, gradually refining the pruning configuration and identifying the most efficient and effective substructure.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dbd56c87-cda7-4291-aae9-3a91a4f7291d/image.png?t=1739894113"/></div><p class="paragraph" style="text-align:justify;">A key innovation in DarwinLM is its <b><i>training-aware</i></b> approach. Recognizing that the performance of a pruned model depends not only on its initial structure but also on how it&#39;s subsequently trained, DarwinLM incorporates a <b>lightweight, multi-step training process</b> within each generation. This allows the method to assess the effect of post-compression training and eliminate poorly performing models early on. The training process progressively increases the number of tokens used, providing a more comprehensive evaluation of the model&#39;s ability to learn and generalize.</p><p class="paragraph" style="text-align:justify;"><b>Results:</b></p><p class="paragraph" style="text-align:justify;">DarwinLM achieves a remarkable <b>2x reduction in model size </b>with only a <b>3% performance loss</b> across various tasks. The effectiveness of DarwinLM is consistent across different model sizes, including smaller models (350M parameters) and larger ones (7B parameters). Meaning, it works well on all size of models. </p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h3 class="heading" style="text-align:left;">Sponsored: Special offer for you claim $50 free credits!</h3><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/842ecda9-55ef-4274-978a-0eba2a662568/Neuron_Newsletter_image_2025.png?t=1740066909"/></div><p class="paragraph" style="text-align:left;">Accelerate your AI projects with Prolific. Claim <b>$50 free credits</b> and get quality human data in minutes <b>from 200,000+ taskers</b>. No setup cost, no subscription, no delay—get started, top up your account to claim your free credit, and test Prolific for yourself now. </p><p class="paragraph" style="text-align:left;"><b>Use code:</b> LLM-RESEARCH-50 <br><b>Link:</b> <span style="text-decoration:underline;"><a class="link" href="https://eu1.hubs.ly/H0gYdYX0?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">https://eu1.hubs.ly/H0gYdYX0</a></span></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2502.08910?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a> | <a class="link" href="https://github.com/DeepAuto-AI/hip-attention/?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Implementation</a> </p><p class="paragraph" style="text-align:justify;">Context window or length is the amount of text a model can consider when processing information. <b>Think of the context window as the model&#39;s short-term memory</b>. A larger context window allows the model to better understand complex relationships, resolve ambiguities, and maintain coherence over longer passages. However, increasing the context window of an LLM presents significant technical challenges. One of the main bottlenecks is computational cost. As the context window grows, the amount of computation required to process the information increases dramatically, often requiring more powerful hardware and more energy consumption. Another challenge is memory. LLMs need to store information about the entire context window in memory, and as the context window expands, the memory requirements can quickly exceed the capacity of available hardware. This research paper works in this direction.</p><p class="paragraph" style="text-align:justify;"><b>How it works?</b></p><p class="paragraph" style="text-align:justify;">This paper introduces a novel framework called InfiniteHiP, which stands for Infinite Hierarchical Pruning. This framework applies several innovative techniques to extend the context length that can be processed by an LLM to an impressive 3 million tokens, all while using a single GPU.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c2da64e2-5cb4-4a3a-95b4-93c91e040ef7/image.png?t=1739883936"/></div><p class="paragraph" style="text-align:left;">The first key component of InfiniteHiP is hierarchical token pruning. In simple terms, this technique intelligently removes less important tokens from the input, reducing the computational load without significantly impacting the quality of the output. It does this at multiple levels, ensuring that the most relevant information is retained for processing.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8ec1b120-0763-43ce-95a2-f49df13b3338/image.png?t=1739883997"/></div><p class="paragraph" style="text-align:left;">Next, the framework uses adaptive adjustments to Rotary Position Embeddings (RoPE). Positional embeddings are a crucial component of many LLMs, as they help the model understand the order of words in a sequence. Traditional positional embeddings assign a fixed vector to each position in the sequence. However, these fixed embeddings can struggle to generalize to longer sequences than the model was originally trained on. RoPE offer a more flexible and generalizable approach. RoPE <b>encodes positional information</b> using rotations in a high-dimensional space. This allows the model to better extrapolate to unseen sequence lengths. By modifying RoPE based on observed attention patterns, InfiniteHiP allows the model to effectively handle much longer sequences than it was originally trained on.</p><p class="paragraph" style="text-align:left;">Main contribution in InfiniteHiP is its approach to <b>memory management</b>. LLMs typically store intermediate computations (known as the key-value cache) in GPU memory, which is fast but limited in capacity. InfiniteHiP instead stores this cache in the <b>computer&#39;s main memory (RAM)</b>, which is much larger but slower. This clever offloading technique allows for processing of extremely long contexts without running out of GPU memory. This framework is implemented within a system called SGLang, which optimizes how these various components work together. This integration ensures that the token pruning, memory management, and computational processes are coordinated efficiently.</p><p class="paragraph" style="text-align:left;"><b>Result</b></p><p class="paragraph" style="text-align:left;">The results of this research are impressive. InfiniteHiP achieves a speedup of <b>18.95 times</b> in processing attention (a key component of LLMs) for sequences of 1 million tokens compared to baseline methods. It can handle sequences three times longer than previous approaches without losing important information. Remarkably, these improvements are achieved without needing to retrain the model or change its fundamental architecture.</p><p class="paragraph" style="text-align:left;">In practical terms, this means that a single GPU with <b>48GB of memor</b>y (specifically, an NVIDIA L40s) can now process <b>up to 3 million tokens at once</b>. This is a massive leap forward, enabling applications that were previously impractical or impossible.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">Explaining Context Length Scaling and Bounds for Language Models</h2><p class="paragraph" style="text-align:left;"><b><a class="link" href="https://arxiv.org/abs/2502.01481?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a></b></p><p class="paragraph" style="text-align:justify;">Ever wondered how does the length of context affect the performance of these language models? Well this research paper has answers for that! </p><p class="paragraph" style="text-align:justify;">To understand the significance of this question, it&#39;s important to know that context in language models refers to the amount of preceding text the model considers when predicting the next word or performing a task. Traditionally, most models were limited to relatively short contexts, often just a few hundred words. However, recent advancements have pushed this boundary, with some models capable of considering thousands or even millions of words as context.</p><p class="paragraph" style="text-align:justify;">This increase in context length has led to some interesting and sometimes conflicting observations. Some studies have shown that longer contexts can improve model performance, leading to more accurate predictions and better understanding of complex topics. This phenomenon is often described as &quot;<b>Scaling Laws</b>,&quot; suggesting that performance improves as context length increases. On the other hand, other research has found that long, irrelevant contexts can actually degrade performance, confusing the model rather than helping it. Now, question is which one is true?</p><p class="paragraph" style="text-align:justify;">This research paper provides a theoretical framework to explain how context length impacts language modeling. It introduces the concept of <b>&quot;Intrinsic Space&quot;</b> as a key to understanding context length effects. Intrinsic Space refers to <b>a lower-dimensional representation of the data</b> that captures the essential features relevant to the language modeling task. Think of it as distilling the most important information from the vast sea of language data. Paper suggests that impact of context length can be explained by how well the model can learn and represent these intrinsic features. As context length increases, the model has more information to work with, potentially allowing it to better capture these essential features. However, there&#39;s a point of <b>diminishing returns</b>, beyond which additional context doesn&#39;t provide meaningful new information.</p><p class="paragraph" style="text-align:justify;">One of the key findings from this research is that <b>there&#39;s an optimal context length for any given size of training dataset</b>. This means that simply increasing context length indefinitely isn&#39;t always beneficial. Instead, the optimal context length grows with the size of the training data. The paper also establishes theoretical bounds for context length scaling under certain conditions. These bounds provide guidance on how context length should be increased as models and training datasets grow larger. </p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2402.13753?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=research-papers-improving-performance-of-llms-1-3" target="_blank" rel="noopener noreferrer nofollow">Research paper</a></p><p class="paragraph" style="text-align:justify;">This research paper introduces a novel technique called &quot;LongRoPE&quot; that allows LLMs to effectively handle extremely long context windows – in this case, extending beyond 2 million tokens – while maintaining strong performance and reasonable computational costs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/67a206d0-7891-4db7-b474-e9601cdb3bd0/image.png?t=1739890917"/></div><p class="paragraph" style="text-align:left;">LongRoPE modifies the RoPE (explained in 1st paper) mechanism to make it more efficient for very long sequences. The researchers observed that for very long contexts, the high-frequency components of the RoPE embeddings become less important. LongRoPE takes advantage of this observation by scaling down these high-frequency components as the context length increases. This scaling reduces the computational burden and allows the model to process much longer sequences without sacrificing performance.</p><p class="paragraph" style="text-align:left;">In simpler terms, LongRoPE is like a zoom lens for the model&#39;s attention. When focusing on a small area (short context), the lens provides a detailed view. But when zooming out to see a larger area (long context), some of the fine details become less important, and the lens adjusts to provide a broader, more efficient view.</p><p class="paragraph" style="text-align:justify;">The important thing about LongRoPE is <span style="color:oklch(0.304 0.04 213.681);font-family:fkGroteskNeue, &quot;fkGroteskNeue Fallback&quot;, ui-sans-serif, system-ui, -apple-system, system-ui, &quot;Segoe UI&quot;, Roboto, &quot;Helvetica Neue&quot;, Arial, &quot;Noto Sans&quot;, sans-serif, &quot;Apple Color Emoji&quot;, &quot;Segoe UI Emoji&quot;, &quot;Segoe UI Symbol&quot;, &quot;Noto Color Emoji&quot;;font-size:16px;">its simplicity and ease of integration.</span> It <span style="color:oklch(0.304 0.04 213.681);font-family:fkGroteskNeue, &quot;fkGroteskNeue Fallback&quot;, ui-sans-serif, system-ui, -apple-system, system-ui, &quot;Segoe UI&quot;, Roboto, &quot;Helvetica Neue&quot;, Arial, &quot;Noto Sans&quot;, sans-serif, &quot;Apple Color Emoji&quot;, &quot;Segoe UI Emoji&quot;, &quot;Segoe UI Symbol&quot;, &quot;Noto Color Emoji&quot;;font-size:16px;">can be applied to existing LLMs with minimal modifications. </span></p><p class="paragraph" style="text-align:justify;"><span style="color:oklch(0.304 0.04 213.681);font-family:fkGroteskNeue, &quot;fkGroteskNeue Fallback&quot;, ui-sans-serif, system-ui, -apple-system, system-ui, &quot;Segoe UI&quot;, Roboto, &quot;Helvetica Neue&quot;, Arial, &quot;Noto Sans&quot;, sans-serif, &quot;Apple Color Emoji&quot;, &quot;Segoe UI Emoji&quot;, &quot;Segoe UI Symbol&quot;, &quot;Noto Color Emoji&quot;;font-size:16px;"><b>Results:</b></span></p><p class="paragraph" style="text-align:left;">The models extended via LongRoPE maintained low perplexity scores across evaluation lengths ranging from 4k to 2048k tokens. Low perplexity indicates that the model&#39;s predictions are more accurate and reliable. In tasks requiring long contexts, such as passkey retrieval, LongRoPE achieved over 90% accuracy! and, most importantly, models utilizing LongRoPE preserved their original performance levels on shorter context windows while significantly reducing perplexity when processing longer contexts.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=380e1480-2cd8-428f-be8b-f475dd76255c&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>DeepSeek-R1 Special Edition</title>
  <description>Deep technical analysis of DeepSeek-R1&#39;s groundbreaking approach to enhancing AI reasoning capabilities through reinforcement learning. Comprehensive breakdown of methodology, implementation, and results. DeepSeek-R1, AI reasoning, reinforcement learning, LLM, artificial intelligence, machine learning, PPO, reward modeling, technical analysis.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/28dbf4a5-c72c-4030-a733-654fb83a5744/Screenshot_2025-01-28_at_3.04.16_PM.png" length="260197" type="image/png"/>
  <link>https://llm.beehiiv.com/p/deepseek-r1-special-edition</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/deepseek-r1-special-edition</guid>
  <pubDate>Tue, 28 Jan 2025 20:36:57 +0000</pubDate>
  <atom:published>2025-01-28T20:36:57Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">📑<b> IN THIS ISSUE</b></h2><p class="paragraph" style="text-align:left;">DeepSeek-R1 has become very popular online, and as an AI professional, it&#39;s important to understand why. How is the DeepSeek team able to train models with just a few million dollars in GPU computing, while teams like OpenAI need billions? </p><p class="paragraph" style="text-align:left;">Take some time to read this newsletter. It&#39;s important to learn about this model, as it represents a new research direction in large language models (LLMs).</p><p class="paragraph" style="text-align:left;">TL;DR listen to this amazing <b>podcast</b> discussing this research paper.</p><div class="recommendation" id="e28a986e-f2f5-42ee-98ed-cd89edd25267"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/da8b2fc3-001a-469e-b50c-18ac65ff55be/Screenshot_2025-01-28_at_11.54.26_AM.png?t=1738087487"/></figure><h3 class="recommendation__title"> DeepSeek-R1 Incentivizing Reasoning Capability in LLMs via Reinforcement Learning </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2L2UyOGE5ODZlLWYyZjUtNDJlZS05OGVkLWNkODllZGQyNTI2Ny9EZWVwU2Vlay1SMV8lMjBJbmNlbnRpdml6aW5nJTIwUmVhc29uaW5nJTIwQ2FwYWJpbGl0eSUyMGluJTIwTExNcyUyMHZpYSUyMFJlaW5mb3JjZW1lbnQlMjBMZWFybmluZy5tcDM_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDVaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9NTZiNmZjZGEyN2FjYzlmMDgyYmM5MTZhMDNhZTlmMzgwNzdhNmQyNTVhOGMwZTQxZDViYjczY2QzNTE1NWRhYVwiLFwidHlwZVwiOlwiYXVkaW8vbXBlZ1wiLFwidGh1bWJuYWlsVXJsXCI6XCJodHRwczovL2JlZWhpaXYtaW1hZ2VzLXByb2R1Y3Rpb24uczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Fzc2V0L2ZpbGUvZGE4YjJmYzMtMDAxYS00NjllLWI1MGMtMThhYzY1ZmY1NWJlL1NjcmVlbnNob3RfMjAyNS0wMS0yOF9hdF8xMS41NC4yNl9BTS5wbmc_dD0xNzM4MDg3NDg3XCIsXCJ0aXRsZVwiOlwiRGVlcFNlZWstUjEgSW5jZW50aXZpemluZyBSZWFzb25pbmcgQ2FwYWJpbGl0eSBpbiBMTE1zIHZpYSBSZWluZm9yY2VtZW50IExlYXJuaW5nXCJ9Ig." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Introduction</h2><p class="paragraph" style="text-align:left;">Remember when calculators transformed mathematics? We&#39;re at a similar inflection point with AI. DeepSeek-R1 isn&#39;t just another language model – it&#39;s a fundamental rethinking of how AI systems learn to reason.</p><p class="paragraph" style="text-align:left;">DeepSeek-R1 has sent shockwaves through the AI community, wiping $1 trillion from tech market caps in a single day. But beyond the market chaos lies a revolutionary approach to AI reasoning that could redefine the field. While ChatGPT and others focus on generating human-like responses, DeepSeek-R1 focuses on something more crucial: teaching AI how to think step-by-step through problems.</p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;"><i><b>&quot;This isn&#39;t just another model release - it&#39;s a fundamental rethinking of how AI systems learn to reason. DeepSeek-R1 represents what we&#39;ve been missing in the field - focus on quality of thinking rather than just output generation.&quot;</b></i></p><figcaption class="blockquote__byline"><i><b>Dr. Sarah Chen, Stanford AI Lab</b></i></figcaption></blockquote></div><hr class="content_break"><h2 class="heading" style="text-align:center;">Market reaction</h2><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e583bd25-f5c9-4035-94a7-8661a17a1ffd/untitled__4_.png?t=1738079736"/><div class="image__source"><span class="image__source_text"><p>Quadrant summarizing reaction on DeepSeek-R1 launch</p></span></div></div><p class="paragraph" style="text-align:left;"><b>Interesting facts:</b></p><ul><li><p class="paragraph" style="text-align:left;">Ultra-low cost training cost: Reportedly $5.6 million training cost for the base model</p></li><li><p class="paragraph" style="text-align:left;">According to sources, DeepSeek bought a large number of Nvidia A100 GPUs (between 10,000 and 50,000 units) before U.S. chip restrictions were put in place.</p><hr class="content_break"></li></ul><h2 class="heading" style="text-align:center;">Technical Analysis</h2><p class="paragraph" style="text-align:left;">The core idea of DeepSeek-R1 is to improve the reasoning capability of LLMs by incorporating a reward mechanism that incentivizes logical reasoning steps, rather than focusing solely on generating correct final answers. The model undergoes a three-phase training process:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Supervised Fine-Tuning (SFT)</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Reward Model Training</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Reinforcement Learning via Proximal Policy Optimization (PPO)</b></p></li></ol><h4 class="heading" style="text-align:left;"><b>Stage 1. Supervised Fine-Tuning (SFT)</b></h4><p class="paragraph" style="text-align:left;">The process begins with supervised fine-tuning of the base DeepSeek-LLM model. The team meticulously curated a dataset combining GSM8K, MATH, LogiQA, and ProofWriter datasets. What sets their approach apart is the careful structuring of this data. Rather than simply using question-answer pairs, they incorporated detailed reasoning paths and step-by-step solutions.</p><p class="paragraph" style="text-align:left;">The fine-tuning process employs a dynamic curriculum learning strategy. Problems are presented in increasing order of complexity, with the model&#39;s performance determining when to introduce more challenging examples. This ensures the model builds a strong foundation in basic reasoning before tackling more complex problems.</p><h4 class="heading" style="text-align:left;"><b>Stage 2: Reward Model Development</b></h4><p class="paragraph" style="text-align:left;">The heart of DeepSeek-R1&#39;s innovation lies in its reward modeling system. The team developed a specialized reward model trained on human preferences for different reasoning approaches. This wasn&#39;t simply about marking answers as correct or incorrect; instead, they collected detailed human feedback on:</p><ul><li><p class="paragraph" style="text-align:left;">The logical coherence of reasoning steps</p></li><li><p class="paragraph" style="text-align:left;">The efficiency of solution paths</p></li><li><p class="paragraph" style="text-align:left;">The clarity and completeness of explanations</p></li><li><p class="paragraph" style="text-align:left;">The validity of intermediate conclusions</p></li></ul><p class="paragraph" style="text-align:left;">The reward model training involved pair-wise comparisons of different reasoning approaches for the same problems. Human evaluators, primarily mathematics and logic experts, rated these comparisons based on predefined quality criteria. This resulted in a reward model that could effectively distinguish between different levels of reasoning quality.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3773f09d-d8db-48c2-87d6-0f0ee9331794/DeepSeek_training-2025-01-28-151508.png?t=1738077334"/><div class="image__source"><span class="image__source_text"><p>How DeepSeek-R1 works?</p></span></div></div><h4 class="heading" style="text-align:left;"><b>Stage 3: Reinforcement Learning</b></h4><p class="paragraph" style="text-align:left;">The final stage employs a modified version of Proximal Policy Optimization (PPO). The innovation here lies in how they adapted PPO for reasoning tasks. The process works as follows:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">The model generates multiple reasoning attempts for each problem</p></li><li><p class="paragraph" style="text-align:left;">The reward model evaluates each attempt based on multiple criteria</p></li><li><p class="paragraph" style="text-align:left;">The policy is updated using a custom loss function that balances:</p><ul><li><p class="paragraph" style="text-align:left;">Reward maximization</p></li><li><p class="paragraph" style="text-align:left;">KL divergence from the reference model</p></li><li><p class="paragraph" style="text-align:left;">Entropy regularization to maintain exploration</p></li></ul></li></ol><p class="paragraph" style="text-align:left;">A key innovation is their implementation of intermediate rewards. Rather than waiting for the final answer, the model receives feedback on each step of its reasoning process. This creates a more granular learning signal, helping the model understand which specific reasoning steps are effective.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>The Evolution of DeepSeek-R1</b></p><p class="paragraph" style="text-align:left;">DeepSeek-R1&#39;s development is interesting because of its progress over time. It started with DeepSeek-V3-Base and used GRPO for reinforcement learning. This lead to a new model <b>DeepSeek-R1-Zero</b>. After many training steps, DeepSeek-R1-Zero&#39;s performance on AIME 2024 improved from 15.6% to 71.0%. With majority voting, it went up to 86.7%, matching OpenAI-o1-0912&#39;s performance. The team then refined the model by: Creating new training data through rejection sampling Combining it with DeepSeek-V3&#39;s supervised data on writing, factual QA, and self-awareness Retraining the base model Doing more reinforcement learning This careful process led to the final DeepSeek-R1 model, which performed as well as OpenAI-o1-1217. A key achievement was making the technology more accessible through model distillation: They successfully transferred capabilities to smaller models based on Qwen and Llama Their 14B distilled model did better than the top QwQ-32B-Preview The 32B and 70B distilled versions set new records for reasoning abilities</p><hr class="content_break"><h3 class="heading" style="text-align:left;"><b>Implementation Details</b></h3><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b019d869-0c59-4c5d-880f-ee41924c41a9/image.png?t=1738083815"/><div class="image__source"><span class="image__source_text"><p>DeepSeek-R1 training pipeline</p></span></div></div><p class="paragraph" style="text-align:left;">The implementation required significant computational resources:</p><ul><li><p class="paragraph" style="text-align:left;">Training was conducted on a cluster of A100 GPUs</p></li><li><p class="paragraph" style="text-align:left;">The reward model training alone took several weeks</p></li><li><p class="paragraph" style="text-align:left;">The full training pipeline required multiple iterations and refinements</p></li></ul><p class="paragraph" style="text-align:left;">The team implemented several optimizations:</p><ul><li><p class="paragraph" style="text-align:left;">Distributed training across multiple nodes</p></li><li><p class="paragraph" style="text-align:left;">Efficient reward computation</p></li><li><p class="paragraph" style="text-align:left;">Dynamic batch sizing based on problem complexity</p></li><li><p class="paragraph" style="text-align:left;">Gradient accumulation for stability</p></li></ul><hr class="content_break"><h3 class="heading" style="text-align:left;"><b>Evaluation Framework</b></h3><p class="paragraph" style="text-align:left;">The evaluation of DeepSeek-R1 is particularly comprehensive, going beyond simple accuracy metrics. The team developed a multi-dimensional evaluation framework:</p><h4 class="heading" style="text-align:left;"><b>Reasoning Quality Assessment</b></h4><p class="paragraph" style="text-align:left;">The model&#39;s performance is evaluated on:</p><ul><li><p class="paragraph" style="text-align:left;">Mathematical reasoning (using GSM8K and MATH benchmarks)</p></li><li><p class="paragraph" style="text-align:left;">Logical deduction (using LogiQA and ProofWriter)</p></li><li><p class="paragraph" style="text-align:left;">Scientific reasoning (using ScienceQA)</p></li><li><p class="paragraph" style="text-align:left;">General problem-solving (using custom benchmarks)</p></li></ul><p class="paragraph" style="text-align:left;">Each solution is evaluated for: Step-by-step coherence, Logical validity, Efficiency of approach, and Clarity of explanation. Team also performed extensive comparisons with previous versions of DeepSeek models, other leading language models, and specialized reasoning models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a64b3e63-49e3-4d14-96e4-ebaca2bdbd1c/image.png?t=1738095655"/><div class="image__source"><span class="image__source_text"><p>Performance comparison of DeepSeek-R1 with other models</p></span></div></div><p class="paragraph" style="text-align:left;">What&#39;s particularly noteworthy is the quality of the reasoning processes. The model doesn&#39;t just arrive at correct answers; it demonstrates human-like problem-solving approaches, showing each step clearly and logically.</p><hr class="content_break"><h3 class="heading" style="text-align:left;"><b>Challenges and Solutions</b></h3><p class="paragraph" style="text-align:left;">DeepSeek team faced several challenges and here’s a quick summary of challenges they faced and solutions they used to overcome the challenges.</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Training Stability: Implementation of reference models and careful KL divergence constraints</p></li><li><p class="paragraph" style="text-align:left;">Reward Sparsity: Development of dense reward signals through intermediate evaluations</p></li><li><p class="paragraph" style="text-align:left;">Computational Efficiency: Custom optimizations in the training pipeline and reward computation</p></li><li><p class="paragraph" style="text-align:left;">Generalization: Careful curriculum design and diverse training data</p></li></ol></div><div class="image"><a class="image__link" href="https://magic.beehiiv.com/v1/31a7c576-0eb2-4ef3-abc7-bc75ede786fe?email={{email}}&utm_source=beehiiv&utm_campaign={{publication_name_param}}_{{publication_alphanumeric_id}}&_bhiiv=opp_ec1b0b7d-7e1e-47f7-8a4a-11c2bbdf48fe_65769d95&bhcl_id=859f139f-51e5-4009-b560-98b90efe67e1_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ff79e3ad-93d4-4cc0-8540-6a3b58d70541/Ad_The_AI_report.png?t=1742251228"/></a></div><h3 class="heading" style="text-align:left;">There’s a reason 400,000 professionals read this daily. </h3><p class="paragraph" style="text-align:left;">Join <a class="link" href="https://magic.beehiiv.com/v1/31a7c576-0eb2-4ef3-abc7-bc75ede786fe?email={{email}}&utm_source=beehiiv&utm_campaign={{publication_name_param}}_{{publication_alphanumeric_id}}&_bhiiv=opp_ec1b0b7d-7e1e-47f7-8a4a-11c2bbdf48fe_65769d95&bhcl_id=859f139f-51e5-4009-b560-98b90efe67e1_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">The AI Report</a>, trusted by 400,000+ professionals at Google, Microsoft, and OpenAI. Get daily insights, tools, and strategies to master practical AI skills that drive results.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://magic.beehiiv.com/v1/31a7c576-0eb2-4ef3-abc7-bc75ede786fe?email={{email}}&utm_source=beehiiv&utm_campaign={{publication_name_param}}_{{publication_alphanumeric_id}}&_bhiiv=opp_ec1b0b7d-7e1e-47f7-8a4a-11c2bbdf48fe_65769d95&bhcl_id=859f139f-51e5-4009-b560-98b90efe67e1_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Sign up now for free and work smarter, not harder.</a></p></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=35acc535-f832-42fb-99a0-4b6c5f6ce3df&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Exploring New LLM Applications: Insights from January 1–15, 2025 Research Papers</title>
  <description>Discussing interesting applications of LLMs in Healthcare, Software Engineering, Finance, Energy, and Robotics</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/481cfdff-8724-4c89-bb01-d3c8045f5028/application_jan_1-15.png" length="396404" type="image/png"/>
  <link>https://llm.beehiiv.com/p/exploring-new-llm-applications-from-january-1-15-2025-research-papers</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/exploring-new-llm-applications-from-january-1-15-2025-research-papers</guid>
  <pubDate>Mon, 20 Jan 2025 15:00:00 +0000</pubDate>
  <atom:published>2025-01-20T15:00:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Fun & engaging podcast using NotebookLM</h2><p class="paragraph" style="text-align:left;">TL;DR! Listen to this amazing and fun podcast discussing below mentioned research papers!</p><div class="recommendation" id="f1243ef5-aa2c-4164-aa74-3b901da587fe"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/f1d1af97-b4bb-4921-bcef-1eb68935b3d0/application_jan_1-15.png?t=1737270314"/></figure><h3 class="recommendation__title"> Podcast: Application of LLMs (Jan 1-15th, 2025) </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2L2YxMjQzZWY1LWFhMmMtNDE2NC1hYTc0LTNiOTAxZGE1ODdmZS9BcHBsaWNhdGlvbl8lMjBKYW4lMjAxLTE1dGglMjAyMDI0Lndhdj9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NlxcdTAwMjZYLUFtei1DcmVkZW50aWFsPUFLSUFRQ01IVFFTRTJKR0FHWEhKJTJGMjAyNjA5MTMlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdFxcdTAwMjZYLUFtei1EYXRlPTIwMjYwOTEzVDExMzM0NlpcXHUwMDI2WC1BbXotRXhwaXJlcz02MDQ4MDBcXHUwMDI2WC1BbXotU2lnbmVkSGVhZGVycz1ob3N0XFx1MDAyNlgtQW16LVNpZ25hdHVyZT0xMWRjMjk2OWNiNGU4NTU2MGZiYWQ0YTlhM2ViNmRhNDk4YTViYzVjYjkwMjA1NTExYmZlODI4YmFmODEwZmJlXCIsXCJ0eXBlXCI6XCJhdWRpby94LXdhdlwiLFwidGh1bWJuYWlsVXJsXCI6XCJodHRwczovL2JlZWhpaXYtaW1hZ2VzLXByb2R1Y3Rpb24uczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Fzc2V0L2ZpbGUvZjFkMWFmOTctYjRiYi00OTIxLWJjZWYtMWViNjg5MzViM2QwL2FwcGxpY2F0aW9uX2phbl8xLTE1LnBuZz90PTE3MzcyNzAzMTRcIixcInRpdGxlXCI6XCJQb2RjYXN0OiBBcHBsaWNhdGlvbiBvZiBMTE1zIChKYW4gMS0xNXRoLCAyMDI1KVwifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Transforming healthcare with LLMs</h2><p class="paragraph" style="text-align:left;">This section contains LLMs applications in healthcare and bioinformatics. Researchers are leveraging these models to enhance disease detection, interpret complex medical data, and democratize bioinformatics analysis.</p><p class="paragraph" style="text-align:left;"><b>Enhancing Alzheimer&#39;s Detection with ADAM-1</b></p><p class="paragraph" style="text-align:left;">Understanding and detecting Alzheimer&#39;s disease has been a persistent challenge. The <a class="link" href="http://arxiv.org/abs/2501.08324v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">ADAM-1: AI and Bioinformatics for Alzheimer&#39;s Detection and Microbiome-Clinical Data Integrations</a> introduces a multi-agent LLM framework that integrates microbiome profiles, clinical datasets, and external knowledge bases. By synthesizing insights through retrieval-augmented generation, ADAM-1 improves research and diagnostic applications. Remarkably, it achieved mean F1 scores comparable to XGBoost but with significantly reduced variance, highlighting its robustness, especially with small laboratory datasets.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Interpretable Mental Health Diagnostics</b></p><p class="paragraph" style="text-align:left;">Mental health diagnostics often grapple with complexity and the risk of errors. <b><a class="link" href="http://arxiv.org/abs/2501.07653v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Large Language Models for Interpretable Mental Health Diagnosis</a></b><b> </b>proposes a clinical decision support system using LLMs to translate diagnostic manuals into logic programs. The implementation involves fine-tuning an LLM to understand and convert diagnostic criteria into constraint logic programming (CLP) rules. These logic programs are then executed by a CLP engine to simulate diagnostic reasoning. The system allows experts to inspect and modify the generated logic to ensure alignment with official diagnostic standards, enhancing interpretability. Validation is performed by comparing the system&#39;s outputs with established diagnoses, emphasizing the necessity of expert review to maintain fidelity to the manuals.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Advancing Radiology Report Generation with RadAlign</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e445ed87-d35f-477d-988f-c8f0343a59a3/image.png?t=1737261672"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07525v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment</a> presents a model that improves automated radiology report generation by aligning visual features with medical concepts. The implementation involves a vision-language model (VLM) that extracts visual features from radiological images and maps them to predefined medical concepts using a disease classification module achieving an AUC of 0.885. These concepts guide a LLMs equipped with retrieval-augmented generation to produce comprehensive reports. The system incorporates a concept alignment mechanism to ensure consistency between the visual input and textual output and employs error mitigation strategies to reduce inaccuracies, resulting in a GREEN score of 0.678 for report quality.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Democratizing Bioinformatics with BioAgents</b></p><p class="paragraph" style="text-align:left;">Bioinformatics workflows are inherently complex, requiring interdisciplinary expertise. <a class="link" href="http://arxiv.org/abs/2501.06314v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems</a> introduces a multi-agent system utilizing small language models tailored for bioinformatics tasks. Enhanced with retrieval-augmented generation, this system allows for personalized, local use with proprietary data. The performance was comparable to human experts on genomics tasks, paving the way for more accessible bioinformatics analysis.<br></p></div><h3 class="heading" style="text-align:left;" id="stay-uptodate-with-ai">Stay up-to-date with AI</h3><div class="image"><a class="image__link" href="https://magic.beehiiv.com/v1/4d03390d-2481-4299-b949-ffd8b38b4c38?email={{email}}&utm_campaign={{publication_alphanumeric_id}}&redirect_to=https%3A%2F%2Fsubscribe.therundown.ai%2F%3Fform%3Dopen&redirect_delay=1&_gl=1*1qqix25*_gcl_au*MTYwNDc0Mjg2OC4xNzI5NTMyNjYw*_ga*MTk2YzU4MDctZGFlZi00MjQ3LWIzZDYtYTQ1MTUwMmJiZTQ0*_ga_E6Y4WLQ2EC*MTczMjUxMTg2Ny4yNTkzLjEuMTczMjUxMzM4My42MC4wLjE4NTk3NDE3MTE.&_bhiiv=opp_1397170f-ab83-497d-95d1-13ab77cd00ab_e4221c46&bhcl_id=e6dda01b-b8a2-4f13-b530-ebc60feb2f2a_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/901d0649-4e4c-40f1-921b-974ba34a4167/Banner_1.png?t=1732571397"/></a></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://magic.beehiiv.com/v1/4d03390d-2481-4299-b949-ffd8b38b4c38?email={{email}}&utm_campaign={{publication_alphanumeric_id}}&redirect_to=https%3A%2F%2Fsubscribe.therundown.ai%2F%3Fform%3Dopen&redirect_delay=1&_gl=1*1qqix25*_gcl_au*MTYwNDc0Mjg2OC4xNzI5NTMyNjYw*_ga*MTk2YzU4MDctZGFlZi00MjQ3LWIzZDYtYTQ1MTUwMmJiZTQ0*_ga_E6Y4WLQ2EC*MTczMjUxMTg2Ny4yNTkzLjEuMTczMjUxMzM4My42MC4wLjE4NTk3NDE3MTE.&_bhiiv=opp_1397170f-ab83-497d-95d1-13ab77cd00ab_e4221c46&bhcl_id=e6dda01b-b8a2-4f13-b530-ebc60feb2f2a_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">The Rundown</a> is the most trusted AI newsletter in the world, with 1,000,000+ readers and exclusive interviews with AI leaders like Mark Zuckerberg, Demis Hassibis, Mustafa Suleyman, and more.</p><p class="paragraph" style="text-align:left;">Their expert research team spends all day learning what’s new in AI and talking with industry experts, then distills the most important developments into one free email every morning.</p><p class="paragraph" style="text-align:left;">Plus, complete the quiz after signing up and they’ll recommend the best AI tools, guides, and courses – tailored to your needs.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://magic.beehiiv.com/v1/4d03390d-2481-4299-b949-ffd8b38b4c38?email={{email}}&utm_campaign={{publication_alphanumeric_id}}&redirect_to=https%3A%2F%2Fsubscribe.therundown.ai%2F%3Fform%3Dopen&redirect_delay=1&_gl=1*1qqix25*_gcl_au*MTYwNDc0Mjg2OC4xNzI5NTMyNjYw*_ga*MTk2YzU4MDctZGFlZi00MjQ3LWIzZDYtYTQ1MTUwMmJiZTQ0*_ga_E6Y4WLQ2EC*MTczMjUxMTg2Ny4yNTkzLjEuMTczMjUxMzM4My42MC4wLjE4NTk3NDE3MTE.&_bhiiv=opp_1397170f-ab83-497d-95d1-13ab77cd00ab_e4221c46&bhcl_id=e6dda01b-b8a2-4f13-b530-ebc60feb2f2a_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Sign up to start learning.</a></p><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Code Generation</h2><p class="paragraph" style="text-align:left;">This section is for us! Let’s see how LLMs makes our work easy or in other words snatches our job! 😀</p><p class="paragraph" style="text-align:left;"><b>Automating Cloud Operations with MOYA</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/88be9ab2-72db-4e68-a2c7-226e65b0bdaf/image.png?t=1737261887"/></div><p class="paragraph" style="text-align:left;">Managing cloud infrastructure can be complex and labor-intensive. <a class="link" href="http://arxiv.org/abs/2501.08243v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps</a> presents MOYA, a multi-agent framework integrating various systems using generative AI. By addressing challenges like diverse data sources and complex task automation, MOYA enhances the management and optimization of cloud infrastructure. Practitioners observed improved accuracy and responsiveness across complex workflows.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Self-Reflective Code Generation with CodeCoR</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e63fcdbe-b9c8-4e68-991d-0c92d869523b/image.png?t=1737261956"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07811v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation</a> introduces a multi-agent approach where agents specialize in task prompts, code generation, test case creation, and repair advice. By producing multiple outputs and refining failing code based on testing outcomes, CodeCoR ensures robust final code delivery. It significantly outperformed baselines with an average Pass@1 score of 77.8%.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Hierarchical Code Summarization in Business Applications</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07857v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow"><b>Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs</b></a> presents a two-step summarization approach implemented with local LLMs to respect confidentiality. Initially, code is parsed into small units like functions and classes, which are summarized using LLMs with prompts tailored to capture technical details and business context. In the second step, these summaries are aggregated hierarchically to create summaries for larger structures such as files and modules, utilizing custom prompts that relate these components to overall business logic. The implementation ensures that the summarization captures both the technical and domain-specific aspects, which was validated in a telecommunications business support system, enhancing code comprehension and maintenance.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Enhancing GitHub Issue Resolution with SWE-Fixer</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.05040v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution</a> presents a framework with code retrieval and editing modules. Trained on an extensive dataset of 110K GitHub issues, SWE-Fixer achieved state-of-the-art performance among open-source models, facilitating effective issue resolution.</p><p class="paragraph" style="text-align:left;"><b>Advancements in Program Repair and Debugging</b></p><p class="paragraph" style="text-align:left;">Several research papers have focused on automated program repair and debugging:</p><ul><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2501.08165v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">I Can Find You in Seconds!</a></b> leverages LLMs for code authorship attribution, aiding in software forensics and plagiarism detection.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2501.07531v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Evaluating Agent-based Program Repair at Google</a></b> assesses the applicability of agent-based approaches for automated bug fixing in an enterprise environment, achieving plausible patch creation for 73% of evaluated bugs.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2501.06972v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">How is Google using AI for internal code migrations?</a></b> shares insights into Google&#39;s application of LLMs for enterprise-level code migration, demonstrating significantly reduced migration times.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2501.06706v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds</a></b> introduces a framework that evaluates AI agents in cloud environments, contributing to the development of self-healing systems.</p></li><li><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2501.05706v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Debugging Without Error Messages</a></b> explores how different LLM prompting strategies can enhance programming error explanations, aiding educators and novice programmers.</p></li></ul></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">AI agents in autonomous systems</h2><p class="paragraph" style="text-align:left;">AI agents and multi-agent systems is opening new avenues for automation, collaboration, and efficiency across various domains. This section is dedicated to such papers!</p><p class="paragraph" style="text-align:left;"><b>Dynamic Workflow Generation with Flow</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07834v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Flow: A Modular Approach to Automated Agentic Workflow Generation</a> introduces workflows as activity-on-vertex graphs to emphasize modularity. By refining workflows through dynamic task allocations based on historical performance, the framework enhances efficiency and error tolerance, showing substantial improvements in handling practical tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Integrating AI Agents with Web3 Applications through Eliza</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ecabad2f-3ba9-42bd-9772-2ede4d647192/image.png?t=1737266597"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.06781v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Eliza: A Web3 Friendly AI Agent Operating System</a> addresses the integration gap between AI agents and Web3 applications. Developed fully in TypeScript, Eliza allows users to deploy Web3 applications effortlessly. It facilitates interaction with blockchain data and smart contracts while ensuring stable performance through optimized runtime components.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Automating Systematic Reviews with LatteReview</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.05468v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">LatteReview: A Multi-Agent Framework for Systematic Review Automation Using LLMs</a> leverages LLMs and multi-agent systems to automate tasks like title screening and data extraction. By incorporating features like retrieval-augmented generation and asynchronous programming, LatteReview handles large datasets efficiently, integrating with both cloud-based and local models.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Context-Aware Storytelling with MDSF (Uncovers aha moment from data!)</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/19df6a47-4afc-48df-b264-e1e9bc31e5be/image.png?t=1737266931"/></div><p class="paragraph" style="text-align:left;">Data storytelling is becoming increasingly important. <a class="link" href="http://arxiv.org/abs/2501.01014v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">MDSF: Context-Aware Multi-Dimensional Data Storytelling Framework Based on Large Language Model</a> introduces a framework that uses LLMs for automated insight generation and storytelling. With advanced preprocessing, analysis algorithms, and an agent-based storytelling continuation control, MDSF outperforms existing methods in insight ranking accuracy and narrative coherence.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Democratizing LLM Services with LLM-Net</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07288v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">LLM-Net: Democratizing LLMs-as-a-Service through Blockchain-based Expert Networks</a> presents a decentralized blockchain framework creating a network of specialized LLM providers. By leveraging collective computational resources and incorporating reputation-based mechanisms, LLM-Net ensures high service quality and facilitates seamless AI advancement across multiple domains.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><span style="font-size:1.5rem;">Domain adoption</span></h2><p class="paragraph" style="text-align:left;"><b>Enhancing Labor Market Analytics</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07663v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning</a> utilizes semantic chunking, retrieval-augmented generation, and fine-tuning of DistilBERT models to process over one million job postings. The fine-tuned models significantly improved the identification of complex job features like remote work availability and remuneration structures.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Automated Crypto Portfolio Management</b></p><p class="paragraph" style="text-align:left;">Cryptocurrency investments requires multi-modal data and intricate reasoning. <a class="link" href="http://arxiv.org/abs/2501.00826v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">LLM-Powered Multi-Agent System for Automated Crypto Portfolio Management</a> employs specialized LLM agents that collaborate for tasks like data analysis and investment decision-making. Fine-tuned with historical data and enhanced with unique collaboration mechanisms, the framework outperformed single-agent models in classification, asset pricing, portfolio management, and explainability.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Optimizing Home Energy Management</b></p><p class="paragraph" style="text-align:left;">Simplifying home energy management systems (HEMS) for non-technical users is important. <a class="link" href="http://arxiv.org/abs/2501.07919v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Large Language Model Interface for Home Energy Management Systems</a> presents an LLM-based interface that transforms poorly formatted user input into well-structured parameters. Using the Reason and Act method and few-shot prompting, the system achieved an average parameter retrieval accuracy of 88%, enhancing user interaction with power systems.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Power Grid Optimization with SafePowerGraph-LLM</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f6922e23-26c5-4803-a5a6-df88f43a3055/image.png?t=1737267253"/></div><p class="paragraph" style="text-align:left;">Modern power grids face increasing complexity. <a class="link" href="http://arxiv.org/abs/2501.07639v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">SafePowerGraph-LLM: Novel Power Grid Graph Embedding and Optimization with Large Language Models</a> introduces a framework that leverages LLMs combined with graph and tabular representations to solve Optimal Power Flow problems. Utilizing in-context learning and fine-tuning protocols, the framework reliably handled realistic grid components and constraints, demonstrating the impact of LLM architecture on performance.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Improving Time Series Forecasting with Pre-trained LLMs</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.06386v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow"><b>Using Pre-trained LLMs for Multivariate Time Series Forecasting</b></a> leverages the capabilities of pre-trained LLMs by mapping multivariate inputs into the LLM&#39;s token embedding space. Utilizing a novel multivariate patching strategy, the approach embeds features into decoder-only pre-trained transformers. The results are competitive with state-of-the-art forecasting models, showcasing the potential of LLMs in complex data analysis tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Innovative FPGA Architecture for Efficient Computing</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/29db4cbd-e3ad-4996-be8b-00198fd77d6f/image.png?t=1737267380"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.06921v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow"><b>Monolithic 3D FPGAs Utilizing Back-End-of-Line Configuration Memories</b></a> introduces a novel FPGA architecture that significantly improves area, latency, and power efficiency. By using back-end-of-line (BEOL) transistors for configuration memory and pass gates, and integrating n-type and p-type amorphous oxide semiconductors (AOS) under development, the researchers developed physics-based models interfaced with Verilog-to-Routing tools. Demonstrated at 7 nm technology, the proposed M3D FPGA design reduces the area-time squared product by 3.4x and decreases critical path latency by 27%. This advancement has broad implications for industries relying on efficient computing, including the deployment of LLMs.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Robotics/VLLMs application/improvement</h2><p class="paragraph" style="text-align:left;"><b>Advancing 3D Scene Understanding with 3UR-LLM</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/13199a92-4535-48fd-ba9a-7a92fd7a7656/image.png?t=1737267446"/></div><p class="paragraph" style="text-align:justify;"><a class="link" href="http://arxiv.org/abs/2501.07819v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding</a> addresses this by building the 3DS-160K dataset and introducing a model that processes 3D point cloud data. Utilizing a 3D compressor module, the model efficiently handles computation demands, outperforming previous state-of-the-art models by 7.1% on relevant benchmarks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>VLMs as Operator Agents in Space</b></p><p class="paragraph" style="text-align:justify;"><a class="link" href="http://arxiv.org/abs/2501.07802v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Visual Language Models as Operator Agents in the Space Domain</a> explores the application of Vision-Language Models in autonomous control and decision-making for space missions. In both software simulations and hardware contexts, VLMs effectively processed visual inputs to perform orbital maneuvers and inspect space objects. These models compete with or outperform traditional methods, showing promise for future space exploration.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Compositional Text-to-Video Generation with BlobGEN-Vid</b></p><p class="paragraph" style="text-align:left;">Generating videos that accurately reflect complex text prompts is a very challenging task. <a class="link" href="http://arxiv.org/abs/2501.07647v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations</a> introduces a model that uses blob-grounded video diffusion to control object motions and appearances. With enhanced regional consistency and semantic control, BlobGEN-Vid demonstrates superior zero-shot video generation capabilities, surpassing proprietary generators in compositional accuracy.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b>Affective Tactile Interaction Driven by LLMs</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d15c29c1-fb07-45be-806a-98ac7ebcea60/image.png?t=1737267624"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2501.07224v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=exploring-new-llm-applications-insights-from-january-1-15-2025-research-papers" target="_blank" rel="noopener noreferrer nofollow">Touched by ChatGPT: Using an LLM to Drive Affective Tactile Interaction</a> leverages LLMs to generate tactile signals that convey emotions. Using a wearable sleeve equipped with vibration motors, unique patterns associated with specific emotions were generated. Participants accurately recognized intended emotions, demonstrating the effectiveness of LLMs in communicating emotion via tactile signals.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=4e57dc96-c798-4f4e-82a5-c5e961af92e2&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published in December 2024</title>
  <description>Discussing Key Innovations and Breakthroughs Transforming Large Language Models (LLMs)</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0d536763-0315-4505-85ef-08f8b4738932/Jnan_13th.png" length="301058" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-in-december-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-in-december-2024</guid>
  <pubDate>Sun, 12 Jan 2025 15:00:00 +0000</pubDate>
  <atom:published>2025-01-12T15:00:00Z</atom:published>
    <dc:creator>LLMs Research</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑 takeaway from today’s newsletter</h2><ul><li><p class="paragraph" style="text-align:left;"><b>Tokens are so yesterday!</b> The Byte Latent Transformer ditches tokens for dynamic byte patches, making models faster and more efficient.</p></li><li><p class="paragraph" style="text-align:left;"><b>Less is more!</b> TrimLLM trims unnecessary layers, boosting speed without sacrificing smarts. It&#39;s like a transformer on a diet!</p></li><li><p class="paragraph" style="text-align:left;"><b>Now you cache it, now you don&#39;t!</b> Slashing KV cache memory usage to just <b>20%</b>, it&#39;s the Houdini of memory optimization.</p></li><li><p class="paragraph" style="text-align:left;"><b>Now you cache it, now you don&#39;t!</b> Slashing KV cache memory usage to just <b>20%</b>, it&#39;s the Houdini of memory optimization.</p></li><li><p class="paragraph" style="text-align:left;"><b>From drone dances to AR cooking!</b> See how LLMs are shaking things up in creative ways you never imagined.</p></li></ul></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:left;"><b>TL;DR - No time to read entire newsletter? </b>Don’t worry We’ve got you covered! Listen <b>🎧 </b>to this fun and informative podcast <span style="color:color(display-p3 0.988 0.992 1 / 0.937);font-family:Inter, system-ui, -apple-system, &quot;Segoe UI&quot;, Roboto, Ubuntu, Cantarell, &quot;Noto Sans&quot;, sans-serif, &quot;Segoe UI&quot;, Roboto, Ubuntu, Cantarell, &quot;Noto Sans&quot;, sans-serif;font-size:14px;">🎙️</span> using Google’s NotebookLM which discuss important research papers from this newsletter in good detail with real world references! </p><div class="recommendation" id="9b0c0c21-d5b5-4ece-8571-ebeff3c62f4c"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/965f16ed-3a5d-45db-9530-9c1d0d143f39/Jnan_13th.png?t=1736660908"/></figure><h3 class="recommendation__title"> Podcast: December&#39;24 research summary </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzliMGMwYzIxLWQ1YjUtNGVjZS04NTcxLWViZWZmM2M2MmY0Yy9EZWNlbWJlciUyMDIwMjQlMjBzdW1tYXJ5Lndhdj9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NlxcdTAwMjZYLUFtei1DcmVkZW50aWFsPUFLSUFRQ01IVFFTRTJKR0FHWEhKJTJGMjAyNjA5MTMlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdFxcdTAwMjZYLUFtei1EYXRlPTIwMjYwOTEzVDExMzM0NlpcXHUwMDI2WC1BbXotRXhwaXJlcz02MDQ4MDBcXHUwMDI2WC1BbXotU2lnbmVkSGVhZGVycz1ob3N0XFx1MDAyNlgtQW16LVNpZ25hdHVyZT03Njk3ZjgwZGIyZDFlNmFmMjZmMjc0ODg5ODQ3NGNjZWNkZDBiYmU2NTU5ZjZhZTJiZTdhMjA5NGRmZGZhY2Y3XCIsXCJ0eXBlXCI6XCJhdWRpby94LXdhdlwiLFwidGh1bWJuYWlsVXJsXCI6XCJodHRwczovL2JlZWhpaXYtaW1hZ2VzLXByb2R1Y3Rpb24uczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Fzc2V0L2ZpbGUvOTY1ZjE2ZWQtM2E1ZC00NWRiLTk1MzAtOWMxZDBkMTQzZjM5L0puYW5fMTN0aC5wbmc_dD0xNzM2NjYwOTA4XCIsXCJ0aXRsZVwiOlwiUG9kY2FzdDogRGVjZW1iZXInMjQgcmVzZWFyY2ggc3VtbWFyeVwifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Architecture & Token Processing</h2><p class="paragraph" style="text-align:left;">Research papers from this section explore fundamental architectural improvements and token processing methods, achieving better scalability and efficiency through novel approaches to model architecture and token handling.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.09871v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Byte Latent Transformer: Patches Scale Better Than Tokens</a> 🔥</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7fd541c-2b98-4fc7-a155-8d680b9b1ce0/image.png?t=1736656219"/></div><p class="paragraph" style="text-align:left;">The <b>Byte Latent Transformer (BLT)</b> operates at the byte level rather than on tokens. Instead of fixed-size tokens, BLT forms dynamically sized <b>byte patches</b> determined by the entropy (unpredictability) of the next byte. High-entropy regions, which contain more information, receive more computational focus.</p><ul><li><p class="paragraph" style="text-align:left;"><b>Dynamic Patching:</b> BLT groups sequences of bytes into variable-length patches based on predictability, efficiently capturing patterns without relying on a predefined vocabulary.</p></li><li><p class="paragraph" style="text-align:left;"><b>Entropy-Based Processing:</b> By allocating computational resources according to the complexity of the data (measured by entropy), BLT ensures that challenging inputs receive more attention.</p></li></ul><p class="paragraph" style="text-align:left;">The researchers conducted a FLOP-controlled scaling study, training BLT models with<b> up to 8 billion parameters on 4 trillion bytes of data</b>. This study assessed the efficiency and performance of byte-level models compared to traditional token-based LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06106v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enhanced Computationally Efficient Long LoRA Inspired Perceiver Architectures for Auto-Regressive Language Modeling</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/179e715b-82a3-4f7c-8a43-cf0eec41bf8f/image.png?t=1736656313"/></div><p class="paragraph" style="text-align:left;">This research builds upon the <b>Perceiver AR</b> architecture, designed to handle longer sequences efficiently. By integrating ideas from <b>Long Low-Rank Adaptation (Long-LoRA)</b>, the researchers develop an enhanced architecture called the <b>Long LoRA Perceiver</b>. This model introduces architectural enhancements that balance computational efficiency and performance. It applies low-rank approximations to the attention matrices, reducing computational complexity from quadratic to linear with respect to sequence length. Additionally, it utilizes hierarchical attention mechanisms to capture both local and global dependencies efficiently, enabling the model to focus on relevant information across different scales within the sequence. By carefully optimizing model depth, width, and other hyperparameters, the Long LoRA Perceiver maximizes performance while keeping computational demands manageable.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.09952v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Llama 3 Meets MoE: Efficient Upcycling</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1dde146d-577b-4e2a-aa8c-30c907fac792/image.png?t=1736656343"/></div><p class="paragraph" style="text-align:left;">The research introduces an efficient method to &quot;upcycle&quot; a pre-trained dense model into an MoE architecture. Specifically, an <b>8-Expert Top-2 MoE</b> model is trained using pre-trained dense checkpoints from Llama 3-8B. This approach leverages existing knowledge in the dense model and integrates MoE layers with minimal additional training. By adding MoE layers where each expert specializes in different data subsets, the model increases its capacity. During inference, only the top two most relevant experts are activated per input token, keeping computational costs similar to the original dense model. The upcycling process requires less than one epoch of training, significantly reducing time and resources compared to training an MoE model from scratch. Initializing the MoE model with pre-trained dense weights helps maintain performance while benefiting from the increased capacity.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The upcycled MoE model achieves a <b>2×</b> increase in model capacity compared to the original Llama 3-8B dense model</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Advertisement</h2><p class="paragraph" style="text-align:left;">Please support my work by checking what this advertise offer. Maybe it helps you!👇</p><h3 class="heading" style="text-align:left;" id="want-to-get-the-most-out-of-chat-gp">Want to get the most out of ChatGPT?</h3><div class="image"><a class="image__link" href="https://offers.hubspot.com/using-chatgpt-at-work?utm_medium=email-media-newsletter&utm_source={{publication_alphanumeric_id}}&utm_campaign=creator&utm_content=beehiiv&utm_term=10-1-2024&_bhiiv=opp_ac431f71-5e08-4265-af98-36c8b4c08567_b942af4d&bhcl_id=0b31e92b-4120-4981-a4bf-f39350e17d2f_{{subscriber_id}}_{{email_address_id}}" rel="noopener" target="_blank"><img class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b923243c-67a6-4d20-9aca-e21973b24866/1.jpg?t=1740768834"/></a></div><p class="paragraph" style="text-align:left;">ChatGPT is a superpower if you know how to use it correctly.</p><p class="paragraph" style="text-align:left;">Discover how <a class="link" href="https://offers.hubspot.com/using-chatgpt-at-work?utm_medium=email-media-newsletter&utm_source={{publication_alphanumeric_id}}&utm_campaign=creator&utm_content=beehiiv&utm_term=10-1-2024&_bhiiv=opp_ac431f71-5e08-4265-af98-36c8b4c08567_b942af4d&bhcl_id=0b31e92b-4120-4981-a4bf-f39350e17d2f_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">HubSpot&#39;s guide to AI</a> can elevate both your productivity and creativity to get more things done.</p><p class="paragraph" style="text-align:left;">Learn to automate tasks, enhance decision-making, and foster innovation with the power of AI.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://offers.hubspot.com/using-chatgpt-at-work?utm_medium=email-media-newsletter&utm_source={{publication_alphanumeric_id}}&utm_campaign=creator&utm_content=beehiiv&utm_term=10-1-2024&_bhiiv=opp_ac431f71-5e08-4265-af98-36c8b4c08567_b942af4d&bhcl_id=0b31e92b-4120-4981-a4bf-f39350e17d2f_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Download the free guide with actionable prompts and tips to make your work more productive.</a></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Model Efficiency & Compression</h2><p class="paragraph" style="text-align:justify;">Recent advancements have introduced innovative techniques to compress and accelerate LLMs. current methods such as quantization and layer pruning achieve substantial speedups (2× to 5×) and memory savings while maintaining, or even enhancing, model performance. This section covers papers improving the LLMs efficiency and their compression. </p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.11242v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/288f2b05-d62d-469d-a2a2-c66a931df7bf/image.png?t=1736573099"/></div><p class="paragraph" style="text-align:left;">TrimLLM introduces a progressive layer dropping technique inspired by the layer-wise specialization phenomenon observed in LLMs. The method begins with a thorough analysis of each transformer&#39;s layer contribution to the model&#39;s performance on domain-specific tasks. By measuring changes in loss or accuracy when a layer is temporarily removed, less critical layers are identified.</p><p class="paragraph" style="text-align:left;">The progressive layer dropping involves iteratively removing these least important layers. After each layer is dropped, the model is evaluated on validation data to ensure performance remains within acceptable bounds. Domain-specific fine-tuning is applied after each pruning step, allowing the model to adjust to its new, slimmer architecture and mitigate any performance loss. Also, this methodology does not require specialized hardware, making it accessible for a wide range of users.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> TrimLLM achieved 2.1-5.7× inference speedup on consumer GPUs and up to 3.1× on A100 GPUs compared to state-of-the-art compression methods, while maintaining no loss in accuracy at 50∼60% model compression ratio.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.07902v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Low-Rank Correction for Quantized LLMs</a> 🔥</p><p class="paragraph" style="text-align:left;"><b>Microsoft</b> research team introduced a method to mitigate quantization errors by introducing low-rank weight matrices in full precision that act on unquantized activations. The approach augments the quantized weight matrix with a low-rank correction matrix, resulting in a corrected weight matrix that combines both. A joint optimization objective is defined to minimize the discrepancy between the outputs of the corrected quantized model and the original full-precision model.</p><p class="paragraph" style="text-align:left;">During fine-tuning, quantization-aware training is performed where both the quantized weights and correction matrices are updated. The quantized weights are adjusted using quantized gradients, while the correction matrices are updated using full-precision gradients, effectively compensating for the precision loss in aggressive quantization. By correcting the activations through these low-rank matrices, the method restores model accuracy without significantly increasing model size, as the correction matrices require minimal additional storage.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Proposed approach reduced the accuracy gap by <b>over 50%</b> compared to quantized models without correction.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06419v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LLM-BIP: Structured Pruning for Large Language Models with Block-Wise Forward Importance Propagation</a></p><p class="paragraph" style="text-align:left;">LLM-BIP introduces a structured pruning technique that calculates the importance of parameters using block-wise forward importance propagation. The model is divided into blocks, such as attention heads or MLP layers. The method simulates the impact of pruning each block by approximating its effect on the model&#39;s output using a forward pass and leveraging Lipschitz continuity to ensure stability.</p><p class="paragraph" style="text-align:left;">By assessing how changes in a block affect the loss function, the importance of different structures is accurately evaluated. Blocks are ranked based on these importance scores, and the least important ones are pruned while ensuring the model remains functional. After pruning, fine-tuning is conducted to recover any potential performance degradation, using learning rate schedules that stabilize training.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The method increases accuracy by 3.26% compared to previous pruning methods at similar compression rates.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.00648v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation</a></p><p class="paragraph" style="text-align:left;">DFRot reduces outliers in LLM quantization by refining weight matrices through rotation to produce outlier-free and evenly distributed activation values. The method applies randomized Hadamard transforms to the weight matrices, spreading out large values and reducing the chance of large activations. A weighted loss function is introduced to penalize activations exceeding a certain threshold, effectively handling long-tail distributions.</p><p class="paragraph" style="text-align:left;">The rotation matrices are further refined using orthogonal Procrustes optimization, ensuring the transformed weights remain close to the original weights while achieving desired activation properties. This results in a more uniform activation distribution, facilitating better quantization without sacrificing precision for other values.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> DFRot improved perplexity by 0.25 on W4A4KV4 and 0.21 on W4A4KV16, demonstrating effectiveness in reducing outliers and massive activations in quantized LLaMA3-8B.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.09250v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuning</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c2d0497b-be4f-4de9-8c2d-188a82e7d8a2/image.png?t=1736654878"/></div><p class="paragraph" style="text-align:left;">GeLoRA introduces a geometry-inspired approach that dynamically adjusts the ranks of Low-Rank Adaptation (LoRA) matrices based on the intrinsic dimensionality of hidden state representations in each layer. By computing the singular values of hidden representations, the method estimates the intrinsic dimensionality, with layers processing more complex information receiving higher ranks.</p><p class="paragraph" style="text-align:left;">Analyzing the decay of singular values allows GeLoRA to adjust ranks to match each layer&#39;s complexity, ensuring efficient parameter allocation. This leverages geometric principles to determine a theoretical lower bound for optimal rank selection, balancing efficiency and expressiveness.</p><p class="paragraph" style="text-align:left;">During fine-tuning, only the LoRA parameters are updated while the base model remains frozen. This parameter-efficient strategy enables task-specific adaptation without requiring full model training, making it suitable for environments with limited resources.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:left;">Memory Management & KV Cache Optimization</h2><p class="paragraph" style="text-align:left;">The Key-Value (KV) cache is very important for efficient transformer model inference, storing past hidden states to avoid redundant computations. However, the KV cache can consume significant memory resources, leading to scalability issues and high computational costs.</p><p class="paragraph" style="text-align:left;">Recent research has focused on optimizing KV cache memory usage without compromising model performance. By developing sophisticated compression and pruning techniques, researchers aim to reduce memory footprints by up to 80-90%, enabling LLMs to handle longer sequences and be deployable in resource-constrained environments.</p><p class="paragraph" style="text-align:left;">In this section, we explore innovative approaches that address memory efficiency in LLMs, particularly focusing on KV cache optimization through dynamic compression, sparse coding, adaptive strategies, and novel pruning methods.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.09036v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty</a></p><p class="paragraph" style="text-align:left;">ZigZagkv introduces a dynamic KV cache compression method that allocates memory budgets per layer based on layer uncertainty. The key idea is to assess the necessity of memory allocation for each layer individually, rather than applying a uniform compression strategy across all layers. The method estimates uncertainty scores for each layer by analyzing attention patterns and the variability of hidden state outputs. Layers with higher uncertainty—indicating greater importance for accurate predictions—are assigned larger memory budgets to retain more precise KV representations. Conversely, layers with lower uncertainty have their KV caches compressed to save memory without significantly impacting performance.</p><p class="paragraph" style="text-align:left;">By dynamically adjusting the compression ratio for the KV cache in each layer, the model retains essential information in critical layers while substantially reducing overall memory usage. This approach ensures efficient memory utilization, enabling LLMs to handle longer contexts without exceeding memory limitations.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The proposed method reduces KV cache memory usage to approximately <b>20%</b> of the original size during inference, significantly alleviating memory bottlenecks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.08890v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/23ad707b-3261-4c5b-a3f9-261cdcddb39a/image.png?t=1736655579"/></div><p class="paragraph" style="text-align:left;">Lexico uses a sparse coding approach using a universal dictionary to approximate the KV caches with sparse linear combinations. The universal dictionary, consisting of approximately <b>4,000</b> atoms (basis vectors), captures essential patterns in KV caches across various inputs, tasks, and model architectures. By representing each key and value vector as a sparse combination of these atoms, the method compresses the KV cache significantly.</p><p class="paragraph" style="text-align:left;">Orthogonal Matching Pursuit (OMP) is utilized to find the optimal sparse representation for each vector, ensuring accurate reconstruction with minimal coefficients. The sparsity level can be adjusted to control the compression ratio, allowing flexible trade-offs between memory savings and performance. This approach maintains the model&#39;s ability to reconstruct necessary information during inference, achieving extreme compression without significant loss.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Lexico maintains <b>90-95%</b> of the original model performance despite significant compression. </p><hr class="content_break"><p class="paragraph" style="text-align:left;"><br><a class="link" href="http://arxiv.org/abs/2412.08521v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">EMS: Adaptive Evict-then-Merge Strategy for Head-wise KV Cache Compression Based on Global-Local Importance</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f2d16523-1b74-49b8-8da9-d11fe41e50a3/image.png?t=1736655626"/></div><p class="paragraph" style="text-align:left;">EMS introduces an adaptive evict-then-merge strategy for KV cache compression by calculating a Global-Local Importance (GLI) score for each token. The GLI score integrates global accumulated attention scores—reflecting how much a token is attended to throughout the entire sequence—with local attention scores that capture importance within specific contexts or attention heads. This combined score provides a more accurate estimation of token significance.</p><p class="paragraph" style="text-align:left;">Using the GLI scores, less important tokens are first evicted to reduce the KV cache size. Then, redundant or similar tokens are merged by aggregating their representations, further saving memory. This process is applied on a per-head basis, recognizing that different attention heads may focus on different aspects of the input. A zero-class mechanism is introduced to handle evicted tokens consistently during inference without adverse effects.</p><p class="paragraph" style="text-align:left;">By adapting the compression strategy based on the GLI scores and processing it head-wise, EMS effectively balances memory savings with performance preservation, allowing efficient handling of long-context inputs in LLMs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> EMS achieves state-of-the-art performance with low perplexity, improving metrics by over <b>1.28 points</b> on four LLMs evaluated on the LongBench benchmark under a <b>256 token</b> cache budget.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.04652v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Cross-Self KV Cache Pruning for Efficient Vision-Language Inference</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1ee83383-3f82-42a0-93ec-8cfcb5367f29/image.png?t=1736655668"/></div><p class="paragraph" style="text-align:left;">The proposed method decomposes attention scores into intra-modality (within the same modality) and inter-modality (across modalities) components. By separately analyzing these components, the model can more precisely estimate the importance of tokens in both visual and textual modalities.</p><p class="paragraph" style="text-align:left;">Intra-modality attention captures relationships among tokens within the same modality, helping to identify essential visual features or textual elements. Inter-modality attention examines interactions between modalities, ensuring that tokens critical for cross-modal understanding are retained. Additionally, an n-softmax function is introduced to maintain the smoothness of attention scores after pruning, adjusting attention distributions to account for the pruned tokens and preventing abrupt changes that could destabilize performance.</p><p class="paragraph" style="text-align:left;">By employing this tailored pruning approach, the model effectively reduces the KV cache size without losing important information from either modality, leading to more efficient inference in vision-language tasks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The Cross-Self Pruning (CSP) method matches the performance of models with full KV caches and outperforms previous pruning methods, achieving up to a 41% memory reduction.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06263v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models</a></p><p class="paragraph" style="text-align:left;">iLLaVA introduces a method to optimize LVLMs by merging redundant tokens in image representations. Utilizing a precise and rapid algorithm, the method identifies tokens within image inputs that carry redundant or less significant information. These redundant tokens are merged, effectively reducing the number of tokens required to represent an image.</p><p class="paragraph" style="text-align:left;">To ensure important visual information is not lost, the method recycles information from the removed tokens by integrating their contributions into the remaining tokens. This process preserves essential features and relationships necessary for the model&#39;s performance. By reducing the token count for images, iLLaVA decreases memory usage and doubles the inference throughput, as the model has fewer tokens to process. The approach achieves these efficiency gains with minimal modifications to the model architecture, facilitating easy integration into existing systems.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> iLLaVA doubles throughput and halves memory consumption with only a 0.2% drop in performance, demonstrating efficiency gains without sacrificing accuracy.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Training & Optimization Methods</h2><p class="paragraph" style="text-align:left;">These papers improve LLM training efficiency and effectiveness through enhanced optimization techniques, distributed training methods, and novel adaptation approaches, achieving better convergence and reduced computational requirements.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.07210v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7276e4df-aeef-4567-9186-e7541649ffbc/image.png?t=1736655713"/></div><p class="paragraph" style="text-align:left;">EDiT introduces an efficient distributed training framework that integrates a specific variant of Local Stochastic Gradient Descent (Local SGD) with model sharding. By allowing workers to perform multiple local updates before synchronizing, the method reduces communication frequency and overhead. The key innovation lies in using layer-wise synchronization during the forward pass. This approach overlaps computation and communication, minimizing idle time and enhancing efficiency.</p><p class="paragraph" style="text-align:left;">To stabilize training, EDiT introduces a pseudo-gradient penalty that acts as a regularizer, mitigating the divergence that can occur with delayed updates in Local SGD. Furthermore, recognizing the heterogeneity of cluster environments, the researchers present A-EDiT, an asynchronous version of the method. A-EDiT accommodates varied computational resources and network conditions by allowing workers to operate without strict synchronization, further improving scalability and robustness in diverse training environments.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.05270v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">APOLLO: SGD-like Memory, AdamW-level Performance</a></p><p class="paragraph" style="text-align:left;">APOLLO addresses this challenge by proposing an optimizer that retains the memory efficiency of SGD while achieving performance comparable to AdamW. The key idea is to approximate the per-parameter adaptive learning rates of AdamW through a structured update rule that relies on an auxiliary low-rank optimizer state. By leveraging random projection techniques, APOLLO maintains a low-rank approximation of the gradients, capturing essential variance information without the need to store full-size moment estimates.</p><p class="paragraph" style="text-align:left;">This approach reduces memory usage significantly since the auxiliary optimizer state requires far less storage than the full moment vectors. APOLLO adjusts learning rates adaptively based on the low-rank approximation, effectively emulating the benefits of AdamW&#39;s adaptive updates. Additionally, the researchers introduce APOLLO-Mini, a variant designed for even lower memory environments, making it suitable for training on devices with limited resources.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> APOLLO and its variant, APOLLO-Mini, deliver competitive or superior performance compared to AdamW, with significant memory savings: 3x throughput and 4x larger batch sizes on an 8xA100-80GB setup, enabling model scalability and low-end GPU training.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06071v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">KaSA: Knowledge-Aware Singular-Value Adaptation of Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/76733cf8-3aa6-48ef-8077-69129f31df29/image.png?t=1736655907"/></div><p class="paragraph" style="text-align:left;">KaSA introduces a novel PEFT method that employs Singular Value Decomposition (SVD) with knowledge-aware singular values to adapt LLMs more efficiently. The method decomposes weight matrices into singular vectors and singular values, allowing fine-grained control over the model&#39;s parameter adaptation. By analyzing task relevance, KaSA dynamically activates singular values corresponding to components of the model that are most pertinent to the target task.</p><p class="paragraph" style="text-align:left;">This knowledge-aware activation reduces noise by de-emphasizing irrelevant parameters, effectively focusing the model&#39;s capacity on task-specific knowledge. The approach improves expressiveness without significantly increasing the computational burden. By updating only the most relevant singular values, KaSA minimizes memory overhead and accelerates the fine-tuning process.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> KaSA outperforms 14 popular PEFT baselines and FFT across 16 benchmarks and 4 synthetic datasets, demonstrating its efficacy in tasks such as NLU, NLG, and commonsense reasoning.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06858v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization</a></p><p class="paragraph" style="text-align:left;">Paper introduces Noise Perturbation Fine-Tuning (NPFT), a method designed to diminish the impact of sensitive weights during quantization. NPFT involves applying random weight perturbations during parameter-efficient fine-tuning (PEFT). By adding controlled noise to the weights identified as sensitive (those with high influence on the loss Hessian trace), the method encourages the model to distribute importance more evenly across parameters.</p><p class="paragraph" style="text-align:left;">This process reduces the reliance on outlier weights, effectively smoothing the loss landscape and enhancing the model&#39;s robustness to quantization errors. NPFT operates without requiring special treatment for outliers or altering the model architecture, making it a straightforward addition to existing fine-tuning procedures. By integrating NPFT with PEFT, the approach maintains parameter efficiency while improving quantized model performance.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> NPFT enhanced the performance of OPT and LLaMA models across uniform and non-uniform quantizers. On the LLaMA2-7B-4bits benchmark, RTN matched the performance of GPTQ, achieving better inference efficiency.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Knowledge Management & Reasoning</h2><p class="paragraph" style="text-align:left;">This section focuses on research that helps LLMs to acquire, maintain, and utilize knowledge through innovative approaches such as memory compression, knowledge graph completion, interpretability techniques, and novel prompt optimization methods. These advancements not only improve the models reasoning capabilities but also address challenges like continual learning, interpretability, and efficient adaptation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.07393v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be19d2bb-44a6-40e6-8374-14471e3dcb66/image.png?t=1736656026"/></div><p class="paragraph" style="text-align:left;">The CMT method emulates human memory by compressing and storing information from new documents into an external memory bank. Instead of updating the entire model, CMT retrieves relevant compressed memories during query processing. It introduces three key techniques to enhance performance:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Memory-Aware Objective:</b> This encourages effective utilization of the compressed memory by adjusting the training objective, balancing reliance on stored knowledge versus internal parameters.</p></li><li><p class="paragraph" style="text-align:left;"><b>Self-Matching:</b> It improves the quality of stored memories by matching and integrating similar information within the memory bank, enhancing representational coherence.</p></li><li><p class="paragraph" style="text-align:left;"><b>Top-Aggregation:</b> During inference, it retrieves and aggregates the most relevant memories, allowing the model to access and reason over pertinent information without altering its original parameters.</p></li></ul><p class="paragraph" style="text-align:left;">By leveraging these techniques, CMT reduces catastrophic forgetting—the loss of previously learned knowledge when new information is added—allowing the model to incorporate new knowledge seamlessly.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The CMT method demonstrates improved adaptability and robustness, achieving enhancements like +4.07 EM and +4.19 F1 on the StreamingQA dataset with Llama-2-7b.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.11016v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">A Contextualized BERT model for Knowledge Graph Completion</a></p><p class="paragraph" style="text-align:left;">The study introduces a contextualized BERT model that leverages contextual information from neighboring entities and relationships within the knowledge graph. Unlike traditional KGC approaches that rely heavily on entity descriptions or require negative triplet sampling, this model incorporates structural context directly into the embeddings. By encoding the neighborhood information of entities using BERT&#39;s transformer architecture, the model captures both local and global semantics, enabling it to infer missing links more accurately. This approach reduces computational demands and mitigates semantic inconsistencies by eliminating the need for negative sampling.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The model outperforms existing methods, achieving a <b>5.3%</b> improvement in standard evaluation metrics such as Mean Reciprocal Rank (MRR) and Hits@N.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.07334v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation</a></p><p class="paragraph" style="text-align:left;">Researchers propose the Frame Representation Hypothesis, modeling multi-token words as frames—a sequence of vectors representing each token in the word. Analyzing these frames uncovers how LLMs internally represent complex concepts composed of multiple tokens. By interpreting concepts using the average of word frames that share the same concept, the methodology facilitates manipulation of these concepts within the model&#39;s latent space. This enables detection and modification of biases, as well as concept-guided text generation where specific concepts can be emphasized or de-emphasized in generated text. The approach was validated on models such as Llama 3.1, Gemma 2, and Phi 3, demonstrating its ability to identify biases related to gender and language and to remediate them through frame manipulation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.09722v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0bdcb5f3-271d-4061-a17f-a53625223ee3/image.png?t=1736656127"/></div><p class="paragraph" style="text-align:left;">GReaTer introduces a novel approach that utilizes task loss gradients for prompt optimization in smaller language models. Instead of relying solely on textual feedback or guidance from larger models, GReaTer allows the smaller model to introspect its own performance by analyzing gradients computed during task-specific losses. By leveraging these gradients, the model adjusts its prompts to improve performance iteratively. This self-improvement mechanism enables the model to enhance its reasoning capabilities without external assistance or significant computational overhead, making it practical for deployment in resource-constrained environments.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;"><b>Creative ways to use LLMs!!</b></h2><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.11043v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Semantic Steganography: A Framework for Robust and High-Capacity Information Hiding using Large Language Models</a> - Paper<b> </b>constructs a semantic space to map secret messages using ontology-entity trees, enhancing robustness, transmission reliability, and indistinguishability of stego texts.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.12201v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Embracing Large Language Models in Traffic Flow Forecasting</a> - Paper uses LLMs to improve traffic flow prediction by capturing complex spatio-temporal relationships through graph and hyper graph structures, advancing intelligent transportation systems.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.10872v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">IntelEX: A LLM-driven Attack-level Threat Intelligence Extraction Framework</a> - Paper uses LLMs to<b> </b>convert unstructured cyber threat intelligence reports into structured formats, facilitating better security analysis and response planning.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.10582v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models</a> [Marvel Fans check this paper!] 🔥 - Paper applies<b> </b>zero-shot meta-prompting with GPT-4 to generate branching narratives from a prewritten story. Starting with a linear plot, it creates branches at key decision points by prompting the LLM to consider major plot points, ensuring coherent alternate storylines. The branching plot is stored in a graph for tracking and structuring the interactive fiction system.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.08428v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">SwarmGPT-Primitive: A Language-Driven Choreographer for Drone Swarms Using Safe Motion Primitive Composition</a> - Paper achieves smooth drone swarm choreographies using natural language commands powered by LLMs, integrating motion planning and safety constraints.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.06294v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects</a> - Paper Creates an agent named Installamatic that autonomously installs Python project dependencies by interpreting and following documentation, improving software installation processes.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.05393v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">HiVeGen -- Hierarchical LLM-based Verilog Generation for Scalable Chip Design</a> - Paper uses hierarchical LLMs to generate complex hardware designs in Verilog by structuring tasks into smaller submodules, enhancing scalability and reducing errors in chip design.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.04806v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Rethinking Time Series Forecasting with LLMs via Nearest Neighbor Contrastive Learning</a> - Paper adapts LLMs for time series forecasting using Nearest Neighbor Contrastive Learning to align textual embeddings with time series data, excelling in few-shot and long-term predictions.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/95e0baa7-1dda-407f-acd8-2eefa3728541/image.png?t=1736657841"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2412.00627v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-december-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">ARChef: An iOS-Based Augmented Reality Cooking Assistant Powered by Multimodal Gemini LLM</a> - Paper combines LLMs with augmented reality to identify ingredients and generate personalized recipes through an iOS application, enhancing the cooking experience and accessibility.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6462a294-6acf-4a84-8318-10b252e71a52/image.png?t=1736657794"/></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=68cea5d3-8417-4238-9d42-ae208f2e127c&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published in November 2024</title>
  <description>Discussing Key Innovations and Breakthroughs Transforming Large Language Models (LLMs)</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dfe8d7f8-f2e6-49e2-9dc8-501d18191176/nov_newsletter.png" length="498268" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-in-november-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-in-november-2024</guid>
  <pubDate>Fri, 27 Dec 2024 15:21:44 +0000</pubDate>
  <atom:published>2024-12-27T15:21:44Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ul><li><p class="paragraph" style="text-align:left;"><b>Smarter Thinking</b>: Fixing logic gaps with <b>Critical Tokens</b> and tackling multi-hop reasoning challenges.</p></li><li><p class="paragraph" style="text-align:left;"><b>Efficient Fine-Tuning</b>: Innovations like <b>LoRA-SB</b> cut costs without compromising performance.</p></li><li><p class="paragraph" style="text-align:left;"><b>Compact Models</b>: Faster, lighter LLMs with breakthroughs like <b>FlexiBit</b> and <b>MixPE</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>Creative Applications</b>: Endless panoramas, AI storytelling, and dynamic simulations powered by LLMs.</p></li><li><p class="paragraph" style="text-align:left;"><b>Sharper Understanding</b>: Syntax tools and self-distillation improve accuracy and versatility.</p></li></ul></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> Fun & engaging podcast using NotebookLM</b></h2><p class="paragraph" style="text-align:left;">Don’t have much time to read entire newsletter? Well, listen to this fun and engaging podcast covering these research papers in a detailed manner!</p><div class="recommendation" id="4678ce19-d492-4a88-ace1-a9cf383ca3de"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/dfe8d7f8-f2e6-49e2-9dc8-501d18191176/nov_newsletter.png?t=1735275248"/></figure><h3 class="recommendation__title"> November 2024 research summary </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOlwiI0ZGRkZGRlwiLFwiYmFja2dyb3VuZFRoZW1lXCI6XCJsaWdodFwiLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzQ2NzhjZTE5LWQ0OTItNGE4OC1hY2UxLWE5Y2YzODNjYTNkZS9Ob3ZlbWJlciUyMDIwMjQlMjByZXNlYXJjaCUyMHN1bW1hcnkud2F2P1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2XFx1MDAyNlgtQW16LUNyZWRlbnRpYWw9QUtJQVFDTUhUUVNFMkpHQUdYSEolMkYyMDI2MDkxMyUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0XFx1MDAyNlgtQW16LURhdGU9MjAyNjA5MTNUMTEzMzQ3WlxcdTAwMjZYLUFtei1FeHBpcmVzPTYwNDgwMFxcdTAwMjZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3RcXHUwMDI2WC1BbXotU2lnbmF0dXJlPWQ2ZWEzNmMwYTdkY2FkZmIwNmIxZTMzYWI5M2Y3MDMzNmZiYzIwMmQ2ZTAyNThjZThiZDIxOTQ2ODgzNDI1MzVcIixcInR5cGVcIjpcImF1ZGlvL3gtd2F2XCIsXCJ0aHVtYm5haWxVcmxcIjpcImh0dHBzOi8vYmVlaGlpdi1pbWFnZXMtcHJvZHVjdGlvbi5zMy5hbWF6b25hd3MuY29tL3VwbG9hZHMvYXNzZXQvZmlsZS9kZmU4ZDdmOC1mMmU2LTQ5ZTItOWRjOC01MDFkMTgxOTExNzYvbm92X25ld3NsZXR0ZXIucG5nP3Q9MTczNTI3NTI0OFwiLFwidGl0bGVcIjpcIk5vdmVtYmVyIDIwMjQgcmVzZWFyY2ggc3VtbWFyeVwifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Token-Level Reasoning and Syntax</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19943v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM&#39;s Reasoning Capability</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f19bf837-b716-4f37-8003-ccf5f1d5360c/image.png?t=1735235952"/></div><p class="paragraph" style="text-align:left;"><b>Why?: </b>Discriminating influential &#39;critical tokens&#39; which leads to incorrect reasoning paths can help improve LLMs reasoning ability.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors propose a method called <b>cDPO</b> (Contrastive Direct Preference Optimization) to identify and correct these troublesome tokens. The method is built on the insight that <b>not all tokens are equally important</b>; some can drastically shift the model’s reasoning trajectory.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Positive vs. Negative Reasoning Trajectories</b></p><ul><li><p class="paragraph" style="text-align:left;">They collect <b>positive</b> (correct) reasoning paths and <b>negative</b> (incorrect) ones, often from a chain-of-thought or step-by-step solution. The key is to figure out which tokens appeared in correct vs. incorrect solutions, and how likely the LLM is to generate those tokens under different conditions.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Contrasting Likelihoods</b></p><ul><li><p class="paragraph" style="text-align:left;">They train two separate LLM “heads” or fine-tuned variants: one favoring <b>positive</b> tokens, the other favoring <b>negative</b> tokens. This allows them to <b>contrast</b> how likely the model is to produce certain “critical tokens” in each scenario. If a token is much more probable in the negative model, it might be a strong predictor of failure.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Token-Level DPO</b></p><ul><li><p class="paragraph" style="text-align:left;"><b>Direct Preference Optimization (DPO)</b> is usually done at the sequence level (i.e., preferring one entire output over another). The authors adapt DPO <b>at the token level</b>—pinpointing precisely where the model’s reasoning diverges. This token-level alignment is the core innovation: rather than treating a whole reasoning chain as “good or bad,” they surgically intervene on the problematic tokens.</p></li></ul></li></ol><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.18885v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Sneaking Syntax into Transformer Language Models with Tree Regularization</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/00b232e0-0a81-43ac-92a8-0563cad1adb6/image.png?t=1735249540"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper addresses the need for syntactic inductive biases in transformer language models to improve their robustness and data efficiency, without limiting model expressivity or increasing inference complexity.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors propose <b>TreeReg</b>, a technique that <b>blends syntactic constraints into a Transformer’s hidden states</b>without altering the core model architecture. Rather than fully retraining a parser, they apply a carefully designed <b>auxiliary loss</b> that leverages <b>syntactic bracketing</b> (e.g., parse trees).</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Syntactic Brackets</b></p><ul><li><p class="paragraph" style="text-align:left;">Traditional syntactic parsers produce bracketed structures indicating phrases, clauses, etc.</p></li><li><p class="paragraph" style="text-align:left;">These brackets encode where one syntactic constituent ends and another begins.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Orthogonality Constraints</b></p><ul><li><p class="paragraph" style="text-align:left;">TreeReg translates bracket information into <b>differentiable constraints</b> on hidden states: if two tokens are in the same syntactic constituent, their vector representations should align differently than tokens in different constituents.</p></li><li><p class="paragraph" style="text-align:left;">By encouraging orthogonality (or closeness) based on bracket positions, the model <b>“feels”</b> the syntactic boundaries directly in its feature space.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>No Architecture Overhaul</b></p><ul><li><p class="paragraph" style="text-align:left;">Crucially, <b>no new parameters</b> or special modules are added. Instead, the model gains a <b>syntax-aware training signal</b> through a supplementary loss that can be combined with standard next-token prediction or classification losses.</p></li></ul></li></ol><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.17679v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Existing tokenization methods, such as Byte-Pair Encoding (BPE), which often obscure the internal character structures within tokens, hindering LLMs&#39; ability to grasp these details and perform effectively on tasks with limited data.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper introduces a method called Token Internal Position Awareness (TIPA), which trains LLMs on reverse character prediction tasks using the tokenizer&#39;s vocabulary. This approach helps models learn internal token structures and character positions, enhancing their understanding.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> LLMs trained with TIPA outperform baseline models in predicting character positions at the token level and show improved performance and faster convergence on the downstream task of Chinese Spelling Correction (CSC).</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.16353v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">The Two-Hop Curse: LLMs trained on A-&gt;B, B-&gt;C fail to learn A--&gt;C</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding the limitations of LLMs in performing internal reasoning without explicit chain-of-thought helps to enhance their reasoning capabilities, which is essential for advancing LLMs applications.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper introduces a controlled experimental setting to assess two-hop reasoning in LLMs by training models on fictional facts and testing them on their ability to perform reasoning tasks both with and without chain-of-thought (CoT) aid. The models, including Llama 3 8B Instruct and GPT-4o, were evaluated on their performance in generalizing and composing learned facts across different documents and prompts.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Models succeeded at two-hop reasoning using CoT but failed when reasoning required internal inference without CoT. The failure was evident when learned facts were in separate documents, highlighting LLMs inability for latent multi-hop reasoning without external aids in over half of the cases.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Efficient and Effective Fine-Tuning</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2411.19557?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning</a> - <a class="link" href="https://github.com/RaghavSinghal10/lora-sb?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/217c8209-f795-437f-b01c-b730f095f608/image.png?t=1735250750"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Fine-tuning LLMs often involves updating <b>billions of parameters</b>, which is computationally expensive and requires high-end hardware. <b>LoRA</b> can drastically reduce the number of trainable parameters by decomposing the weight updates into low-rank matrices. However, standard LoRA approaches often lag behind full fine-tuning in terms of performance and may demand extensive hyperparameter tuning to close that gap.</p><p class="paragraph" style="text-align:justify;"><b>How?:</b> The authors observe that while LoRA significantly cuts parameter counts, <b>its performance can be very sensitive</b> to how the low-rank matrices are initialized and scaled. Traditional approaches typically rely on random or naive initializations, requiring lengthy hyperparameter sweeps.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Gradient-Based Approximation</b></p><p class="paragraph" style="text-align:left;">LoRA-SB approximates the initial <b>gradient</b> of a full fine-tuning step to seed the low-rank matrices. By focusing on the directions in parameter space that matter most for the task, it effectively “zeroes in” on the relevant updates right from the start.</p></li><li><p class="paragraph" style="text-align:left;"><b>Constrained Update Space</b></p><p class="paragraph" style="text-align:left;">The low-rank adapters operate in a restricted subspace compared to full model updates. Finding the <b>best possible initialization</b> in that subspace is critical because every parameter or gradient dimension counts more when you have fewer degrees of freedom.</p></li><li><p class="paragraph" style="text-align:left;"><b>No Extra Hyperparameters</b></p><p class="paragraph" style="text-align:left;">A major advantage is that LoRA-SB <b>does not</b> introduce additional hyper parameters. It’s a straightforward re-initialization scheme that can be easily integrated into existing LoRA pipelines.</p></li></ol><p class="paragraph" style="text-align:left;"><b>Results:</b> LoRA-SB outperforms standard LoRA and LoRA-XS, achieving efficient fine-tuning with 27-90x fewer parameters while maintaining high performance across various reasoning and language tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19146v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Puzzle: Distillation-Based NAS for Inference-Optimized LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/01e64d36-a997-4f75-ba32-812e1fe82f74/image.png?t=1735266532"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Deploying LLMs on resource-constrained hardware (e.g., a single GPU or edge devices) can be extremely <b>expensive</b> and <b>slow</b> at inference time.</p><ul><li><p class="paragraph" style="text-align:left;"><b>Neural Architecture Search (NAS)</b> can find specialized, smaller architectures—but typically requires extensive compute and big teacher models.</p></li><li><p class="paragraph" style="text-align:left;"><b>Knowledge distillation</b> helps preserve performance in a smaller model, but it’s often done in a global, all-at-once manner that doesn’t account for hardware constraints at a granular level.</p></li></ul><p class="paragraph" style="text-align:left;">The <b>Puzzle</b> framework aims to <b>co-optimize</b>:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Architecture design</b> (via NAS) and</p></li><li><p class="paragraph" style="text-align:left;"><b>Performance retention</b> (via distillation),<br>all while ensuring the resulting model meets specific <b>hardware constraints</b> (e.g., fits on a single GPU with minimal latency).</p></li></ol><p class="paragraph" style="text-align:left;"><b>How?:</b> Puzzle combines <b>blockwise local knowledge distillation (BLD)</b> with <b>mixed-integer programming</b> to systematically prune and restructure an LLM, step by step, into a smaller yet high-performing network.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Blockwise Local Distillation (BLD)</b></p><ul><li><p class="paragraph" style="text-align:left;">Instead of training a smaller model from scratch or distilling globally all at once, the approach <b>partitions the original model</b> into “blocks.” Each block is distilled locally, ensuring the sub-architecture within that block preserves knowledge from the teacher. This local focus can make distillation more stable and more aligned with the final architecture changes.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Mixed-Integer Programming for NAS</b></p><ul><li><p class="paragraph" style="text-align:left;">Puzzle sets <b>hardware constraints</b> (e.g., model size, memory footprint, or inference time) as <b>optimization objectives</b>. It then systematically <b>selects or prunes</b> blocks and layers, guided by a global objective function. Mixed-integer programming finds an <b>optimal combination</b> of blocks that fit the constraints while aiming for maximal performance.</p></li></ul></li><li><p class="paragraph" style="text-align:left;"><b>Iterative Refinement</b></p><ul><li><p class="paragraph" style="text-align:left;">The framework iterates, refining architecture choices and distilling block by block. Each iteration zeros in on a final architecture that best balances size, speed, and accuracy.</p></li></ul></li></ol><p class="paragraph" style="text-align:left;"><b>Results:</b> Nemotron-51B, derived from Puzzle, achieves 2.17x inference speedup fitting on a single NVIDIA H100 GPU, retaining <b>98.4%</b> of the teacher model’s performance.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.16991v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d9d7c48-bd8e-4aec-9056-6045aa8c1f96/image.png?t=1735267410"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose a model-agnostic self-distillation method, DynSDPB, which learns from the model&#39;s own previous mini-batch outputs. This approach dynamically adjusts distillation influence and temperature to improve early iteration accuracy. It is a fine-tuning policy that integrates self-correction and self-training methods without architectural modification.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Previous Mini-Batch Outputs: </b>Instead of referencing an external teacher, the model looks at its own predictions (logits, soft labels) from the <b>previous mini-batch</b>. These predictions become the <b>pseudo “teacher”</b> for the next iteration.</p></li><li><p class="paragraph" style="text-align:left;"><b>Dynamic Adjustments: </b>The distillation influence (how strongly the model trusts its past predictions) and temperature (softening or sharpening the pseudo labels) are <b>adjusted on the fly</b> during fine-tuning. This dynamic tuning ensures the self-distillation signal remains <b>useful and stable</b> as the model learns.</p></li><li><p class="paragraph" style="text-align:left;"><b>No Extra Architectural Changes: </b>Importantly, no additional layers or model capacity is required—just a revised <b>fine-tuning policy</b> that leverages the previous batch’s logits.</p></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">A word from today’s sponsor!</h2><div class="image"><a class="image__link" href="https://tldr.tech/signup?utm_source=LllmsResearch&utm_campaign=LllmsResearch-cpa-campaign&utm_medium=newsletter-sponsorship" rel="noopener" target="_blank"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/61a76fd7-15d4-4376-93e5-8a1ef5514f8b/image.png?t=1735278744"/></a></div><p class="paragraph" style="text-align:left;">Love Hacker News but don’t have the time to read it every day? Try <b>TLDR’s</b> free daily newsletter. <b>TLDR</b> covers the best tech, startup, and coding stories in a quick email that takes 5 minutes to read. No politics, sports, or weather (we promise). And it&#39;s read by over <b>1,250,000 people!</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://tldr.tech/signup?utm_source=LllmsResearch&utm_campaign=LllmsResearch-cpa-campaign&utm_medium=newsletter-sponsorship" target="_blank" rel="noopener noreferrer nofollow">Subscribe for free</a> now and you&#39;ll get our next newsletter tomorrow morning.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Quantization: Memory & Speed Gains</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.17691v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1404f4ad-231c-4585-8304-c4df8ee0358d/image.png?t=1735269442"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper explores how low-bit quantization interacts with the training level of LLMs, uncovering implications for model efficiency and future training strategies.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study involved analyzing over 1500 quantized LLM checkpoints with variations in model size and training levels. Scaling laws were derived to explore the relationship between quantization-induced degradation (QiD) and factors like training tokens, model size, and bit width.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study revealed that undertrained models are less susceptible to QiD impacts compared to fully trained small models. These insights provide benchmarks and predictions for quantization performance in massive future models expected to be trained with 100 trillion tokens.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.17525v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Pushing the Limits of Large Language Model Quantization via the Linearity Theorem</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Quantization often impacts performance unevenly across layers, but identifying which layers tolerate fewer bits while maintaining overall model quality is a challenge. This research introduces a <b>systematic framework</b> to minimize this trade-off, optimizing layer-specific quantization.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors introduce the &#39;linearity theorem&#39; to correlate layer-wise reconstruction error with increased perplexity due to quantization. They developed the HIGGS method employing Hadamard rotations and MSE-optimal grids and proposed an efficient dynamic programming approach for non-uniform per-layer quantization.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The proposed methods enhance accuracy-compression trade-offs for Llama-3.1, Llama-3.2, and Qwen models, outperforming existing data-free approaches like NF4.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.16158v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">MixPE: Quantization and Hardware Co-design for Efficient LLM Inference</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f42d6ab4-2539-4afa-85e3-f603fdb17e1c/image.png?t=1735269676"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research tackles the computational and memory challenges inherent in deploying LLMs by introducing a more efficient quantization approach.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> MixPE introduces a mixed-precision processing element to enhance inference efficiency. It minimizes dequantization operations using two innovations: performing dequantization post per-group mixed-precision matrix multiplication and replacing traditional multipliers with shift add operations, thus boosting computational and energy efficiency.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> MixPE achieves a 2.6× speedup and 1.4× energy reduction over current quantization accelerators.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.18065v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b5bb9375-8c00-4ebb-9264-42b7359e0694/image.png?t=1735270684"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Current hardware accelerators are rigid, often supporting only standard precisions (e.g., FP16 or INT8). This limitation hampers the potential of custom mixed-precision arithmetic, which could better exploit model-specific needs and improve performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> FlexiBit proposes <b>FlexiBit</b>, a bit-parallel accelerator architecture which supports flexible precision and dynamic formats. It overcomes the limitations of bit-serial designs by enabling <b>arbitrary mixed-precision computations</b>, allowing models to adapt precision based on layer or task-specific requirements.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> FlexiBit achieved 1.66x to 3.9x higher performance per area compared to existing architectures on GPT-3 using FP6 precision.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Infrastructure, Caching, and Serving Optimizations</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19379v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Marconi: Prefix Caching for the Era of Hybrid LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8e77e569-a8d0-4a76-a9ba-e9c0f118e3cb/image.png?t=1735270953"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Hybrid LLMs, which combine attention and recurrent layers, often waste compute and memory when caching prefixes. Traditional caching systems rely on <b>exact-match caching</b>, which is inefficient for hybrid models that process partially overlapping prefixes.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Marconi introduces novel cache admission and eviction policies that evaluate potential entries based on recency, predicted reuse likelihood, and computational savings versus memory costs. This approach allows efficient prefix caching for hybrid models, correcting traditional systems that require exact-match caches.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Marconi achieves up to 34.4× higher token hit rates and reduces time-to-first-token by 617 ms compared to existing systems, indicating significant performance improvements.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.18077v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ed7e9f92-1637-4a9c-9246-84ead445c21f/image.png?t=1735270990"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The KV (key-value) cache is a critical bottleneck in long-context LLM inference due to its <b>high memory requirements</b>. Standard 8-bit compression methods are insufficient for scaling to larger models and contexts.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study introduces MiniKV, a KV cache optimization method using a novel 2-bit layer-discriminative approach, preserving accuracy while significantly reducing cache size. Specialized CUDA kernels were developed to ensure compatibility with FlashAttention, tested across various long-context tasks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experiments demonstrated reduction in KV cache size by <b>86%</b> with minimal accuracy loss.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.18424v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> In multi-user LLM-serving systems, inefficient <b>context switching</b> leads to fairness issues, where certain users face delays due to resource contention or idling GPUs. Improving <b>context-switching efficiency</b> ensures fairer service-level objectives (SLOs) for all users.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers developed FastSwitch, a system that maintains memory allocation while reducing context-switching overhead through enhancements like better I/O utilization and minimizing GPU idleness, responding to identified inefficiencies in current systems.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> FastSwitch demonstrated significant performance improvements over existing systems like vLLM, with speedups ranging from 1.4 to 11.2 times in various metrics.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.13820v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">InstCache: A Predictive Cache for LLM Serving</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0e9735bd-2506-450c-960e-18ab252e5d57/image.png?t=1735271194"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper proposes InstCache, which predicts user instructions using an instruction-aligned LLM. They implement an instruction pre-population algorithm based on the negative log likelihood to optimize cache size and hit rate. InstCache is executed as a hash table to minimize lookup latency, allowing for quick deployment.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> InstCache attains a <b>51.3% improvement</b> in cache hit rate.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19477v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?: </b>The computational cost of <b>test-time inference</b> grows with model size and sequence length, often creating bottlenecks in scalable deployments. Reducing test-time compute without sacrificing accuracy is crucial for scaling LLMs efficiently.</p><p class="paragraph" style="text-align:left;"><b>How?: </b>Paper<b> </b>proposes a <b>two-stage algorithm</b>: (1) Generate <b>N candidate solutions</b> using the LLM. (2) Select the best solution via a <b>K-round knockout tournament</b> that compares pairs of solutions.</p><p class="paragraph" style="text-align:left;">The algorithm leverages <b>parallelization</b> and requires only <b>N × (K + 1)</b> LLM calls.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Hallucination and attribution</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.15102v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding how context influences LLM behavior is crucial for improving interpretability and efficiency. However, calculating context attribution is computationally expensive.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper introduces AttriBoT, using cached activations to avoid redundant operations, hierarchical attribution for reduced computation, and proxy models to emulate the behavior of larger models. This approach approximates the LOO error more efficiently than previous methods.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> AttriBoT achieves over 300x speedup in computing context attributions, making it 30x faster than generating the response, while maintaining fidelity to target model&#39;s LOO error.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.04996v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/11770965-0720-4307-99b8-a7e22a64a781/image.png?t=1735271976"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Training dense multi-modal transformers for handling text, images, and speech demands <b>substantial computational resources</b>, which limits their scalability and accessibility. An efficient architecture is required to reduce <b>floating-point operations (FLOPs)</b> while maintaining high performance across modalities.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces the Mixture-of-Transformers (MoT), a sparse multi-modal transformer architecture. MoT separates model parameters by modality while maintaining global self-attention, thus allowing modality-specific processing. This separation reduces the computational cost by utilizing fewer floating-point operations (FLOPs) than dense models.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Research showcase a comparable performance to dense multi-modal models while using <b>55.8% fewer FLOPs</b>, demonstrating significant resource savings.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Advanced Pruning, Layer Slicing, and Attention Mechanisms</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.14055v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Pruning LLMs can reduce their size and computational cost but often results in uneven performance across domains. This research addresses these imbalances by introducing a pruning method that maintains robust performance across diverse datasets and tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces DRPruning, integrating distributionally robust optimization to improve structured pruning. This method refines pruning and pretraining processes, automatically finding optimal reference losses and data ratios to prevent biased performance across different domains.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experiments in both monolingual and multilingual settings indicate DRPruning outperforms similarly sized models in pruning metrics such as perplexity, downstream tasks, and instruction tuning, enhancing robustness against distribution shifts.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.07191v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">The Super Weight in Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d6c26178-cc22-47b3-a618-666c7d73fd1f/image.png?t=1735273662"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> A small fraction of LLM parameters, termed <b>&quot;super weights,&quot;</b> are disproportionately critical for model performance. Understanding these parameters enables more efficient pruning and quantization while preserving accuracy.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study identifies &#39;super weights&#39; using a data-free approach involving a single forward pass through the model. The impact of pruning these weights is examined by evaluating changes in perplexity and zero-shot accuracy. The researchers also examine how preserving &#39;super activations&#39; affects quantization strategies.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Pruning a single &#39;super weight&#39; dramatically increases perplexity and reduces accuracy. Preserving these weights allows simple quantization methods to match state-of-the-art performance and enables larger quantization block sizes.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.17116v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Star Attention: Efficient LLM Inference over Long Sequences</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8a1c67ff-759e-46bc-a739-29526920a839/image.png?t=1735273821"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Transformers face <b>quadratic complexity</b> in their attention mechanism, making inference on long sequences computationally inefficient. Star Attention addresses this by reducing memory and compute requirements.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> A two-phase block-sparse approximation called Star Attention is introduced. Phase one uses blockwise-local attention processed in parallel across multiple hosts. Phase two applies sequence-global attention for query and response tokens attending to cached tokens. This method integrates seamlessly with global attention-trained LLMs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The method reduces memory requirements and inference times by up to 11x while maintaining 95-100% accuracy.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.03493v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LASER: Attention with Exponential Transformation</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b598eb7b-1e9e-45d5-8cda-543d3c90ca3a/image.png?t=1735274061"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Transformer attention mechanisms suffer from poor <b>gradient signal backpropagation</b>, which can hinder learning and slow convergence. LASER addresses this issue by improving gradient flow in the attention mechanism.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose LASER, an attention mechanism with larger gradient signals than standard attention, implemented with minor modifications to existing setups. They conducted experiments with autoregressive LLMs up to 2.2 billion parameters.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.15558v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Reassessing Layer Pruning in LLMs: New Insights and Methods</a></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.09266v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception</a></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19921v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">SIMS: Simulating Human-Scene Interactions with Real World Script Planning</a> - Generate scripts for human-scene interactions in physics-based animations, enabling more dynamic and realistic character motion for games and simulations.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19869v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">AIDetx: a compression-based method for identification of machine-learning generated text</a> - <b>AIDetx</b> leverages data compression techniques to accurately distinguish between human-written and AI-generated text, achieving exceptional detection accuracy with minimal computational cost.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.19635v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Build An Influential Bot In Social Media Simulations With Large Language Models</a> - A novel framework integrates LLMs into <b>agent-based social media simulations</b>, replicating opinion dynamics and influencer behavior to better understand public opinion formation.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.18764v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding</a> - <b>CoVis</b> combines segmentation networks with LLM-based content generation to provide detailed and holistic graphic visual interpretations, improving the accessibility of complex visual data.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.17933v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Automated Test Transfer Across Android Apps Using Large Language Models</a> - <b>LLMigrate</b> simplifies and accelerates UI test transfer between Android apps, significantly reducing the time and effort required for mobile app testing.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.15867v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs</a> </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0edfe39d-1f37-4dc2-ac69-9eb25ef2e2d1/image.png?t=1735274569"/></div><p class="paragraph" style="text-align:left;"><b>PanoLlama</b> transforms panoramic image generation by reimagining it as a next-token prediction task, enabling endless, coherent panoramas with enhanced scalability and precision.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.16341v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">From CISC to RISC: language-model guided assembly transpilation</a> - <b>CRT</b> automates the translation of x86 assembly code to ARM, facilitating the transition to more energy-efficient architectures while ensuring high accuracy and performance gains.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2411.14672v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-in-november-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Multiverse of Greatness: Generating Story Branches with LLMs</a> - Generate branching storylines with enhanced coherence and creativity, revolutionizing AI-driven storytelling for visual novels and dynamic narratives.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=edd7de56-1521-4a4b-8786-2fe55ebc256b&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published from September 15-19, 2024</title>
  <description>From surgical edits to self-evolving minds read the summary of key research papers published between September 15-19, 2024</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c3d8a83e-842c-4c14-8bbb-26033f7690de/sept_15-19_collage.png" length="319094" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-from-september-15-19-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-from-september-15-19-2024</guid>
  <pubDate>Sat, 12 Oct 2024 20:00:00 +0000</pubDate>
  <atom:published>2024-10-12T20:00:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><p class="paragraph" style="text-align:left;">🧠<b> LLM Surgery:</b> Learn how researchers can update or remove specific knowledge in LLMs without retraining the entire model, making AI more efficient and up-to-date.</p><p class="paragraph" style="text-align:left;">🔀<b> Adaptive Transformers:</b> Meet new Transformer models that adjust their depth dynamically, processing simpler inputs faster by skipping unnecessary layers.</p><p class="paragraph" style="text-align:left;"><b>⚡ FP8 Training Acceleration:</b> Discover how using FP8 precision speeds up training and reduces memory usage in LLMs, leading to faster and more efficient models without sacrificing performance.</p><p class="paragraph" style="text-align:left;">🎨<b> Creative LLM Applications:</b> See how LLMs are being used in innovative ways—from assisting in Alzheimer&#39;s detection to generating recipes and food images.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">Don’t have much time?</h2><p class="paragraph" style="text-align:left;">No worries! Jump onto your car or subway for commute and listen to this amazing podcast covering these papers!</p><div class="recommendation" id="4a7ff9df-02d2-4413-8080-f76b2b96ecfa"><figure class="recommendation__logo"><img src="https://beehiiv-images-production.s3.amazonaws.com/uploads/asset/file/6051aa1b-2c41-4cef-8ec3-78cd209d7dae/1_FLfsM9mhZkRq8mhj_5OI7Q__1_.png?t=1728755517"/></figure><h3 class="recommendation__title"> Podcast discussing papers improving core performance of LLMs [september 15-19, 2024] </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzRhN2ZmOWRmLTAyZDItNDQxMy04MDgwLWY3NmIyYjk2ZWNmYS9jb3JlJTIwc2VwdCUyMDE1LTE5Lndhdj9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NlxcdTAwMjZYLUFtei1DcmVkZW50aWFsPUFLSUFRQ01IVFFTRTJKR0FHWEhKJTJGMjAyNjA5MTMlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdFxcdTAwMjZYLUFtei1EYXRlPTIwMjYwOTEzVDExMzM0OFpcXHUwMDI2WC1BbXotRXhwaXJlcz02MDQ4MDBcXHUwMDI2WC1BbXotU2lnbmVkSGVhZGVycz1ob3N0XFx1MDAyNlgtQW16LVNpZ25hdHVyZT03MmFmMDhmNTdkOGM3YjkyNTE3YTVkNjUwZjY2OGRkZjRjMTI4M2ZiNzAyYjkyMjRiOWNiN2M4OGRjOTJhZGJiXCIsXCJ0eXBlXCI6XCJhdWRpby94LXdhdlwiLFwidGh1bWJuYWlsVXJsXCI6XCJodHRwczovL2JlZWhpaXYtaW1hZ2VzLXByb2R1Y3Rpb24uczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Fzc2V0L2ZpbGUvNjA1MWFhMWItMmM0MS00Y2VmLThlYzMtNzhjZDIwOWQ3ZGFlLzFfRkxmc005bWhaa1JxOG1oal81T0k3UV9fMV8ucG5nP3Q9MTcyODc1NTUxN1wiLFwidGl0bGVcIjpcIlBvZGNhc3QgZGlzY3Vzc2luZyBwYXBlcnMgaW1wcm92aW5nIGNvcmUgcGVyZm9ybWFuY2Ugb2YgTExNcyBbc2VwdGVtYmVyIDE1LTE5LCAyMDI0XVwifSI." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.13054v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models</a></p><p class="paragraph" style="text-align:left;">Updating large language modes (LLMs) with new information and removing outdated or problematic knowledge without retraining from scratch is a significant challenge. Traditional fine-tuning methods are computationally expensive and risk overwriting valuable existing knowledge.</p><p class="paragraph" style="text-align:left;"><b>How?</b><br>The paper introduces <b>LLM Surgery</b>, a framework that efficiently modifies an LLM&#39;s knowledge base using a three-component objective function:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Unlearning Objective (Reverse Gradient):</b> Removes outdated knowledge by performing gradient ascent (reverse gradient) on an unlearning dataset D<sub>unlearn​</sub>. This effectively increases the model&#39;s loss on this data, causing it to forget specific information.</p></li><li><p class="paragraph" style="text-align:left;"><b>Learning Objective (Gradient Descent):</b> Incorporates new knowledge by performing gradient descent on an update dataset D<sub>update</sub>.</p></li><li><p class="paragraph" style="text-align:left;"><b>Retention Objective (KL Divergence Minimization):</b> Ensures that the model&#39;s outputs remain similar to the original model on a retain dataset D<sub>retain</sub> by minimizing the Kullback-Leibler (KL) divergence between the original and updated models&#39; output distributions.</p></li></ol><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dd92ce61-a9ec-4f98-b221-f885b883d8c3/image.png?t=1728751652"/></div><p class="paragraph" style="text-align:left;"><b>Results</b><br>Applying LLM Surgery to the LLaMA2-7B model, the researchers achieved:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Significant forgetting</b> on the unlearn dataset.</p></li><li><p class="paragraph" style="text-align:left;">A <b>20% improvement</b> on the update dataset.</p></li><li><p class="paragraph" style="text-align:left;"><b>Minimal performance degradation</b> on the retain dataset.</p></li></ul><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.12517v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Scaling FP8 Training to Trillion-Token LLMs</a></p><p class="paragraph" style="text-align:left;">Training LLMs using lower-precision formats like FP8 can significantly reduce computational costs and memory usage. However, scaling FP8 training to large datasets (trillion tokens) reveals instabilities that need to be addressed for practical deployment.</p><p class="paragraph" style="text-align:left;"><b>How?</b><br>The researchers identified that the <b>SwiGLU activation function</b> amplifies outlier gradients over long training durations, leading to numerical instabilities in FP8 training. To mitigate this, they introduced:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Smooth-SwiGLU Activation Function:</b> A modification that stabilizes FP8 training by smoothing the activation function, preventing gradient explosions without altering its functional behavior.</p></li><li><p class="paragraph" style="text-align:left;"><b>FP8 Quantization of Adam Optimizer Moments:</b> Applied FP8 quantization to the first and second moments in the Adam optimizer to maintain stability and efficiency.</p></li></ul><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9b11d9aa-a636-40c5-92b0-df5b5e4683a5/image.png?t=1728751839"/></div><p class="paragraph" style="text-align:left;"><b>Results</b><br>Training a 7B parameter model on 2 trillion tokens using 256 Intel Gaudi2 accelerators:</p><ul><li><p class="paragraph" style="text-align:left;">Achieved <b>comparable performance</b> to BF16 baselines.</p></li><li><p class="paragraph" style="text-align:left;">Realized up to a <b>34% speedup</b> and <b>43% memory reduction</b> during training.</p></li><li><p class="paragraph" style="text-align:left;">Demonstrated that FP8 training is feasible and efficient for large-scale LLMs when addressing activation-induced instabilities.</p></li></ul><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.10870v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Adaptive Large Language Models by Layerwise Attention Shortcuts</a><br><br>Traditional Transformer architectures process inputs sequentially layer by layer, which can be inefficient and may not adapt computations based on input complexity. Introducing adaptability can improve efficiency and performance.</p><p class="paragraph" style="text-align:left;"><b>How?</b><br>Paper proposes modifications to the Transformer architecture by adding <b>attention shortcuts</b> from the final layer to all intermediate layers:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Layerwise Attention Shortcuts:</b> The final layer attends to the outputs of all previous layers, allowing it to directly access lower-level features.</p></li><li><p class="paragraph" style="text-align:left;"><b>Dynamic Depth Adaptation:</b> The model can adjust its computational depth based on input complexity, effectively skipping unnecessary layers for simpler inputs.</p></li></ul><p class="paragraph" style="text-align:left;">The modified attention mechanism at the final layer LL becomes:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4faba701-7f0e-466e-8eb1-d5fc37ea72c7/image.png?t=1728752009"/></div><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.10516v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval</a></p><p class="paragraph" style="text-align:left;">Handling long contexts in LLMs is computationally intensive due to the quadratic complexity of self-attention mechanisms. Reducing inference time and memory consumption is critical for deploying LLMs in practical applications.</p><p class="paragraph" style="text-align:left;"><b>How?</b><br>The paper introduces <b>RetrievalAttention</b>, a method that accelerates inference by:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Key-Value (KV) Caching with Vector Retrieval:</b> Instead of processing the entire context, the model retrieves only the most relevant key-value pairs from the KV cache using Approximate Nearest Neighbor Search (ANNS).</p></li><li><p class="paragraph" style="text-align:left;"><b>Attention-Aware Vector Search:</b> Adapts the retrieval process to the distribution of query vectors, ensuring that relevant context is efficiently retrieved.</p></li></ul><p class="paragraph" style="text-align:left;">for a query vector <b>q</b>, the attention is computed over a subset of keys <b>{k</b><sub><b>i</b></sub><b>}</b> and values <b>{v</b><sub><b>i</b></sub><b>}</b>:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9399894a-766e-46d4-a7d0-90870e6f7f95/image.png?t=1728752900"/></div><p class="paragraph" style="text-align:left;">By selecting only the <b>top-k</b> keys <b>k</b><sub><b>i</b></sub> that are most similar to <b>q</b>, the computation is significantly reduced.</p><p class="paragraph" style="text-align:left;"><b>Results</b></p><p class="paragraph" style="text-align:left;">The method accesses only a small fraction of the KV cache (1-3% KV cache access) during inference. It achieves up to a <b>4x speedup</b> in inference time with minimal impact on model performance!</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.10715v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Self-Attention Limits Working Memory Capacity of Transformer-Based Models</a></p><p class="paragraph" style="text-align:left;">By understanding the limitations of LLMs working memory capacity their ability to handle tasks requiring long-term dependencies can be improved.</p><p class="paragraph" style="text-align:left;"><b>How?</b><br>Paper investigate the working memory capacity of Transformers by training models on N-back tasks to test their ability to remember information NN steps back. It observes that as NN increases, the entropy of the attention distribution increases, leading to dispersed attention and reduced capacity.</p><p class="paragraph" style="text-align:left;">Paper proposes that self-attention mechanisms inherently limit working memory capacity due to this dispersion. It shows that<span style="color:#000000;font-family:-webkit-standard;font-size:medium;"> </span>attention mechanisms contribute to capacity limits due to increased entropy. </p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.12914v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Defending against Reverse Preference Attacks is Difficult</b></a><b> - </b>shows how adversarial reinforcement learning can induce harmful behaviours in safety-aligned LLMs through Reverse Preference Attacks (RPAs). Explores defence strategies using Constrained Markov-Decision Processes, finding that &#39;online&#39; defences with controlled loss functions are effective, while &#39;offline&#39; defences are less so.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11844v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow"><b>MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts</b></a><b> </b>proposes MEOW, a gradient descent-based method for unlearning sensitive information in LLMs. Generates inverted facts and utilizes the MEMO metric to quantify memorization, effectively improving forgetfulness without sacrificing model utility.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11690v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow"><b>LLM-Powered Text Simulation Attack Against ID-Free Recommender Systems</b></a><b> </b>Introduces a text poisoning attack using LLMs to manipulate textual information of target items in ID-free recommender systems. Leverages popular item characteristics to create promotional descriptions that mimic popular items.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11445v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Jailbreaking Large Language Models with Symbolic Mathematics</b></a><b> </b>shows how encoding harmful prompts into mathematical problems can bypass LLM safety measures. It encodes harmful natural language prompts into mathematical problems.The MathPrompt technique achieves 73.6% attack success rate in top 13 LLMs!</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12541v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Profiling Patient Transcripts Using LLMs for Alzheimer&#39;s Detection</a></b><b> </b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/91a7d165-7c09-44c8-b941-d39c5c7c390e/image.png?t=1728712815"/></div><p class="paragraph" style="text-align:left;">Introduces a framework that utilizes LLMs for profiling linguistic deficits in patient transcripts. Shows an 8.51% improvement over state-of-the-art methods in Alzheimer&#39;s disease detection.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12010v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">ChefFusion: Multimodal Foundation Model for Recipe and Food Image Generation</a></b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cb6a9c5b-f431-4f27-aaeb-7ae53fd418b7/image.png?t=1728713022"/></div><p class="paragraph" style="text-align:left;">Paper develops ChefFusion, a foundation model integrating LLMs with pre-trained image encoders and decoders. Performs tasks like recipe generation and food image generation, advancing multimodal food computing.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12561v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Leveraging LLMs for Automated Framing Analysis in TV Shows</a></b><b> </b>uses prompt engineering to guide LLMs in identifying framing in spoken content from TV shows. Achieves agreement rates up to 43% with human annotations, offering a support tool for framing analysis.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12471v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Arena 4.0: A ROS2 Platform for Human-Centric Navigation</a></b><b> </b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fc1cb8a8-5b9d-447c-bc3b-2cda807a314d/image.png?t=1728738378"/></div><p class="paragraph" style="text-align:left;">Paper<b> </b>introduces Arena 4.0, utilizing LLMs and diffusion models to generate human-centric environments from text prompts or floorplans. Offers a 3D model database and migrates to ROS 2 for hardware compatibility, validated through a user study.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12274v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Hierarchical LLMs In-the-Loop Optimization for Multi-Robot Target Tracking</a></b><b> </b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dece5164-0546-4298-ae9c-c028ff979402/image.png?t=1728738440"/></div><p class="paragraph" style="text-align:left;">Paper introduces a hierarchical LLM-based framework for optimizing multi-robot task allocation in real-time. Validated through simulations and real-world experiments, enhancing adaptability in hazardous environments.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.10027v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">E2Map: Self-Reflective Robot Navigation with LLMs</a></b><b> </b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ddbe36ec-1343-496a-9994-06fd2d3ee295/image.png?t=1728738489"/></div><p class="paragraph" style="text-align:left;">Paper<b> </b>creates an &#39;Experience-and-Emotion Map&#39; integrating LLM knowledge with real-world experiences. Allows robots to adjust behavior dynamically in stochastic environments.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">What Would You Ask When You First Saw a</a></b><sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">2</a></b></sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">+b</a></b><sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">2</a></b></sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">=c</a></b><sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">2</a></b></sup><b><a class="link" href="http://arxiv.org/abs/2409.17172v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">? Evaluating LLM on Curiosity-Driven Questioning</a></b><b> </b>Introduces a framework to assess LLMs ability to generate curious questions when presented with new scientific statements. Surprisingly it found out that models like GPT-4 and Mistral 8x7b generate coherent questions, with smaller models matching or exceeding their effectiveness.</p><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.12866v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications</a></b><b> </b>introduces SpecEval, a framework using program specifications to assess LLMs code comprehension. It designs four tasks utilizing formal specifications, counterfactual analysis, and consistency checks.</p><p class="paragraph" style="text-align:left;"><b><a class="link" href="http://arxiv.org/abs/2409.10955v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow">Investigating Context-Faithfulness in Large Language Models</a></b><b> </b>studies factors influencing context-faithfulness in LLMs by quantifying memory strength and evaluating different evidence styles. Uses datasets with well-known and long-tail questions.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.09629v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Confidence Estimation for LLM-Based Dialogue State Tracking</b></a><b> </b><i>e</i>xplores methods to estimate confidence scores in dialogue state tracking. It finds that fine-tuning open-weight LLMs enhances calibration accuracy and improves performance.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.18996v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">From Linguistic Giants to Sensory Maestros: A Survey on Cross-Modal Reasoning with Large Language Models</a> - <span style="color:#000000;font-family:Arial, sans-serif;font-size:medium;">survey of methodologies for cross-modal reasoning using LLMs.</span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11650v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-15-19-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Art and Science of Quantizing Large-Scale Models: A Comprehensive Overview</a> - <span style="color:#000000;font-family:Arial, sans-serif;font-size:medium;">A review of variations of quantization methods like post-training quantization (PTQ) and quantization-aware training (QAT).</span></p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=13edfcc6-a74f-4528-be7b-d6644b8658f1&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published from September 9-13, 2024</title>
  <description>From Accelerated Speech Recognition to Advanced Reasoning - weekly summary of papers improving LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/132466e5-9e7e-4f99-b056-7819cc095cd9/9-13_september.png" length="320250" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-from-september-9-13-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-from-september-9-13-2024</guid>
  <pubDate>Sun, 29 Sep 2024 20:13:28 +0000</pubDate>
  <atom:published>2024-09-29T20:13:28Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;">Dear readers,</p><p class="paragraph" style="text-align:left;">This edition is long so you may need to click on “view entire content” button at the end to read full edition. Not my fault, it’s Google!</p><p class="paragraph" style="text-align:left;">Speaking of Google, I have tried NotebookLM to create an interesting podcast discussing this edition’s research papers. If you don’t have time to read, LISTEN!😄</p><p class="paragraph" style="text-align:left;">Share you thoughts if you liked this podcast format or want any modifications! This podcast is yours, treat it like one! - shower the comments for improvements.👏🏻</p><div class="recommendation" id="452adb6f-1909-4d18-a649-cfefb75bbf66"><figure class="recommendation__logo"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" fill="currentColor"><path d="M14.8287 7.75737L9.1718 13.4142C8.78127 13.8047 8.78127 14.4379 9.1718 14.8284C9.56232 15.219 10.1955 15.219 10.586 14.8284L16.2429 9.17158C17.4144 8.00001 17.4144 6.10052 16.2429 4.92894C15.0713 3.75737 13.1718 3.75737 12.0002 4.92894L6.34337 10.5858C4.39075 12.5384 4.39075 15.7042 6.34337 17.6569C8.29599 19.6095 11.4618 19.6095 13.4144 17.6569L19.0713 12L20.4855 13.4142L14.8287 19.0711C12.095 21.8047 7.66283 21.8047 4.92916 19.0711C2.19549 16.3374 2.19549 11.9053 4.92916 9.17158L10.586 3.51473C12.5386 1.56211 15.7045 1.56211 17.6571 3.51473C19.6097 5.46735 19.6097 8.63317 17.6571 10.5858L12.0002 16.2427C10.8287 17.4142 8.92916 17.4142 7.75759 16.2427C6.58601 15.0711 6.58601 13.1716 7.75759 12L13.4144 6.34316L14.8287 7.75737Z"></path></svg></figure><h3 class="recommendation__title"> Summary of papers from today&#39;s edition </h3><iframe src="https://audio.beehiiv.com?token=eyJhbGciOiJub25lIn0.IntcImJhY2tncm91bmRDb2xvclwiOm51bGwsXCJiYWNrZ3JvdW5kVGhlbWVcIjpudWxsLFwic3JjXCI6XCJodHRwczovL2JlZWhpaXYtcHVibGljYXRpb24tZmlsZXMuczMuYW1hem9uYXdzLmNvbS91cGxvYWRzL2Rvd25sb2FkYWJsZXMvYTE1YTk2NzktMThlYy00M2U2LThlNDUtNmMyYjA2NGY5NWQ2LzQ1MmFkYjZmLTE5MDktNGQxOC1hNjQ5LWNmZWZiNzViYmY2Ni9jb3JlLSUyMFNlcHRlbWJlciUyMDl0aC0xM3RoJTJDJTIwMjAyNC53YXY_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTZcXHUwMDI2WC1BbXotQ3JlZGVudGlhbD1BS0lBUUNNSFRRU0UySkdBR1hISiUyRjIwMjYwOTEzJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3RcXHUwMDI2WC1BbXotRGF0ZT0yMDI2MDkxM1QxMTMzNDlaXFx1MDAyNlgtQW16LUV4cGlyZXM9NjA0ODAwXFx1MDAyNlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdFxcdTAwMjZYLUFtei1TaWduYXR1cmU9ZjQxNmZmNGUzODRhOWYwZTQ5YzUyYmM5ZjUyNDljNTRmNDZmNGUyNzlhNDZmYjczZTQxNmU1YzEwY2IxYjNjZVwiLFwidHlwZVwiOlwiYXVkaW8veC13YXZcIixcInRodW1ibmFpbFVybFwiOm51bGwsXCJ0aXRsZVwiOlwiU3VtbWFyeSBvZiBwYXBlcnMgZnJvbSB0b2RheSdzICBlZGl0aW9uXCJ9Ig." frameborder="0" width="100%" height="162" allow="encrypted-media"></iframe></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>AIPO refines LLMs to deliver concise and meaningful responses</b> - enhancing alignment with human preferences for better AI interactions.</p></li><li><p class="paragraph" style="text-align:left;"><b>CPL introduces Critical Planning Step Learning, boosting LLM reasoning and significantly improving problem</b> - solving accuracy.</p></li><li><p class="paragraph" style="text-align:left;"><b>Faster Speech-LLaMA accelerates speech recognition by predicting multiple tokens at once</b> - reducing inference time without sacrificing accuracy.</p></li><li><p class="paragraph" style="text-align:left;"><b>AdaCAD</b> <b>dynamically balances contextual and parametric knowledge in LLMs</b>, resolving conflicts for clearer and more accurate outputs.</p></li><li><p class="paragraph" style="text-align:left;"><b>E2LLM extends LLMs&#39; capabilities to handle longer contexts efficiently</b> - enhancing understanding and reasoning over extended inputs.</p></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08845v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">AIPO: Improving Training Objective for Iterative Preference Optimization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Improving the alignment of LLMs through iterative preference optimization is important for their effective operation and scaling, as preference optimization is becoming a popular alternative to proximal policy optimization.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> AIPO introduces a new training method that encourages LLMs to produce concise and meaningful responses, addressing the length exploitation problem. It does so by:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Carefully selecting training data:</b> AIPO focuses on using synthetic data, particularly instructions created by humans and responses generated by the LLM itself.</p></li><li><p class="paragraph" style="text-align:left;"><b>Iterative training:</b> The model is trained in multiple rounds, incorporating feedback from a reference model (a more powerful LLM) in each round. This helps refine the model&#39;s responses over time.</p></li><li><p class="paragraph" style="text-align:left;"><b>Agreement-aware training:</b> AIPO uses a special coefficient to adjust the training process based on how well the LLM&#39;s responses align with the reference model&#39;s preferences. This ensures that the LLM is rewarded for producing responses that are both concise and aligned with human-like preferences.</p></li></ol><p class="paragraph" style="text-align:left;"><b>In simpler terms, imagine you&#39;re teaching a child to write good stories.</b> Instead of just telling them to write longer stories, you:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Give them specific examples</b> of good stories and let them learn from their own attempts.</p></li><li><p class="paragraph" style="text-align:left;"><b>Provide feedback in stages</b>, gradually refining their writing style and content.</p></li><li><p class="paragraph" style="text-align:left;"><b>Reward them not just for length, but also for clarity, conciseness, and creativity</b>, using a teacher&#39;s (reference model&#39;s) opinion as a guide.</p></li></ol><p class="paragraph" style="text-align:left;">AIPO applies a similar concept to LLMs, encouraging them to produce responses that are not just long, but also high-quality and aligned with human expectations.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08642v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">CPL: Critical Planning Step Learning Boosts LLM Generalization in Reasoning Tasks</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Reasoning improvement of LLMs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> This paper introduces Critical Planning Step Learning (CPL) that uses Monte Carlo Tree Search (MCTS) to identify diverse planning steps in multi-step reasoning tasks. CPL focuses on learning step-level planning preferences based on long-term outcomes to improve planning capabilities. It incorporates Step-level Advantage Preference Optimization (Step-APO), which integrates an advantage estimate for step-level preference pairs using MCTS within Direct Preference Optimization (DPO) techniques, capturing fine-grained supervision at each reasoning step. Experiments were conducted on GSM8K and MATH datasets.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The CPL method significantly improved model performance on GSM8K +10.5% and MATH by 6.5%.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08148v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Faster Speech-LLaMA Inference with Multi-token Prediction</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/784e0477-996f-4562-aeac-c12e65f42d43/image.png?t=1727616514"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Reducing the inference time of Speech-LLaMA models.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Meta team proposes predicting multiple tokens per decoding step instead of the traditional sequential approach. They explore various model architectures and implement threshold-based and verification-based inference strategies. Additionally, they introduce a prefix-based beam search decoding method for efficient minimum word error rate training. The models are then evaluated across several public speech recognition benchmarks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The proposed models achieve a  3.2x reduction in decoder calls while maintaining or improving word error rate (WER) performance.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.07394v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/efc9ed68-bff9-4868-acf0-d62c828d7909/image.png?t=1727616793"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The paper tries to resolve knowledge conflict in LLMs, where discrepancies between contextual and parametric information can impair performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper introduces AdaCAD, a dynamic approach that calculates the degree of conflict between contextual and parametric knowledge using the Jensen-Shannon divergence. This calculated divergence helps to dynamically adjust the decoding process at an instance level, rather than using static methods. Tests involved comparing AdaCAD to other contrastive methods across four models on various QA and summarization tasks, focusing on how well it handles varying degrees of conflict.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> AdaCAD yielded an average accuracy gain of 14.21% over a static contrastive baseline.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06691v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Geometric-Averaged Preference Optimization for Soft Preference Labels</a> </p><p class="paragraph" style="text-align:start;">This research paper argues that limiting human preferences to binary choices is an oversimplification and proposes using <b>soft preference labels</b>, represented by a probability (p̂), to reflect the nuanced nature of human preferences. To incorporate these soft labels into LLMs alignment, the authors introduce <b>weighted geometric averaging</b> into the loss function of Direct Preference Optimization (DPO) and related algorithms like IPO and ROPO. This method weights the likelihoods of the chosen and rejected responses based on the soft preference label—strong preferences (p̂ closer to 1) result in larger gradient updates during training than weak preferences (p̂ closer to 0.5). This weighting scheme allows the model to prioritize learning from high-confidence preference pairs while mitigating the effects of noisy data points and the tendency to over-optimize for response length.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06679v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/232860ca-0cac-4c8f-9e9d-80682eae25fe/image.png?t=1727637010"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Extends context length</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper introduces E2LLM, a model that enhances long-context capabilities by segmenting long inputs into chunks, which are compressed into embedding vectors using a pre-trained text encoder. These embeddings are then aligned with a decoder-only LLM through an adapter. The training objectives focus on reconstructing the encoder’s output and fine-tuning for long-context instructions, aiding LLMs in interpreting soft prompts.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06411v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Length Desensitization in Directed Preference Optimization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Current LLMs which are trained with Direct Preference Optimization (DPO) suffer from excessive verbosity. </p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper performs a theoretical analysis of DPO and its tendency to correlate rewards with data length, leading to verbosity. </p><p class="paragraph" style="text-align:left;">To handle this, paper introduces <b>LD-DPO</b>, a novel algorithm designed to mitigate length sensitivity in DPO by decoupling the preference for verbosity from other human-like preferences.</p><p class="paragraph" style="text-align:left;">LD-DPO works by:</p><ul><li><p class="paragraph" style="text-align:left;">Identifying the &quot;public length&quot; (lp) of a response pair—the number of tokens shared by both responses.</p></li><li><p class="paragraph" style="text-align:left;">Decomposing the likelihood of the longer response into the product of the likelihood of the public-length portion and the likelihood of the &quot;excessive&quot; portion.</p></li><li><p class="paragraph" style="text-align:left;">Introducing a hyperparameter α (ranging from 0 to 1) to control the influence of the excessive portion on the likelihood. A lower α weakens the impact of the excessive length, reducing DPO&#39;s sensitivity to it.</p></li><li><p class="paragraph" style="text-align:left;">By adjusting α, LD-DPO aims to find a balance between reducing verbosity and preserving the LLM&#39;s ability to learn other valuable preferences from the training data.</p></li></ul><p class="paragraph" style="text-align:left;"><b>Results:</b> LD-DPO consistently reduced response lengths by 10-40% compared to DPO.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06328v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Extracting Paragraphs from LLM Token Activations</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper explores the inner workings of LLMs beyond token-level predictions.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Paper analyze the role of the &quot;\n\n&quot; double newline token in single-token activations. The authors explore how these activations encode information about the content of subsequent paragraphs. They employ a method of patching these activations to assess their information transfer capabilities regarding contextual understanding in paragraph formation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.05746v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LLMs Will Always Hallucinate, and We Need to Live With This</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research claims that hallucination is not a bug in LLMs that can be solved instead it is a nature of LLM and is inevitable.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper uses mathematical concepts and logic to argue that several factors make structural hallucinations inevitable:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Incomplete Training Data:</b> It&#39;s impossible to capture all of human knowledge in a dataset, even a massive one. This incompleteness means that LLMs will always encounter situations or facts they haven&#39;t seen before, leading to potential hallucinations.</p></li><li><p class="paragraph" style="text-align:left;"><b>Difficulty Finding the Right Information:</b> Even with a complete dataset, finding the <i>exact</i> information needed for a specific query can be extremely difficult for an LLM. Think of it like searching for a needle in a haystack—there&#39;s just too much data to sift through perfectly.</p></li><li><p class="paragraph" style="text-align:left;"><b>Misinterpreting What We Mean:</b> Natural language is ambiguous, meaning the same phrase can have different meanings depending on the context. LLMs often struggle to correctly interpret the true meaning behind our prompts, leading to hallucinations.</p></li><li><p class="paragraph" style="text-align:left;"><b>Unpredictable Outputs:</b> The paper draws a parallel between how LLMs generate text and a famous problem in computer science called the &quot;Halting Problem.&quot; Basically, it&#39;s impossible to perfectly predict what output an LLM will generate beforehand, just like it&#39;s impossible to always know if a computer program will run forever or eventually stop. This unpredictability makes hallucinations unavoidable.</p></li><li><p class="paragraph" style="text-align:left;"><b>Fact-Checking Isn&#39;t Foolproof:</b> While fact-checking can help, the paper argues that it&#39;s mathematically impossible to have a fact-checking system that catches every single hallucination.</p></li></ul><p class="paragraph" style="text-align:left;"><b>The Bottom Line:</b> We need to accept that LLMs will always have the potential to hallucinate and focus on developing strategies to manage and mitigate this risk.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.13724v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Logically Consistent Language Models via Neuro-Symbolic Integration</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper improves reasoning ability in LLMs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a neuro-symbolic methodology that integrates neuro-symbolic reasoning into the LLM training process. This involves developing a loss function based on logical consistency with a predefined set of facts and rules, allowing the LLM to be trained to adhere to these constraints. The approach is balanced between full fine-tuning and offloading reasoning tasks to external tools, aiming to teach LLMs logical consistency with minimal fine-tuning. It also facilitates the inclusion and integration of multiple logical constraints efficiently, enabling the models to handle a wide range of logical requirements.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>LLMOps & GPU level optimization</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06646v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Optimal Workload Placement on Multi-Instance GPUs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research manages LLMs inferencing workloads utilizing the partitioning feature, Multi-Instance GPU (MIG).</p><p class="paragraph" style="text-align:left;"><b>How?:</b> This paper developed two approaches: an optimization method and a heuristic method. Both aim to efficiently place or migrate workloads to minimize the number of GPUs used and reduce memory and compute wastage.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The research demonstrated up to a 2.85x improvement in the number of GPUs used and up to a 70% reduction in GPU wastage over baseline heuristics.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.09086v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Deploying multimodal large language models (MLLMs) on edge devices.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper introduce Inf-MLLM, a framework for efficient streaming inference of MLLMs on a single GPU. They identify a novel attention pattern called &#39;attention saddles&#39; that enables the system to maintain a size-constrained key-value (KV) cache. The cache dynamically stores recent and relevant tokens, reducing memory usage. The framework also implements &#39;attention bias&#39; to capture long-term dependencies. This approach facilitates the inference of texts up to 4 million tokens long in multi-turn conversations and long videos.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Inf-MLLM achieves stable performance on texts with 4 million tokens and multi-round conversations, offering superior streaming reasoning quality and 2x speedup compared to H2O.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.05404v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> To address bandwidth limitations in LLM training.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> This paper introduced DFabric, a two-tier interconnect architecture that enhances data communication across multiple hosts. First, it disaggregates computing units with a CXL fabric at the rack level for efficient intra-rack communication. Second, it disaggregates NICs from hosts into a centralized NIC pool using CXL fabric to enhance communication across racks. To tackle local memory access bottlenecks, DFabric creates a memory pool by disaggregating and expanding host memory.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"> <span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08264v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale</a></p><p class="paragraph" style="text-align:left;">Windows Agent Arena, a benchmark environment for LLMs operating in the Windows OS. It utilizes the OSWorld framework to create 150+ tasks that test agents&#39; planning, screen understanding, and tool usage within real Windows applications. The benchmark is scalable and can be evaluated quickly using parallelized computing on Azure. A new multi-modal agent called Navi is also introduced to showcase the capabilities of the benchmark.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.07587v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Exploring LLMs for Malware Detection: Review, Framework Design, and Countermeasure Approaches</a></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.09030v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Agents in Software Engineering: Survey, Landscape, and Vision</a></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06857v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">What is the Role of Small Models in the LLM Era: A Survey</a></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08846v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">FP-VEC: Fingerprinting Large Language Models via Efficient Vector Addition</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research provides a way <b>to protect the intellectual property of LLMs</b>.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces FP-VEC, a technique that creates a fingerprint vector representing a confidential signature embedded in the LLM. This method enables the integration of the fingerprint into an unlimited number of LLMs simply through vector addition. The approach is lightweight, allowing fingerprinting using CPU-only devices, and maintains the model&#39;s typical behavior while being scalable with a single training process.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.07353v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research propose a safety mechanism for the critical vulnerability of Large Vision-Language Models (LVLMs) against adversarial and jailbreak attacks, which can lead to misleading or harmful outputs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces Sim-CLIP+, a defense mechanism that uses a Siamese architecture to adversarially fine-tune the CLIP vision encoder. The method maximizes cosine similarity between perturbed and clean samples to enhance robustness against adversarial attacks. Sim-CLIP+ is designed as a plug-and-play solution without requiring structural changes to the existing LVLM frameworks and maintains minimal computational overhead.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.13745v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Context-Aware Membership Inference Attacks against Pre-trained Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Membership inference attacks can reveal whether specific data points were part of the training dataset, compromising user privacy.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers developed a novel approach to membership inference attacks tailored to LLMs by leveraging the perplexity dynamics of subsequences within a token sequence. This method departs from traditional loss-based attacks that are inadequate for LLMs. By focusing on context-dependent variations, the researchers adapt statistical tests to uncover how these models memorize and reveal training data.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06927v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Representation Tuning</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research makes models more reliable and honest without needing real-time interventions.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research identifies activation vectors related to honesty in Llama-2-13b-chat. These vectors are added to residual stream activations to control model honesty. The novelty lies in directly fine-tuning these vectors into the model through a dual loss function combining cosine similarity of activations and a standard token-based loss. This method called &#39;representation tuning,&#39; is compared to traditional fine-tuning and online steering.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Fine-tuning vectors using the dual loss function increased honesty more effectively than online steering and generalized better than using a standard token-based loss.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.07503v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs</a> - <b><a class="link" href="https://github.com/Yummy416/AdaPPA?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></b></p><p class="paragraph" style="text-align:left;">Paper develop <b>an adaptive position pre-fill jailbreak attack approach</b> that first uses the LLM to generate safe output content. Then, by exploiting the model&#39;s ability to follow instructions and shift narratives, the attack prompts the LLM to produce harmful content. The approach was evaluated through extensive black-box experiments, targeting a well-known secure model, Llama2, effectively aligning with the model&#39;s differential output-stage alignment protection capabilities.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> This method increased the attack success rate by 47% on Llama2.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08795v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment</a> - Providing feedback in music performance for music lessons</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/451a3fde-370b-4070-a191-12b61e68800a/image.png?t=1727587526"/></div><p class="paragraph" style="text-align:left;">The research proposes LLaQo, a query-based coaching model that evaluates music performances by processing audio data to provide insightful feedback. The system uses instruction-tuned query-response datasets covering music aspects like pitch accuracy and technique. It utilizes AudioMAE encoder and Vicuna-7b as a backend to predict performance ratings, understand contextual aspects like piece difficulty, and answer open-ended questions. LLaQo was trained to align responses with teacher assessments, achieving SOTA in predicting performance ratings.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08596v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions</a> - LLMs in multi-talker speech transcription</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c70971c8-fd40-44c2-909a-3556385108f2/image.png?t=1727587608"/></div><p class="paragraph" style="text-align:left;">The study utilizes WavLM and Whisper encoders to capture detailed speech representations that consider <b>speaker-specific traits and the context.</b> These encoded features are then input into the large language model. The LLM is fine-tuned using the Low-Rank Adaptation (LoRA) method, enhancing its ability to comprehend and transcribe speech amidst multiple speakers. The model is tailored to follow various instructions and respond to differences such as language, speaker characteristics, and keywords.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.07829v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enabling Cost-Effective UI Automation Testing with Retrieval-Based LLMs: A Case Study in WeChat</a> - LLMs for UI automation</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42f3fbba-033e-4621-b8c8-d5c69c5f6f7d/image.png?t=1727587784"/></div><p class="paragraph" style="text-align:left;">The researchers developed CAT, a system combining LLMs and machine learning. CAT uses Retrieval Augmented Generation (RAG) to gather contextual examples for few-shot learning, guiding LLMs in generating actions for UI tasks. These tasks are then mapped onto the UI using machine learning, with LLMs optimizing the process. CAT was tested on WeChat, showcasing its integration into the application’s existing testing platform.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06299v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enhancing Long Video Understanding via Hierarchical Event-Based Memory</a> - Long video understanding in LLMs</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/35a03ccc-335e-4c18-a64b-c6bf2162aeba/image.png?t=1727587856"/></div><p class="paragraph" style="text-align:left;">The researchers propose a Hierarchical Event-based Memory-enhanced LLM (HEM-LLM) to better handle long video understanding by developing an adaptive sequence segmentation scheme that divides long videos into separate events. Each event is then individually modeled to reduce redundancy and preserve key semantics. Information from a prior event is compressed and incorporated while modeling the current event to maintain inter-event dependencies over the long term. The approach aims to improve video comprehension by maintaining contextual connections within events and across the video as a whole.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.06336v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Towards Agentic AI on Particle Accelerators</a></p><p class="paragraph" style="text-align:left;">The research proposes a decentralized, multi-agent framework for the control of particle accelerators, powered by LLMs. The framework envisions a system where autonomous agents are responsible for high-level tasks and communication, with each agent specializing in controlling individual accelerator components. The system allows for self-improvement through experience and human interaction. Two examples are provided to demonstrate the feasibility of this architecture, raising questions about future AI applications in this field and the implications of integrating human feedback into the system.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.15343v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Advertiser Content Understanding via LLMs for Google Ads Safety</a></p><p class="paragraph" style="text-align:left;">Google researchers introduces a method using LLMs to comprehend advertiser intent in relation to content policy violations. It constructs advertiser content profiles by aggregating signals from ads, domains, and targeting information and utilizes LLMs for classification. Additionally, LLMs leverage prior knowledge about advertisers to predict policy violations. Minimal prompt tuning was conducted to enhance model performance.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.08493v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Intelligent LiDAR Navigation: Leveraging External Information and Semantic Maps with LLM as Copilot</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/089b7acf-f87e-4780-9f0a-2aefd9a8df71/image.png?t=1727586882"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research enhance robot navigation by using human-like understanding of LLMs, allowing robots to utilize external and experiential information akin to humans.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study proposes osmAG, a semantic topometric hierarchical map format, to integrate LLMs into robot navigation systems. This involves using the LLM as a copilot, which allows for the assimilation of diverse informational inputs. The methodology bridges the contextual understanding of LLMs with the spatial capabilities of traditional systems like ROS move_base, enhancing the flexibility and adaptability of robotic navigation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11424v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-from-september-9-13-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d06a436a-d55f-405f-a01d-2a9c0f6083ab/image.png?t=1727586926"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The deploys LLMs on resource-constrained embedded devices, improving their performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers designed an FPGA-based accelerator to enhance LLM inference on embedded devices. They used post-training quantization to reduce memory usage and optimized for off-chip memory bandwidth. The architecture features asynchronous computation and full pipelining for efficient matrix-vector multiplication. The experiments evaluated the accelerator using the TinyLlama 1.1B model on a Xilinx ZCU102 platform.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experiments demonstrated a 14.3-15.8x speedup and a 6.1x power efficiency improvement compared to running solely on the ZCU102 processing system.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=f91ee9f1-a021-4c71-954b-532475ac2b8a&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Summary of LLMs related research paper published on 2-4 September, 2024</title>
  <description>From Visual Superpowers to Digital Sleuths: The Latest Breakthroughs Shaping the Future of Large Language Models</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d225975a-55c1-4f52-9731-8d8617432913/sept_2-4.png" length="261283" type="image/png"/>
  <link>https://llm.beehiiv.com/p/summary-llms-related-research-paper-published-24-september-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/summary-llms-related-research-paper-published-24-september-2024</guid>
  <pubDate>Sat, 21 Sep 2024 14:00:00 +0000</pubDate>
  <atom:published>2024-09-21T14:00:00Z</atom:published>
    <dc:creator>LLMs Research</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>LongLLaVA scales up multimodal LLMs to process nearly 1,000 images efficiently</b> - boosting AI&#39;s visual prowess without overloading your GPU.</p></li><li><p class="paragraph" style="text-align:left;"><b>The Single-Turn Crescendo Attack (STCA) reveals LLM vulnerabilities in just one interaction</b> - underlining the need for stronger AI safeguards.</p></li><li><p class="paragraph" style="text-align:left;"><b>SmileyLlama modifies large language models for chemical exploration</b>- bringing a fresh twist to <b>molecular discoveries</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>Mamba serves as a motion encoder for robotic imitation learning</b> - helping robots mimic human movements more effectively.</p></li><li><p class="paragraph" style="text-align:left;"><b>LLMs hypothesize missing causal variables</b>—acting as digital detectives to fill gaps in scientific understanding.</p></li></ol></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.03021v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">CLUE: Concept-Level Uncertainty Estimation for Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/55be281b-9595-45f1-b422-afd69c1f3526/image.png?t=1726888753"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> This paper addresses the limitations of existing uncertainty estimation methods that focus only on sequence-level uncertainty and fail to separately assess the uncertainty of each component in a sequence.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The proposed CLUE framework transforms LLM-generated sequences into concept-level representations, allowing for individual concept uncertainty estimation. This is achieved by breaking down sequences into individual concepts and mathematically measuring uncertainty for each concept separately. The approach allows researchers to conduct more detailed uncertainty assessments, providing a granular understanding of model outputs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02897v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Current long-context LLMs often lack citations in their responses, creating verification challenges and diminishing trustworthiness due to potential hallucinations.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces LongBench-Cite for evaluating LLM citation performance. A novel CoF pipeline was developed to generate QA instances with precise citations. This pipeline was used to create LongCite-45k, a dataset to train LongCite-8B and LongCite-9B models for generating accurate responses with sentence-level citations.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Models trained using LongCite-45k achieved state-of-the-art citation quality, outperforming proprietary models like GPT-4o.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02889v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture</a></p><p class="paragraph" style="text-align:left;"><b>GitHub:</b> <a class="link" href="https://github.com/FreedomIntelligence/LongLLaVA?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow">https://github.com/FreedomIntelligence/LongLLaVA</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7110fc88-27cf-46b1-8d89-281e575aa8d6/image.png?t=1726888864"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Expands multimodal LLMs long-context capabilities</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research developed a hybrid model architecture combining <b>Mamba and Transformer blocks</b>. It incorporated a data construction approach capturing temporal and spatial dependencies across multiple images and used a progressive training strategy. This combination aimed to balance efficiency and effectiveness, enabling the model to process nearly a <b>thousand images efficiently with high throughput and low memory usage.</b></p><p class="paragraph" style="text-align:left;"><b>Results:</b> LongLLaVA achieved competitive results across various benchmarks while maintaining high throughput and low memory consumption. It could process nearly a thousand images on a single A100 80GB GPU, indicating its efficiency and potential applicability.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02727v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/49e9388c-b2d2-4f2d-912f-6d2b0c8f664b/image.png?t=1726888941"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper optimizes LLM-based embedding models by identifying effective pooling and attention mechanisms, potentially improving performance on tasks like text similarity and information retrieval.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study involved <b>a large-scale experiment</b> with LLM-based embedding models. These models were trained on the <b>same dataset and base model</b>, <b>differing only in pooling and attention strategies</b>. The researchers compared various designs, including <b>bidirectional attention</b>, <b>EOS-last token pooling</b>, and more. A new Multi-Layers Trainable Pooling strategy was introduced, transforming outputs from all hidden layers using a cross-attention network. This experiment aimed to evaluate statistically significant differences in performance across diverse tasks.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The findings indicate no universal optimal design: bidirectional attention and trainable pooling excel in text similarity and retrieval, whereas simpler designs suffice for clustering and classification. The proposed Multi-Layers Trainable Pooling showed statistically superior performance in specific tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02976v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Hallucination Detection in LLMs: Fast and Memory-Efficient Finetuned Models</a></p><p class="paragraph" style="text-align:left;">The authors developed a novel methodology to create <b>LLM ensembles capable of detecting hallucinations without extensive computational and memory resources.</b> They focus on optimizing the training and inference processes, ensuring that only a single GPU is needed, making it accessible and pragmatic for practical applications.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02686v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research aims to improve the reasoning abilities of LLMs, especially in mathematical and physics tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study begins by <b>visualizing text generation at the attention and representation level</b> to assess genuine reasoning. It then uses <b>a causal framework to explain these observations.</b> Following this, the researchers introduce <b>Deconfounded Causal Adaptation (DCA)</b>, <b>a parameter-efficient fine-tuning approach</b> designed to enhance reasoning skills by allowing the model to generalize problem-solving abilities across questions. The method involves modifying only 1.2 million parameters, focusing on extracting and applying overarching problem-solving skills.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experiments demonstrate that DCA outperforms existing fine-tuning methods across various benchmarks, achieving better or comparable results while being parameter-efficient.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01666v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">In Defense of RAG in the Era of Long-Context Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research revisits the relevance of retrieval-augmented generation (RAG) amidst the rise of long-context LLMs, arguing the former can yield higher quality answers by maintaining focus on pertinent information.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces an order-preserving retrieval-augmented generation (OP-RAG) mechanism, which addresses the challenges of long-context LLMs by finely tuning the amount of context supplied to the model. Through extensive experiments on public benchmarks, the OP-RAG approach is evaluated by varying the number of retrieved chunks. This method observes that as more chunks are retrieved, answer quality initially increases before decreasing, establishing optimal points for maximal answer quality with fewer tokens.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01552v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Increasing the performance of Black-Box LLMs like GPT-4 is challenging due to inaccessible parameters, making it essential to focus on prompt quality for better human-like responses.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a self-instructed in-context learning framework utilizing reinforcement learning to refine prompt generation, enabling better alignment with original prompts. Derived prompts are crafted to form contextual environments to enhance learning capability. This approach addresses semantic inconsistencies in prompt refinement and reduces discrepancies by using responses from LLMs integrated with derived prompts for a cohesive contextual demonstration.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Extensive experiments show that the method improves derived prompt reliability and significantly enhances the capability of LLMs, including models like GPT-4, to produce more effective responses.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01495v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">The Compressor-Retriever Architecture for Language Model OS</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important because it aims to transform LLMs from simple chatbots into general-purpose agents capable of interacting with the real world, which involves managing life-long context and statefulness.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a compressor-retriever architecture to manage life-long context for LLMs. This model-agnostic architecture <b>compresses and retrieves context using the base model&#39;s forward function</b>, maintaining end-to-end differentiability. Unlike retrieval-augmented generation, this approach focuses on exclusive usage of the base model, ensuring scalability and adaptability to long-context tasks without requiring extensive modifications to existing model structures.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01369v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Imitating Language via Scalable Inverse Reinforcement Learning</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research explores the use of inverse reinforcement learning (IRL) to enhance the imitation learning aspect of language model training, offering potential improvements in diversity and task performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research investigates the application of inverse reinforcement learning (IRL) for fine-tuning large language models. It reformulates <b>inverse soft-Q-learning </b>as a <b>temporal difference</b> regularized extension of <b>maximum likelihood estimation (MLE)</b>, establishing <b>a new connection between MLE and IRL</b>. This approach seeks to balance the complexity of the model with improved performance and diversity in language generation. The study emphasizes extracting rewards and optimizing sequences directly, which could improve the sequential structure of autoregressive generation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01227v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research proposes a more efficient prompt compression method.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces <b>context-aware prompt compression (CPC)</b>, which uses a <b>context-aware sentence encoder</b> providing relevance scores for sentences related to a given question. A new dataset is created containing questions, positive sentences (relevant) and negative sentences (irrelevant), which are used to train the encoder under a contrastive setup aimed at learning sentence representations.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The method considerably outperforms previous prompt compression techniques on benchmark datasets and achieves <b>up to 10.93 times faster inference</b> compared to token-level compression.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01162v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Balancing Performance and Efficiency: A Multimodal Large Language Model Pruning Method based Image Text Interaction</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/81659629-4c72-4099-ad5c-e8e02d6def2a/image.png?t=1726889137"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> High computational costs of multimodal LLMs restrict their practical application.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a dynamic pruning algorithm for visual tokens in multimodal large language models. Firstly, the visual and CLS token similarity curve is analyzed to identify an inflection point, which helps determine a segmentation point for pruning visual tokens. Next, the concatenated visual and textual tokens are pruned again in the LLM layer. This is achieved by filtering out tokens with low text correlation, balancing efficiency and performance by leveraging interactions between visual and textual features.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The proposed method reduces the token quantity by an average of 22% of the original token quantity.</p></div><h3 class="heading" style="text-align:left;" id="ai-strategies-tools-that-will-skyro">AI Strategies & tools that will skyrocket your Marketing ROI by 50% 🚀</h3><p class="paragraph" style="text-align:left;">You don’t realize it yet, but AI has massive potential for you as a marketer.</p><p class="paragraph" style="text-align:left;">This free 3-hour Masterclass on AI & ChatGPT (worth $399) will help you become a master of 20+ AI tools & prompting techniques. <a class="link" href="https://web.growthschool.io/BHJM/?utm_source=beehiiv&utm_medium=email&utm_campaign={{publication_alphanumeric_id}}&_bhiiv=opp_56ff29be-2807-4de6-a37e-e265b968d7a7_51596c0c&bhcl_id=1abce2f1-d4a2-41c3-89bf-c1ad8b9624dd_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Join it now for $0</a></p><div class="image"><img class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e6a34b4b-c84d-4d8f-b832-d92cfb15df24/3.jpg?t=1721830057"/></div><p class="paragraph" style="text-align:left;">This is for you if you work in any vertical of marketing–<b> writing, designer, campaign managing, influencer marketing, growth marketing, etc.</b></p><p class="paragraph" style="text-align:left;">Ready to shock your team with a 10x boost in revenue & campaign performance? 🚀</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://web.growthschool.io/BHJM/?utm_source=beehiiv&utm_medium=email&utm_campaign={{publication_alphanumeric_id}}&_bhiiv=opp_56ff29be-2807-4de6-a37e-e265b968d7a7_51596c0c&bhcl_id=1abce2f1-d4a2-41c3-89bf-c1ad8b9624dd_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Get it now for absolutely free! </a><a class="link" href="https://web.growthschool.io/BHJM/?utm_source=beehiiv&utm_medium=email&utm_campaign={{publication_alphanumeric_id}}&_bhiiv=opp_56ff29be-2807-4de6-a37e-e265b968d7a7_51596c0c&bhcl_id=1abce2f1-d4a2-41c3-89bf-c1ad8b9624dd_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">🎁</a></p><p class="paragraph" style="text-align:left;">You will join 1 Million+ people who have taken this masterclass to learn how to:</p><ul><li><p class="paragraph" style="text-align:left;">Create 100+ content pieces for reels,blogs, from one single long form video</p></li><li><p class="paragraph" style="text-align:left;">Put data tracking & reporting for your campaigns on autopilot</p></li><li><p class="paragraph" style="text-align:left;">Do predictive analysis and optimize your marketing campaigns for better results</p></li><li><p class="paragraph" style="text-align:left;">Personalize customer experiences by leveraging the power of AI</p></li></ul><p class="paragraph" style="text-align:left;">You’ll wish you knew about this FREE AI masterclass sooner (Btw, it’s rated at 9.8/10 ⭐)</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://web.growthschool.io/BHJM/?utm_source=beehiiv&utm_medium=email&utm_campaign={{publication_alphanumeric_id}}&_bhiiv=opp_56ff29be-2807-4de6-a37e-e265b968d7a7_51596c0c&bhcl_id=1abce2f1-d4a2-41c3-89bf-c1ad8b9624dd_{{subscriber_id}}_{{email_address_id}}" target="_blank" rel="noopener noreferrer nofollow">Register & save your seat now! (valid for next 24 hours only!)</a></p><p class="paragraph" style="text-align:left;"></p><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b> LLMOps & GPU level optimization</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.11155v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">ISO: Overlap of Computation and Communication within Sequence For LLM Inference</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> LLM inference efficiency improvement is crucial due to the resource-intensive nature of transformer models and multi-GPU tensor parallelism, which can underutilize computing resources.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a novel strategy for <b>overlapping computation and communication at the sequence level</b> during LLM inference. This method enhances the overlap degree and reduces application constraints, compared to existing techniques that overlap matrix computations and interleave micro-batches.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The proposed technique decreased time consumption during the prefill stage by 35% on 4090 GPU and by roughly 15% on A800 GPU.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02423v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Accelerating Large Language Model Training with Hybrid GPU-based Compression</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the overhead in large language model (LLM) training due to data-intensive communication routines. Improving training efficiency is critical as models grow larger, requiring optimized communication strategies to handle massive data scales without sacrificing accuracy.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research investigates using compression with MPI collectives in distributed LLM training with 3D parallelism and ZeRO optimizations. It tests basic compression across all communication collectives, followed by a refined approach where compression varies by the type of data. Gradients receive aggressive compression due to their low-rank structure, while activations, optimizer states, and model parameters use milder compression to maintain precision.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The naive compression scheme yields a 22.5% increase in TFLOPS per GPU and a 23.6% increase in samples per second for GPT-NeoX-20B training. Using hybrind appraoch team achieved 17.3% increase in TFLOPS per GPU and a 12.7% increase in samples per second while reaching baseline loss convergence.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02026v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Foundations of Large Language Model Compression -- Part 1: Weight Quantization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the challenge of deploying LLMs on resource-constrained devices, reducing computational costs, and minimizing the environmental impact of AI.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper introduces CVXQ, a novel quantization framework derived from convex optimization principles. It allows for post-training quantization of LLMs with hundreds of billions of parameters to any target size. The method leverages convex optimization to efficiently scale and optimize model weights, ensuring performance is preserved while reducing the model&#39;s size significantly. A reference implementation of CVXQ is available for public use.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01143v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">FlashFlex: Accommodating Large Language Model Training over Heterogeneous Environment</a></p><p class="paragraph" style="text-align:left;"><b>Github: </b><a class="link" href="https://github.com/Relaxed-System-Lab/FlashFlex?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow">https://github.com/Relaxed-System-Lab/FlashFlex</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research showcase heterogeneous GPUs based environment which allows more flexible and efficient utilization of computational resources beyond traditional homogeneous setups.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces FlashFlex, a system that supports asymmetric partitioning of training computations across data, pipeline, and tensor model parallelism. It formalizes allocation as a constrained optimization problem and uses a hierarchical graph partitioning algorithm to efficiently distribute tasks across heterogeneous GPUs. This approach adaptively allocates training workloads to fully leverage available computational resources.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> FlashFlex achieves comparable training MFU for LLMs ranging from 7B to 30B parameters on heterogeneous GPUs to state-of-the-art systems using homogeneous high-performance GPUs, with minimal performance gap as low as 0.30%.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01141v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research tries to reduce computation of MoE and attention layers by optimizing resource allocation for high and low arithmetic operations.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces Duplex, a device that combines xPU tailored for high-Op/B processes and Logic-PIM designed for low-Op/B operations. The system dynamically selects the appropriate processor for each layer of LLMs based on their arithmetic intensity. Logic-PIM enhances data transmission by adding through-silicon vias (TSVs) for high-bandwidth communication, enabling efficient handling of variable Op/B operations found in LLM&#39;s MoE and attention layers.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.00918v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> During LLM training, current GPU cluster setups face challenges with memory and bandwidth, which hinder efficient scaling.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces LuWu, an innovative in-network optimizer that allows 100B-scale model training by offloading optimizer states and parameters from GPU workers onto an in-network optimizer node. It also shifts collective communication from GPU-NCCL to SmartNIC-SmartSwitch co-optimization, reducing interference and CPU overhead. This setup retains the communication pattern of model-sharded data parallelism while ensuring efficient resource utilization.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> LuWu achieves a 3.98x improvement over state-of-the-art systems when training a 175B model on an 8-worker cluster.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02081v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Physical Rule-Guided Convolutional Neural Network</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fa3c23f1-64dc-43d9-856a-6d1cfd61735e/image.png?t=1726888353"/></div><p class="paragraph" style="text-align:left;">This paper introduces Physics-Guided CNNs (PGCNNs), incorporating dynamic, trainable, and automated LLM-generated physical rules into CNN architecture. These rules are integrated as custom layers, addressing challenges such as limited data and low confidence scores. The architecture enhances the CNN model by enabling it to utilize domain-specific knowledge dynamically during training and inference.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.04465v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Here&#39;s Charlie! Realising the Semantic Web vision of Agents in the age of LLMs</a> - Paper develops semi-autonomous web agents that consult users only if lacking context or confidence. A demonstration is provided using Notation3 rules to ensure safety in belief, data sharing, and usage. LLMs are integrated for natural language interaction between users and agents, enabling user training on preferences and serendipitous agent dialogues.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.03093v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Multi-language Unit Test Generation using LLMs</a> - Let’s ask LLMs to generate <span style="color:#000000;font-family:Arial, sans-serif;font-size:medium;">compilable and high-coverage unit tests! </span></p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02711v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Creating a Gen-AI based Track and Trace Assistant MVP (SuperTracy) for PostNL</a> - LLMs in customer service and logistics to improve communication in parcel tracking. LLMs system autonomously manage user inquiries and improve knowledge handling regarding parcel journeys. </p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02604v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Hypothesizing Missing Causal Variables with LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a11d76eb-7944-4765-b1f8-8828b9ea3c1f/image.png?t=1726888259"/></div><p class="paragraph" style="text-align:left;">This paper explores enhancing the scientific discovery process by leveraging LLMs to hypothesize missing variables in causal relationships, a pivotal task for advancing scientific knowledge efficiently. Researchers used partial causal graphs with missing variables and tasked LLMs with hypothesizing about these missing elements. </p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02231v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration</a> - The researchers modified the Llama LLM, an open-source model, using supervised fine-tuning (SFT) and direct preference optimization (DPO). These methods adapted the model to handle chemical SMILES string data, enabling it to generate molecules based on desired properties. This involved training the LLM to respond to specific prompts and provide functional outputs for chemistry and materials applications.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02636v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Mamba as a motion encoder for robotic imitation learning</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a6c147e7-b2dc-4f94-993b-972042d8389f/image.png?t=1726888115"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research introduces Mamba as an effective tool for robotic imitation learning, aiming to enhance robots&#39; dexterity and adaptability through improved contextual information capture.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study proposes the use of Mamba, a novel architecture akin to an encoder, for robotic imitation learning. Mamba operates similarly to an autoencoder by reducing the dimensionality of the state space while preserving essential temporal dynamics for effective motion prediction. It compresses sequential data into state variables, facilitating superior performance in practical tasks. The research includes experimental evaluations on tasks such as cup placing and case loading.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experimental results indicate that Mamba, despite higher estimation errors, achieves superior success rates compared to Transformers in task execution.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01133v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Large Language Models Can Understanding Depth from Monocular Images</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Monocular depth estimation is vital for computer vision, with applications in robotics, autonomous driving, and augmented reality. Leveraging large language models&#39; capabilities can enhance depth estimation, providing efficient solutions across these domains.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces LLM-MDE, a framework using LLMs for depth estimation from monocular images. It applies cross-modal reprogramming to align vision data with text, and an adaptive prompt estimation module that generates automatic prompts for the LLM based on image input. Both strategies integrate with a pretrained LLM to improve depth estimation tasks with minimal supervision. The framework is validated through experiments on real-world monocular depth estimation datasets.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01990v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Contemporary Model Compression on Large Language Models Inference</a> - The survey explores contemporary model compression techniques to make LLMs memory efficient and to achieve faster computation.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02691v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LLM-Assisted Visual Analytics: Opportunities and Challenges</a> - This is a comprehensive survey of current directions in integrating LLMs into Visual Analytics (VA) systems.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02977v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Large Language Model-Based Agents for Software Engineering: A Survey</a> - This paper collects and analyzing 106 papers related to LLM-based agents in SE.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.02718v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Alignment-Aware Model Extraction Attacks on Large Language Models</a> - The researchers developed a novel attack algorithm called Locality Reinforced Distillation (LoRD). LoRD leverages a policy-gradient-style training task that uses the responses from victim models to guide preference crafting for the local model. Theoretical analysis ensures that LoRD&#39;s convergence aligns with the LLM&#39;s training alignments, and it reduces query complexity while countering watermark protections. Extensive experiments were conducted on various state-of-the-art LLMs to evaluate the effectiveness of the algorithm.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.03131v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)</a> - Research exposes vulnerabilities in LLMs by highlighting <b>how adversarial inputs can provoke inappropriate outputs</b>, emphasizing the need for stronger safeguards in responsible AI. The research introduces STCA, which combines the escalating techniques of previous multi-turn attacks into a single potent prompt. This method condenses escalation into one interaction, effectively bypassing LLM moderation filters designed to block harmful responses. The approach was developed to showcase the weaknesses in existing filter mechanisms of LLMs, which typically monitor dialogues over several interactions.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01630v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems</a> - The research introduced SafeEmbodAI, a framework composed of secure prompting, state management, and safety validation mechanisms to better integrate mobile robots in embodied AI systems. It aims to secure and assist LLMs in processing multi-modal data and validating robotic behaviors through comprehensive safety checks. This framework&#39;s effectiveness is tested through a newly designed metric focusing on mission-oriented exploration in simulated environments.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2409.01380v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-paper-published-on-2-4-september-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Membership Inference Attacks Against In-Context Learning</a> - The researchers developed the first membership inference attack specifically for ICL, using only generated texts. They crafted four attack strategies optimized for different restrictive scenarios and conducted experiments on four major LLMs. The study also evaluates a hybrid attack strategy that combines the strengths of the four individual strategies. Furthermore, they examined three defence mechanisms focusing on data handling, instructional adjustments, and output modifications to mitigate privacy risks. The attack strategies achieved a 95% accuracy advantage in most cases.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=f89e6ce0-6d50-46cf-ad61-d77e94c6c967&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published between 26th august to 1st September</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a89076fd-3097-425b-80b6-8118f29398f1/26th_aug_to_1st_sept.png" length="291017" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-26th-august-1st-september</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-26th-august-1st-september</guid>
  <pubDate>Mon, 16 Sep 2024 11:00:00 +0000</pubDate>
  <atom:published>2024-09-16T11:00:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><h2 class="heading" style="text-align:center;">🔑<b> takeaway from today’s newsletter</b></h2><ul><li><p class="paragraph" style="text-align:left;"><b>From Ambiguous to Accurate: </b>LLMs Now Tackle Vague Questions with Sharp Precision!</p></li><li><p class="paragraph" style="text-align:left;"><b>Guardians of the Algorithm: </b>New Shields Protect AI from Sneaky Prompt Hacks and Data Poisoning!</p></li><li><p class="paragraph" style="text-align:left;"><b>AI Gets a Creative Spark: </b>Mixing Logic with LLMs for Stories You&#39;ve Never Dreamed Of!</p></li><li><p class="paragraph" style="text-align:left;"><b>Overconfidence Overruled: </b>Entropic Steering Makes AI Agents Explore More and Guess Less!</p></li><li><p class="paragraph" style="text-align:left;"><b>Cracking the Cultural Code: </b>Can LLMs Master Multilingual Chats and Hidden Contexts?</p></li></ul></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00509v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Empirical influence functions to understand the logic of fine-tuning</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2788f99b-0c8d-4a3f-9a14-78678300e22d/image.png?t=1726429507"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Improving and interpreting the learning process in neural networks is crucial for enhancing their performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers investigated how individual training samples influence the output of neural networks during fine-tuning. They measured empirical influences and evaluated these influences against criteria such as decreasing influence with semantic distance, sparseness, noise invariance, transitive causality, and logical consistency. The investigation was conducted on both simple convolutional networks and a modern LLM. Additionally, they explored the impact of prompts in potentially mitigating the shortcomings found in these models when they fail to meet the given criteria.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study found that popular models, including simple convolutional networks and modern LLMs, fail to meet the desiderata for influences. However, prompting showed potential in partially rescuing this failure, offering a practical way to quantify how well neural networks learn from fine-tuning stimuli.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00284v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">A Closer Look at Logical Reasoning with LLMs: The Choice of Tool Matters</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding how different symbolic solvers impact LLMs logical reasoning could help improve their performance on complex tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers integrated LLMs with three different symbolic solvers: Z3, Pyke, and Prover9. They then evaluated the combined systems performance on three logical reasoning datasets: ProofWriter, PrOntoQA, and FOLIO. The integration involved designing interfaces for LLMs to communicate effectively with the solvers. Extensive experimentation measured the impact of each solver on the LLMs ability to solve logical reasoning tasks. Performance was assessed based on accuracy and the number of questions each solver executed successfully.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Z3&#39;s overall accuracy performance slightly surpasses Prover9, but Prover9 executed more questions successfully. Pyke&#39;s performance was significantly inferior to both Z3 and Prover9.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00244v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Controlling Large Language Model Agents with Entropic Activation Steering</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7bff3b70-aca2-46b9-a167-11021d9f8bee/image.png?t=1726430114"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research tackles the problem of overconfidence in LLM agents, which is critical for improving their decision-making and explorative behaviors.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Experiments were conducted in controlled sequential decision-making tasks to observe LLM behavior. Researchers identified that <b>LLM agents display overconfidence due to a collapse in the entropy of the action distribution when sampling.</b> Traditional token-level sampling methods were found insufficient for promoting exploration. Therefore, an activation steering method called <b>Entropic Activation Steering (EAST) </b>was introduced. EAST works by calculating a steering vector as an entropy-weighted combination of representations and then manipulating the LLM&#39;s activations during the forward pass to increase uncertainty and exploration in the actions.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> EAST was able to reliably increase the entropy in LLM agent actions, leading to enhanced explorative behavior. It also allows for better control and interpretation of how LLM agents represent uncertainty.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.01633v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots</a></p><p class="paragraph" style="text-align:start;">LLMs often stumble when faced with <b>under-specified queries</b>—questions that lack sufficient detail. This leads to poor or irrelevant responses. Researchers analyzed public chat logs and found this issue to be widespread. To address it, they modeled the problem using <b>Partially Observed Decision Processes (PODPs)</b>. Think of PODPs as a way to make optimal decisions when you don&#39;t have all the information—a bit like navigating a foggy road with limited visibility.</p><p class="paragraph" style="text-align:start;">By framing the chatbot&#39;s decision-making process as a PODP, they could derive improved policies that guide the LLM to ask clarifying questions or provide more useful answers despite the ambiguity. They then recalibrated LLMs using <b>learned control messages</b> based on these policies, effectively teaching the models to handle vagueness better.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00240v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Exploring Vulnerabilities and Protections in Large Language Models: A Survey</a></p><p class="paragraph" style="text-align:left;">LLMs are vulnerable to <b>prompt hacking</b> and <b>adversarial attacks</b>, which can manipulate them into producing harmful or misleading content. For instance:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Prompt Injection Attacks</b>: Crafting inputs that trick the model into revealing confidential information or behaving undesirably.</p></li><li><p class="paragraph" style="text-align:left;"><b>Jailbreaking Attacks</b>: Bypassing the model&#39;s safety protocols to generate prohibited content.</p></li><li><p class="paragraph" style="text-align:left;"><b>Data Poisoning and Backdoor Attacks</b>: Injecting malicious data during training so the model behaves incorrectly when triggered.</p></li></ul><p class="paragraph" style="text-align:start;">A comprehensive survey mapped out these vulnerabilities and highlighted mitigation strategies. Defenses include:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Input Sanitization</b>: Filtering or rephrasing user inputs to remove malicious patterns.</p></li><li><p class="paragraph" style="text-align:left;"><b>Robust Training</b>: Exposing models to adversarial examples during training to make them resilient.</p></li><li><p class="paragraph" style="text-align:left;"><b>Monitoring and Detection</b>: Implementing systems to detect unusual model behaviors indicative of an attack.</p></li></ul><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00548v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">LIDAO: Towards Limited Interventions for Debiasing (Large) Language Models</a></p><p class="paragraph" style="text-align:left;">Bias in LLM outputs can perpetuate stereotypes and unfairness. Traditional debiasing methods often degrade the model&#39;s language fluency. Enter <b>LIDAO</b> (Limited Interventions for Debiasing AI Outputs), a framework based on <b>information theory</b>. Instead of overhauling the entire model, LIDAO makes minimal adjustments to reduce bias.</p><p class="paragraph" style="text-align:start;">Imagine tuning a radio to eliminate static without changing the station. LIDAO fine-tunes the model&#39;s outputs to maintain clarity and expressiveness while ensuring fairness. Tests on models ranging from 0.7 to 7 billion parameters showed that LIDAO effectively debiased outputs even in adversarial scenarios designed to provoke biased responses.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.06558v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enhancing Text Authenticity: A Novel Hybrid Approach for AI-Generated Text Detection</a></p><p class="paragraph" style="text-align:left;">With AI-generated content proliferating, distinguishing it from human-written text is crucial to combat misinformation. Researchers developed a <b>hybrid detection method</b> combining:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Traditional TF-IDF Techniques</b>: Analyzing word importance based on frequency.</p></li><li><p class="paragraph" style="text-align:left;"><b>Advanced Machine Learning Models</b>: Utilizing classifiers like Bayesian models, Stochastic Gradient Descent (SGD), Categorical Gradient Boosting (CatBoost), and multiple instances of <b>DeBERTa-v3-large</b> models.</p></li></ul><p class="paragraph" style="text-align:start;">By merging classic text analysis with deep learning, this approach enhances detection accuracy. Extensive testing confirmed its superiority over existing methods, aiding efforts to maintain content authenticity.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00380v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">The Best of Both Worlds: Toward an Honest and Helpful Large Language Model</a></p><p class="paragraph" style="text-align:left;">LLMs sometimes provide confident answers even when unsure, which can mislead users. To promote honesty without sacrificing helpfulness, researchers introduced <b>HoneSet</b>, a dataset of 930 queries across six categories designed to evaluate model honesty.</p><p class="paragraph" style="text-align:start;">They proposed two methods:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Training-Free Approach</b>: Using <b>curiosity-driven prompting</b> to encourage the model to express uncertainty naturally.</p></li><li><p class="paragraph" style="text-align:left;"><b>Fine-Tuning Approach</b>: A two-stage training process where the model first learns to distinguish honest from dishonest responses, then focuses on being helpful while maintaining honesty.</p></li></ol><p class="paragraph" style="text-align:start;">These methods teach LLMs that it&#39;s acceptable to admit uncertainty—much like a knowledgeable person saying, &quot;I&#39;m not sure, but I can find out.&quot; Experiments on nine different LLMs showed improved trustworthiness in their responses.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.04370v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Large Language Model Confidence Estimation via Black-Box Access</a></p><p class="paragraph" style="text-align:left;">Knowing how confident an LLM is in its response helps users assess reliability. Researchers developed a framework to estimate this confidence using:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Engineered Features</b>: Extracting interpretable data from the model&#39;s outputs, such as response length, specificity, and use of hedging words.</p></li><li><p class="paragraph" style="text-align:left;"><b>Logistic Regression Model</b>: Training a statistical model to predict confidence levels based on these features.</p></li></ul><p class="paragraph" style="text-align:start;">By treating the LLM as a <b>black box</b>, they could apply this method without altering the model&#39;s architecture. Tests using benchmark datasets like <b>TriviaQA</b>, <b>SQuAD</b>, <b>CoQA</b>, and <b>Natural Questions</b> demonstrated that this approach effectively indicates when the model&#39;s answers can be trusted.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00554v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Guiding and Diversifying LLM-Based Story Generation via Answer Set Programming</a></p><p class="paragraph" style="text-align:left;">LLMs excel at generating coherent text but often lack diversity in storytelling. To tackle this, this paper introduced a hybrid approach that combines LLMs with <b>Answer Set Programming (ASP)</b>, a form of symbolic logic programming. By using ASP to generate abstract story structures, the LLMs are guided to produce narratives with greater diversity. This method results in stories that are more varied and creative compared to those generated by LLMs alone, highlighting the benefits of integrating symbolic methods for guiding content generation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00522v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0503e837-3e72-49c2-99d5-3c97639f2018/image.png?t=1726459965"/></div><p class="paragraph" style="text-align:left;">Bridging the gap between speech processing and language models, <b>Wav2Prompt</b> presents an end-to-end approach for integrating spoken input with LLMs. It trains on speech data to learn continuous speech representations, which serve as prompts for the LLM. This method aligns speech and text using a <b>continuous integrate-and-fire mechanism</b>, allowing the model to handle tasks like speech translation and spoken-query-based QA effectively. Notably, Wav2Prompt outperforms traditional cascaded models in few-shot learning scenarios, demonstrating significant improvements in tasks such as English-French speech translation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.07572v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Domain-specific ReAct for physics-integrated iterative modeling: A case study of LLM agents for gas path analysis of gas turbines</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0c33bbfc-db0e-4738-887e-d79d9cd4b698/image.png?t=1726461279"/></div><p class="paragraph" style="text-align:left;">Exploring the capabilities of LLMs in specialized domains, researchers developed a domain-specific <b>ReAct</b> framework for physics-integrated iterative modeling, focusing on <b>gas path analysis of gas turbines.</b> By setting up a dual-agent tool-calling process, the LLMs integrate reasoning with expert knowledge and industry tools. While smaller models struggled, larger LLMs showed promising abilities in handling complex, multi-component problems when fine-tuned appropriately. This underscores the potential of large-scale LLMs in specialized engineering tasks, provided they are equipped with domain knowledge and proper prompt design.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.01631v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">An LLM-based Recommender System Environment</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/994bc39d-60ef-48c0-b789-140bc3986103/image.png?t=1726461823"/></div><p class="paragraph" style="text-align:left;">Training recommender systems often suffers from limited online data and evaluation challenges. An innovative solution is the creation of an <b>LLM-based synthetic environment</b> that simulates human behavior for training and evaluating recommender systems using reinforcement learning (RL). This framework allows for more comprehensive testing and development of recommendation algorithms by providing a controlled environment that mimics user interactions, as demonstrated in experiments with movie and book recommendations.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.06556v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End Approach</a></p><p class="paragraph" style="text-align:left;">Generating presentation slides from lengthy documents can be labor-intensive. A multi-staged end-to-end model combines LLMs with <b>Vision-Language Models (VLMs)</b> to automate this process. The LLM first extracts and summarizes key content, while the VLM identifies and incorporates relevant visual elements. Through iterative refinement, the model produces cohesive and visually appealing presentations, outperforming state-of-the-art prompting methods in both automated metrics and human evaluations.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00333v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">A Practice-Friendly Two-Stage LLM-Enhanced Paradigm in Sequential Recommendation</a></p><p class="paragraph" style="text-align:left;">In sequential recommendation systems (SRS), incorporating LLMs can be inefficient, especially with sparse textual data. The <b>TSLRec</b> paradigm addresses this with a two-stage approach. Initially, it performs user-level supervised fine-tuning to inject collaborative filtering information using a pre-trained SRS model. Then, it enhances item representations by generating embeddings that merge this collaborative information with the LLM&#39;s inference capabilities. This method has been validated on benchmark datasets, showing improved efficiency and effectiveness in recommendations.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.07571v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in Classrooms</a></p><p class="paragraph" style="text-align:left;">Self-reflection enhances learning, but scaling personalized reflection activities is challenging. Randomized field experiments in undergraduate courses demonstrated that LLM-guided reflection can significantly improve student confidence and exam performance. By providing tailored reflection prompts, LLMs help students consolidate knowledge more effectively than traditional methods like reviewing lecture slides, indicating a valuable role for LLMs in educational settings.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00247v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Large Language Models for Relevance Judgment in Product Search</a></p><p class="paragraph" style="text-align:left;">Accurate relevance judgment is essential for effective product search. Researchers explored techniques for fine-tuning LLMs to automate the assessment of <b>query-item pairs (QIPs)</b>. By optimizing hyperparameters and experimenting with different attribute concatenation methods and prompting strategies, they achieved relevance annotations comparable to human evaluators. This advancement has immediate applications in improving product search engines by automating a traditionally labor-intensive process.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00430v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Evaluating Uncertainty-based Failure Detection for Closed-Loop LLM Planners</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/45918996-06cd-4f7b-9f74-40ad27c4d8d5/image.png?t=1726461930"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important as it seeks to enhance the reliability of LLM-based closed-loop planning in robotic manipulation tasks by detecting and addressing planning failures.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors introduce KnowLoop, a framework that incorporates an uncertainty-based failure detection mechanism in closed-loop LLM-based planning. The framework quantifies uncertainty using three methods: token probability, entropy, and self-explained confidence. These methods are evaluated using specific prompting strategies. A custom dataset with various manipulation tasks and an LLM-based robot system is used to test these metrics. The effectiveness of the metrics is measured by filtering out uncertain predictions and actively seeking human help when thresholds are exceeded.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The research demonstrates that token probability and entropy are more reflective metrics compared to self-explained confidence. By setting an appropriate threshold to filter out uncertain predictions, the accuracy of failure detection can be significantly enhanced, thus boosting the effectiveness of closed-loop planning and the overall success rate of tasks.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"> <span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00507v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> To determine the most effective iterative refinement approach in text summarization for improving LLM performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers designed two strategies, Prompt Chaining and Stepwise Prompt, for iterative refinement in text summarization. They implemented Prompt Chaining by using three distinct prompts to sequentially draft, critique, and refine the text. For the Stepwise Prompt method, these three phases were integrated within a single prompt to simulate the refinement process. The effectiveness of each method was evaluated through extensive experiments comparing the quality of the generated summaries.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experimental results show that the prompt chaining method can produce a more favorable outcome. The stepwise prompt might produce a simulated refinement process, according to various experiments. These insights could be extrapolated to other applications, contributing to the broader development of LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00343v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Beyond Metrics: Evaluating LLMs&#39; Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Evaluating LLMs in multilingual and culturally nuanced scenarios is crucial for effective real-world applications.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Researchers conducted a performance evaluation of seven LLMs (Mistral-7b, Mixtral-8x7b, GPT-3.5-Turbo, Llama-2-70b, Gemma-7b, GPT-4, and GPT-4-Turbo) on sentiment analysis tasks using a dataset comprising code-mixed WhatsApp chats with Swahili, English, and Sheng. They employed both quantitative metrics like F1 score and qualitative assessments by analyzing the explanations provided by each LLM for its predictions. The study focused on linguistic and contextual nuances, decision-making transparency, and cultural understanding, comparing the models&#39; effectiveness in these areas.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.06555v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">An Evaluation Benchmark for Autoformalization in Lean4</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important because autoformalization can vastly improve the efficiency and accuracy of formalizing mathematical proofs, thereby advancing scientific research and development.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduced a new evaluation benchmark specifically designed for <b>Lean4, a mathematical programming language</b>. The benchmark was then applied to test the autoformalization capabilities of state-of-the-art LLMs, including GPT-3.5, GPT-4, and Gemini Pro. The evaluation involved analyzing the performance of these LLMs on various mathematical tasks, particularly focusing on more complex areas of mathematics to gauge their autoformalization potential comprehensively.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The comprehensive analysis revealed that, despite recent advancements, these LLMs still exhibit limitations in autoformalization, particularly in more complex areas of mathematics. These findings underscore the need for further development in LLMs to fully harness their potential in scientific research and development.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00257v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-between-26th-august-to-1st-september" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research is important because it explores the capabilities and limitations of LVLMs in chart comprehension and reasoning, a nuanced area that demands both vision-language reasoning and understanding of data visualizations.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers conducted a comprehensive evaluation of LVLMs, including models such as GPT-4V and Gemini, focusing on four major chart reasoning tasks: chart question answering, chart summarization, fact-checking with charts, and general chart comprehension. The evaluation included both quantitative assessments and qualitative analysis across a diverse range of chart types. The qualitative evaluation was aimed at identifying specific strengths and weaknesses of LVLMs. Techniques used involved assessing the fluency and accuracy of generated text, and identifying issues such as hallucinations, factual errors, and data bias.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The findings reveal that LVLMs demonstrate impressive abilities in generating fluent texts covering high-level data insights while also encountering common problems like hallucinations, factual errors, and data bias. The study highlights the key strengths and limitations of chart comprehension tasks, offering insights for future research.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=98df13ca-8431-45c9-8ac8-83a5776e8b7b&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Summary of LLMs related research papers published on 30th May, 2024</title>
  <description>Detailed categorization of research trend to stay informed about latest research in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3fe08579-b7e2-4421-ae9e-35f126c73565/29th_may.png" length="468905" type="image/png"/>
  <link>https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-30th-may-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-30th-may-2024</guid>
  <pubDate>Mon, 22 Jul 2024 15:45:33 +0000</pubDate>
  <atom:published>2024-07-22T15:45:33Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19732v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Two Optimizers Are Better Than One: LLM Catalyst Empowers Gradient-Based Optimization for Prompt Tuning</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important as it explores the integration of gradient-based optimization with LLM-based optimization to enhance the efficacy of solving complex non-convex optimization problems in prompt tuning.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research proposes a collaborative optimization approach that combines gradient-based and LLM-based optimizers in an interleaved manner. The team begins by using a gradient-based optimizer to make locally optimal updates at each step, akin to a diligent doer. Alongside, they introduce an LLM-based optimizer that acts as a high-level instructor by inferring solutions from natural language instructions. Task descriptions and optimization trajectories recorded during the gradient-based optimization process are utilized to instruct the LLMs. The inferred results from these LLMs serve as restarting points for subsequent stages of gradient optimization. This iterative process leverages both the rigorous local updates of gradient-based methods and the high-level deductive reasoning capabilities of LLMs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The combined optimization method consistently yields improvements over competitive baseline prompt tuning methods, demonstrating the synergistic effect of conventional gradient-based optimization and the inference ability of LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20339v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Visual Perception by Large Language Model&#39;s Weights</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research addresses the high computational costs involved in multimodal LLMs by proposing a more efficient method for integrating visual information.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers introduced VLoRA, a new approach leveraging parameter space alignment instead of the conventional input space alignment. They utilized a vision encoder to extract visual features from images, which were then converted into perceptual weights. These perceptual weights were subsequently merged with the LLM&#39;s weights. This approach eliminated the need for visual tokens in the input sequence, thereby reducing its length and enhancing computational efficiency. The primary innovation was the perceptual weights generator that transformed visual features into low-rank perceptual weights, aligning with the structure similar to LoRA (Low-Rank Adaptation). The experimental validation was conducted on various benchmarks for multimodal LLMs to establish the effectiveness of this method.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20314v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Improving the efficiency of speculative decoding in low-memory GPUs is crucial for enabling faster LLM inference on less powerful hardware.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers proposed an approach called Skippy Simultaneous Speculative Decoding (S3D), which leverages simultaneous multi-token decoding combined with mid-layer skipping. This method aims at minimizing the memory overhead while still providing substantial speedup for LLM inference. Multi-token decoding enables the generation of multiple tokens in one step, thus speeding up the decoding process, while mid-layer skipping reduces the computational load by skipping unnecessary layers during the inference. The researchers designed S3D with minimal architectural changes and without requiring extensive new training data. They utilized the memory efficiency provided by S3D to create a high-performing yet smaller SD model based on the Phi-3 architecture and compared its performance against recent open-source SD systems, particularly focusing on the quantized EAGLE model, analyzing speed and memory usage in half-precision mode.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> When compared against recent effective open-source SD systems, the proposed method has achieved one of the top performance-memory ratios while requiring minimal architecture changes and training data. The smaller yet more effective SD model based on Phi-3 was found to be 1.4 to 2 times faster than the quantized EAGLE model and operates in half-precision while using less VRAM.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20252v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the critical dependency of LLM outputs on prompt design, enhancing their effectiveness without human intervention or task-specific training.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers proposed a novel approach named Hierarchical Multi-Agent Workflow (HMAW) that leverages multiple LLMs in a hierarchical structure to autonomously design and optimize prompts. HMAW divides the process into two main steps. First, one or more LLMs collaboratively construct an optimal prompt by refining instructions and wording through multiple iterations. Second, the final prompt generated by the hierarchy is used by another LLM to answer the user&#39;s query. This method is designed to be task-agnostic, meaning it does not require tuning to specific tasks nor does it involve pre-training or task-specific data. The efficacy of HMAW was evaluated through both quantitative and qualitative experiments across various benchmarks, demonstrating its ability to boost LLM performance without significant overhead or complexity.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20175v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">InstructionCP: A fast approach to transfer Large Language Models into target language</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research addresses the challenge of adapting English-focused LLMs to other languages while maintaining their conversational and content filtering abilities.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose a method called Instruction Continual Pre-training (InsCP), which incorporates instruction tags into the continual pre-training (CP) process. This approach aims to prevent the degradation of conversational abilities and content filtering effectiveness when transferring LLMs to new languages. The process starts by collecting 0.1 billion tokens of high-quality instruction-following data in the target language. These tokens contain specific instruction tags that guide the model in maintaining its conversational proficiency. The InsCP method then integrates these tokens into a modified CP regimen, ensuring that language adaptation is achieved without compromising other essential functionalities of the LLM. The efficacy of InsCP is validated through empirical evaluations on benchmarks for language alignment, reliability, and knowledge retention, ensuring that the model performs comparably to its original version while effectively responding in the target language.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Empirical evaluations confirm that InsCP retains conversational and Reinforcement Learning from Human Feedback (RLHF) abilities. The method requires only 0.1 billion tokens of high-quality instruction-following data, thereby reducing resource consumption.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20089v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM Abilities</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> To improve translation quality while preserving desirable abilities in large language models (LLMs).</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research involves an extensive evaluation of the translation capabilities of the LLaMA and Falcon families of models, which range from 7 billion to 65 billion parameters. The authors fine-tuned these models on parallel data to observe its impact on various desirable LLM behaviors, such as steerability, inherent document-level translation, and the ability to produce less literal translations. During this process, they assessed how these abilities degraded or improved post fine-tuning. To address the degradation of certain abilities, the researchers incorporated monolingual data as part of the fine-tuning process. This strategy aimed to balance between enhancing translation quality and maintaining other beneficial LLM properties.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study found that while fine-tuning improved general translation quality, it led to a decline in formality steering, technical translation through few-shot examples, and document-level translation abilities. However, models produced less literal translations. By including monolingual data in the fine-tuning process, these abilities were better preserved while still enhancing overall translation quality.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19893v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research aims to improve Retrieval Augmented Generation (RAG) by addressing the limitations of relying solely on similarity between queries and documents, which can lead to performance degradation in knowledge-intensive tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces MetRag, a multi-layered thought framework that enhances Retrieval Augmented Generation. MetRag incorporates three layers of thoughts: similarity-oriented, utility-oriented, and compactness-oriented thoughts. </p><p class="paragraph" style="text-align:left;">1. <b>Similarity-Oriented Thought</b>: This conventional approach retrieves documents based on similarity metrics to the given query. </p><p class="paragraph" style="text-align:left;">2. <b>Utility-Oriented Thought</b>: A small-scale utility model, supervised by an LLM, evaluates the usefulness of the retrieved documents, adding a layer of relevance beyond mere similarity. </p><p class="paragraph" style="text-align:left;">3. <b>Compactness-Oriented Thought</b>: To deal with the often large set of retrieved documents, an LLM is utilized as a task-adaptive summarizer. </p><p class="paragraph" style="text-align:left;">This step aims to capture commonalities and characteristics within the document set, making the information more compact and manageable.multi-layered thoughts are then combined in the final stage where an LLM generates the knowledge-augmented response. This comprehensive approach aims to tackle the shortcomings of existing RAG methods by integrating not just similarity and relevance, but also conciseness and coherence in the generation process.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19888v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Parrot: Efficient Serving of LLM-based Applications with Semantic Variable</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research aims to enhance the efficiency of LLM-based applications by addressing limitations in current public LLM services, which lead to sub-optimal end-to-end performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces Parrot, a service system designed to improve the end-to-end performance of LLM-based applications. Parrot employs a novel concept called &#39;Semantic Variable,&#39; which serves as a unified abstraction to expose application-level knowledge to public LLM services. Semantic Variables annotate input/output variables in the prompts of requests and create data pipelines when linking multiple LLM requests. This allows for a more integrated workflow and helps uncover correlations across multiple LLM requests through data flow analysis. By performing such an analysis, Parrot can identify optimization opportunities that would not be apparent when handling individual LLM requests separately.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Extensive evaluations demonstrate that Parrot can achieve up to an order-of-magnitude improvement for popular and practical use cases of LLM applications.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19877v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">KNOW: A Real-World Ontology for Knowledge Capture with Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important as it aims to create a structured ontology that enhances LLMs&#39; ability to handle real-world generative AI tasks, especially for applications like personal AI assistants.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> To create the KNOW ontology, the researchers focused on a pragmatic approach that started with established human universals, specifically spacetime (places and events) and social structures (people, groups, organizations). The ontology is designed to capture the everyday knowledge that LLMs can use to augment various applications. The team compared and contrasted KNOW with previous projects like <a class="link" href="https://Schema.org?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow">Schema.org</a> and Cyc, noting how LLMs already internalize much of the commonsense knowledge that older projects painstakingly gathered over decades. They emphasize simplicity and usability by making software libraries available in 12 popular programming languages. These libraries enable developers to leverage the ontology concepts directly in their software engineering efforts, promoting better AI interoperability and enhancing the developer experience. Key criteria for defining concepts in the ontology include universality and utility, ensuring that only the most pragmatically valuable concepts are included.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19874v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Is In-Context Learning Sufficient for Instruction Following in LLMs?</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding how effectively in-context learning (ICL) aligns with long-context language models (LLMs) for following instructions is important for enhancing their performance in real-world tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers began by evaluating the performance of LLMs utilizing URIAL, a method that leverages three in-context examples to align base LLMs, and compared its performance to instruction fine-tuning on standard benchmarks such as MT-Bench and AlpacaEval 2.0 (LC). They observed that adding more ICL demonstrations did not systematically improve instruction following performance for long-context LLMs. To improve this, they introduced a greedy selection approach for ICL examples, which aimed at optimizing the choice of examples to enhance performance. Additionally, the researchers conducted ablation studies to better understand why there was a performance gap between ICL alignment and instruction fine-tuning, specifically looking into aspects of ICL that are unique to the instruction tuning context.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> While effective, ICL alignment with URIAL underperforms compared to instruction fine-tuning, especially with more capable base LMs. Adding more ICL demonstrations does not systematically improve instruction following performance. A greedy selection approach for ICL examples improves performance but does not bridge the gap to instruction fine-tuning. Ablation studies reveal specific reasons behind the remaining gap and aspects of ICL unique to the instruction tuning setting.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19670v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses how to improve retrieval-augmented generation (RAG) in LLMs without compromising their general generation abilities.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research proposes learning scalable and pluggable virtual tokens tailored for retrieval-augmented generation. Instead of fine-tuning the entire LLM, which can affect its general capabilities, the approach optimizes only the embeddings of these virtual tokens. This ensures that LLMs can benefit from updated and factual information retrieved without altering their inherent parameters. The method involves multiple training strategies to ensure the virtual tokens are scalable and can integrate seamlessly with various tasks, thus maintaining the general flexibility and functionality of the original model. The training process emphasizes improving embeddings while keeping the primary LLM parameters intact to preserve overall model performance while enhancing retrieval-augmented capabilities.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20485v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Phantom: General Trigger Attacks on Retrieval Augmented Language Generation</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research is significant because it uncovers new security vulnerabilities in Retrieval Augmented Generation (RAG) systems, which can be exploited to compromise the integrity and safety of chatbot applications that rely on LLMs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research was conducted by designing a framework called Phantom, which executes a two-step attack on RAG systems. First, the researchers created a poisoned document designed to be retrieved by the RAG system. This poisoned document included a specific sequence of words, acting as a backdoor trigger, that ensured it would be among the top-k results for certain queries. During the second step, this adversarial document, once retrieved, included an adversarial string that could trigger various harmful behaviors in the LLM generator. These behaviors included denial of service, reputation damage, privacy violations, and inducing harmful actions in the chatbot. The attacks were validated on multiple LLM architectures, including Gemma, Vicuna, and Llama, demonstrating the broad applicability and effectiveness of the attack method across different systems.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The research successfully demonstrated the Phantom attack framework across multiple LLM architectures, validating its effectiveness and broad applicability. The attacks were able to compromise the RAG system by retrieving malicious documents that triggered adversarial behaviors in the LLM generator, confirming the new security risks posed by RAG systems.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20413v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the vulnerability of Large Language Models (LLMs) to specific prompts designed to bypass moderation guardrails, which can lead to harmful behaviors. Current benchmarks often overlook these moderation triggers, complicating the evaluation of jailbreak strategies.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research proposes JAMBench as a new benchmark to evaluate the moderation guardrails of LLMs by including 160 manually crafted instructions across four risk categories with varying levels of severity. The authors also introduce a new jailbreak method called JAM (Jailbreak Against Moderation). The JAM method involves using tailored jailbreak prefixes to bypass input-level filters and employing a fine-tuned shadow model that is functionally similar to the guardrail model. This shadow model generates cipher characters designed to circumvent output-level filters. The study conducts extensive experiments on four different LLMs to evaluate the efficacy of JAM against existing guardrails.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> JAM achieves higher jailbreak success ( x19.88) and lower filtered-out rates ( x1/6) compared to baseline methods during extensive experiments on four different LLMs.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20319v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>ParSEL: Parameterized Shape Editing with Language</b></a> - Focuses on enabling precise 3D shape edits through natural language instructions, democratizing 3D content creation. The system effectively translates language inputs into geometric editing tasks, outperforming traditional methods in control and precision.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20309v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Large Language Models Can Self-Improve At Web Agent Tasks</b></a> - Examines the self-improvement capabilities of LLMs in navigating and performing complex web agent tasks, using synthetic training data for fine-tuning. The method shows significant enhancements in performance and robustness.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20213v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization</b></a> - Develops a method to automatically transform long multimodal documents into concise and visually appealing posters. Demonstrates significant improvements in content relevance and design aesthetics, enhancing educational and professional presentations.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d0002cef-e568-4e92-b3db-50fd662beeec/image.png?t=1721661492"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20139v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning</b></a> - Combines Graph Neural Networks with LLMs to improve question-answering over Knowledge Graphs. Shows state-of-the-art performance on benchmark tests, significantly enhancing factual question-answering capabilities.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e57a9850-3185-4658-ab74-3ce81c931411/image.png?t=1721662012"/></div></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>LLMs in robotics!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20189v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Nadine: An LLM-driven Intelligent Social Robot with Affective Capabilities and Human-like Memory</b></a> - Enhances human-robot interactions by integrating advanced LLMs into the Nadine social robot platform, giving it human-like memory and emotional appraisal capabilities. The robot system uses multimodal inputs and episodic memory to simulate emotional states and provide natural interaction experiences.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20179v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning CodeLLMs</b></a> - Introduces Robo-Instruct, which enhances the performance of open-weight LLMs in robot programming by integrating simulator-based verification and instruction alignment. This method successfully brings smaller models up to or beyond the capability of larger proprietary models like GPT-3.5-Turbo.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19758v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>InterPreT: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning</b></a><a class="link" href="http://arxiv.org/abs/2405.19758v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-30th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Develops a framework that teaches robots to perform long-horizon planning by learning symbolic predicates from human language feedback. The system translates learned predicates and operators into PDDL for robust task planning in diverse settings, achieving high success rates in real-world robot manipulation scenarios.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/77fedc21-d06c-4120-ab08-e953f388b997/image.png?t=1720851824"/></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=7450b115-be71-42fa-be7a-20c33a28c9e5&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Summary of LLMs related research papers published on 29th May, 2024</title>
  <description>Detailed categorization of research trend to stay informed about latest research in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3330f32e-1e87-4acb-a880-d24fde3130a1/27th_may_24.png" length="263140" type="image/png"/>
  <link>https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-29th-may-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-29th-may-2024</guid>
  <pubDate>Fri, 12 Jul 2024 20:28:11 +0000</pubDate>
  <atom:published>2024-07-12T20:28:11Z</atom:published>
    <dc:creator>LLMs Research</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19534v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Preference Learning Algorithms Do Not Learn Preference Rankings</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding the limitations of preference learning algorithms is crucial for improving how LLMs are trained to align with human preferences.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers conducted a detailed analysis of the ranking accuracy of state-of-the-art preference-tuned models. They evaluated models trained with popular preference learning algorithms like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO). To measure performance, they assessed how often these models assigned higher likelihoods to more preferred outputs compared to less preferred ones, using established preference datasets. Additionally, they derived the idealized ranking accuracy that would be achieved if a preference-tuned LLM optimized the DPO or RLHF objectives perfectly. An efficient formula was introduced to quantify the difficulty of learning from given preference datapoints, and experiments were conducted to explore the relationship between ranking accuracy and the empirically popular win rate metric when the model closely matches the reference model used in learning objectives. This provided insights into the alignment gap and differences between on-policy (e.g., RLHF) and off-policy (e.g., DPO) learning algorithms.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00059v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Integrating LLMs with external tool invocations has significantly increased the complexity and latency of LLM serving workloads.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers introduced Conveyor, an LLM serving system designed to optimize the handling of requests that involve external tools. Conveyor employs a novel interface specifically for tool developers, allowing them to expose partial execution opportunities. This enables portion of tool operations to be executed concurrently alongside LLM decoding. Additionally, the system includes a request scheduler that efficiently manages these partial tool executions. By doing so, Conveyor aims to reduce the latency ordinarily associated with handling LLM requests that require external tool interactions.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19332v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Self-Exploring Language Models: Active Preference Elicitation for Online Alignment</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research focuses on enhancing the alignment of LLMs to better follow human intentions, which is crucial for their practical utility and safety.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors propose a novel approach called Self-Exploring Language Models (SELM) to improve the alignment of LLMs through online preference elicitation. The method involves a bilevel optimization objective that is optimistically biased towards responses that are potentially high-reward but out-of-distribution. The inner-level optimization problem is solved using a reparameterized reward function, eliminating the need for a separate reward model. This iterative method allows the LLM to explore more diverse response spaces efficiently. Experiments were conducted using the Zephyr-7B-SFT and Llama-3-8B-Instruct models, and the SELM approach was tested on various benchmarks, including MT-Bench and AlpacaEval 2.0.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The experimental results indicate that SELM significantly improves performance on instruction-following benchmarks such as MT-Bench and AlpacaEval 2.0, as well as various standard academic benchmarks in different settings. This demonstrates that SELM&#39;s active exploration leads to better alignment and enhanced performance of LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19328v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Normative Modules: A Generative Agent Architecture for Learning Norms that Supports Multi-Agent Cooperation</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0576f743-28b3-4458-8899-b018251845bb/image.png?t=1720815651"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Fostering cooperation between agents and humans in environments with social norms is a critical challenge.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces the concept of &#39;Normative Modules,&#39; an architectural enhancement for generative agents. These modules allow agents to recognize and adapt to normative structures within an environment, focusing particularly on the problem of equilibrium selection in cooperative scenarios. To test this, the researchers designed a new type of environment that supports institutions. Agents equipped with the normative module learn via peer interactions to determine which among several candidate institutions is treated as authoritative by a group. This learning process helps the agents to better coordinate their sanctioning behavior, which in turn influences primary behaviors within a social setting. The study makes use of two key evaluation criteria: (i) the ability of the agent to ignore non-authoritative institutions and (ii) the ability to identify authoritative institutions among multiple candidates. These evaluations are essential to measure the stability and effectiveness of cooperation induced by the normative module.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study demonstrates that agents equipped with normative modules achieve more stable cooperative outcomes compared to baseline agents lacking such modules. This showcases the module&#39;s effectiveness in enabling agents to navigate and adapt to social norms, thereby fostering higher average welfare in multi-agent interactions.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19327v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series</a> - <a class="link" href="https://map-neo.github.io?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow">project page</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research aims to improve the transparency and performance of LLMs by creating an open-sourced, highly capable bilingual LLM.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers designed and trained a new bilingual language model called MAP-Neo with 7 billion parameters. The training was conducted on 4.5 trillion high-quality tokens. Crucially, they open-sourced not only the model&#39;s weights but all the details required to reproduce the model, including the pre-training corpus, data cleaning pipeline, intermediate checkpoints, and the training and evaluation frameworks. This approach allows for complete transparency and offers the research community the tools needed to study, improve, and innovate upon the model.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The MAP-Neo model showed performance comparable to the state-of-the-art LLMs and provided a fully transparent and reproducible framework, enhancing the open research community&#39;s capability to study, innovate, and facilitate further improvements.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19325v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Nearest Neighbor Speculative Decoding for LLM Generation and Attribution</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the limitations of LLMs in hallucination and lack of attribution for their outputs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors introduce Nearest Neighbor Speculative Decoding (NEST), a semi-parametric language modeling approach. Unlike traditional LLMs, NEST performs token-level retrieval at each inference step and computes a semi-parametric mixture distribution to identify promising text span continuations from a non-parametric data store. The approach uses an approximate speculative decoding mechanism to either accept a retrieved text prefix or generate a new token, enhancing text fluency and providing source attribution. This approach refines the output of the base LM, outperforming kNN-LM in both quality and speed. The method was applied to Llama-2-Chat 70B for evaluation.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> NEST significantly enhances the generation quality and attribution rate of the base LM across a variety of knowledge-intensive tasks, surpassing the conventional kNN-LM method and performing competitively with in-context retrieval augmentation. Additionally, NEST substantially improves the generation speed, achieving a 1.8x speedup in inference time when applied to Llama-2-Chat 70B.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19209v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8c508426-1a72-4dd7-bfd7-6b173209a852/image.png?t=1720815772"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> There is a growing need for improving the understanding of long-form videos, as current methods often miss important details and operate inefficiently.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces VideoTree, a hierarchical framework designed to dynamically extract and organize query-related information from long videos for better LLM reasoning. The process involves three stages: 1. Adaptive Frame Selection: Instead of using all frames, VideoTree uses an iterative clustering mechanism based on visual features. It then scores these clusters based on their relevance to the given query, ensuring that only significant frames are selected for captioning. 2. Hierarchical Tree Organization: The selected frames are organized into a tree structure where nodes represent clusters of varying relevance and granularity. This tree structure allows finer details to be encoded where needed most, according to the query&#39;s requirements. 3. Tree Traversal for Answer Generation: The hierarchical tree&#39;s keyframes are traversed, and their captions are sent to the LLM to generate an accurate answer. This ensures that only the most relevant information is processed, optimizing both accuracy and efficiency.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19086v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">MEMoE: Enhancing Model Editing with Mixture of Experts Adaptors</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3341570e-239f-4ce1-b2d8-ff8b0c57e4e8/image.png?t=1720815736"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Enhancing how LLMs can be efficiently and effectively edited improves their adaptability and reliability for various tasks without degrading their overall performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces MEMoE, a novel model editing adapter that uses a Mixture of Experts (MoE) architecture paired with a knowledge anchor routing strategy. The MoE architecture facilitates the updating of the model&#39;s knowledge through a bypass mechanism, which ensures that the original parameters of the LLMs remain unchanged, thus preserving their general capabilities. The knowledge anchor routing component is crucial as it routes inputs that demand similar knowledge to the same expert within the MoE framework, which enhances the generalization of the new, updated knowledge while maintaining local specificity. This method was rigorously tested through different tasks, specifically batch editing and sequential batch editing, to compare its performance against existing model editing techniques.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experimental results show the superiority of our approach over both batch editing and sequential batch editing tasks, exhibiting exceptional overall performance alongside outstanding balance between generalization and locality.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19010v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Evaluating the External and Parametric Knowledge Fusion of Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Integrating both external and parametric knowledge into LLMs can overcome limitations like outdated static memory, potentially improving their performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors propose a four-scenario framework to deconstruct knowledge fusion within LLMs. They developed a systematic pipeline for creating datasets and infusing knowledge, which facilitates controlled experiments. This setup allowed for comprehensive testing of LLMs&#39; ability to blend external and parametric knowledge. The investigation involves assessing the various fusion scenarios to identify strengths and weaknesses in LLMs&#39; knowledge integration processes. Specific methods for enhancing parametric knowledge were applied to test if it would subsequently boost the overall capability of the model in these fusion scenarios.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study demonstrates that enhancing parametric knowledge within LLMs can significantly improve their capability for knowledge integration. However, it also highlights challenges around memorizing and accurately eliciting parametric knowledge and delineating its boundaries. These insights are intended to guide future work in harmonizing external and parametric knowledge in LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18915v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasoners</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> LLMs often generate unfaithful chain-of-thought (CoT) reasoning, impacting their reliability in tasks requiring complex reasoning.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers first analyzed CoT faithfulness at the level of individual reasoning steps and identified two reasoning paradigms: centralized and distributed reasoning, noting their relationship with faithfulness. They conducted a joint analysis examining the causal relevance among context, CoT, and the answer, discovering that the model can recall correct information from context that was missing in the CoT, leading to faithfulness issues. To address this, they proposed an inferential bridging method that employs an attribution method to recall pertinent information as hints during CoT generation and uses semantic consistency and attribution scores to filter out noisy CoTs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18906v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Language Generation with Strictly Proper Scoring Rules</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Improving the effectiveness of language generation models by exploring alternative scoring rules beyond the commonly used logarithmic score.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose using strictly proper scoring rules other than the logarithmic score, specifically the Brier score and the Spherical score, for training language generation models. Instead of relying solely on maximum likelihood estimation (MLE) with log-likelihood loss, they adapt these non-local scoring rules to the context of language generation. The adaptation involves training the models with these alternative loss functions without adjusting other hyperparameters. They evaluate the performance of these models using large language models like LLaMA-7B and LLaMA-13B to determine if the alternative loss functions improve language generation capabilities.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experimental results indicate that simply substituting the loss function, without adjusting other hyperparameters, can yield substantial improvements in the model&#39;s generation capabilities.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18886v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Compressing Large Language Models using Low Rank and Low Precision Decomposition</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/27d81254-b241-4061-8eb4-0f53f6833d8d/image.png?t=1720815913"/></div><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b11fb9b8-d8c3-4724-8fa0-cc1d5a155de5/image.png?t=1720815937"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Large Language Models are immense in size, making them difficult to deploy on memory-constrained edge devices. Reducing their size without significant losses in performance is critical for broader and more efficient applications.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers introduced a post-training LLM compression algorithm named CALDERA, which leverages the low-rank structure of weight matrices. The algorithm approximates a weight matrix 𝐖 using a decomposition of the form 𝐖≈𝐐 + 𝐋𝐑, where 𝐋 and 𝐑 are low-rank matrices, and 𝐐, 𝐋, and 𝐑 are quantized to low precision. The model is compressed by substituting each layer with its 𝐐 + 𝐋𝐑 decomposition. The decomposition is achieved by setting up an optimization problem: min_𝐐,𝐋,𝐑‖(𝐐 + 𝐋𝐑 - 𝐖)𝐗^⊤‖_ F^2, with 𝐗 as calibration data. Constraints ensure that 𝐐, 𝐋, and 𝐑 are in low-precision formats. Additionally, 𝐋 and 𝐑 can be adapted for enhancing zero-shot performance. The research also provides theoretical upper bounds on the approximation error using rank-constrained regression framework and explores the tradeoff between compression ratios and model performance.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Results illustrate that compressing LlaMa-2 7B/70B and LlaMa-3 8B models obtained using CALDERA outperforms existing post-training LLM compression techniques in the regime of less than 2.5 bits per parameter.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"> <span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19444v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions</b></a> - Introduces MathChat, a benchmark for testing LLMs in multi-turn mathematical problem solving. Highlights the significant improvements in model performance after fine-tuning with the MathChat sync dataset, particularly in multi-turn settings.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19285v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>MASSIVE Multilingual Abstract Meaning Representation: A Dataset and Baselines for Hallucination Detection</b></a> - Describes the creation of the MASSIVE-AMR dataset for studying multilingual abstract meaning representation and hallucination detection in AMR tasks. Shows that while LLMs are promising in structured parsing tasks, issues like linguistic diversity and accuracy remain challenges.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18952v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Are You Sure? Rank Them Again: Repeated Ranking For Better Preference Datasets</b></a> - Proposes a Repeat Ranking method to enhance consistency in ranking outputs, improving the quality of datasets for training LLMs with reinforcement learning. Demonstrates that repeated ranking yields better performance and more reliable training data compared to standard methods.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18870v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>LLMs achieve adult human performance on higher-order theory of mind tasks</b></a><a class="link" href="http://arxiv.org/abs/2405.18870v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Evaluates the ability of LLMs like GPT-4 and Flan-PaLM to handle complex theory of mind tasks against adult human benchmarks. Finds that these models reach or exceed adult human performance, underscoring the potential of LLMs in advanced cognitive processing.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19550v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Stress-Testing Capability Elicitation With Password-Locked Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Assessing the safety of large language models (LLMs) by fully eliciting their capabilities is critical for AI developers.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers created password-locked models, which are LLMs fine-tuned to hide certain capabilities unless a specific password is present in the input prompt. The purpose is to see if sophisticated elicitation methods, particularly fine-tuning, can reveal these hidden capabilities. They conducted experiments to test various elicitation techniques. First, they provided a few high-quality demonstrations to the models to see if these could fully reveal the password-locked capabilities. They also tested the efficacy of fine-tuning with the password as well as with different passwords. Further, they evaluated reinforcement learning methods to elicit capabilities when demonstrations were not available.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study found that a few high-quality demonstrations are often sufficient to fully elicit password-locked capabilities. More surprisingly, fine-tuning can also elicit other capabilities that have been locked using the same or different passwords. Furthermore, reinforcement learning methods proved effective in eliciting capabilities when demonstrations were unavailable. Therefore, fine-tuning is an effective but potentially unreliable method for eliciting hidden capabilities, especially when models&#39; hidden capabilities exceed those of human demonstrators.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19544v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">One-Shot Safety Alignment for Large Language Models via Optimal Dualization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Addressing safety concerns in LLM alignment with human preferences is critical for enhancing their usability and reliability.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose a novel method to simplify the constrained alignment problem in LLMs by transforming it into an unconstrained alignment problem using a dualization approach. Instead of relying on computationally expensive and unstable Lagrangian-based primal-dual policy optimization methods, the authors pre-optimize a smooth and convex dual function that has a closed form. This eliminates the need for repeated primal-dual policy iterations. They develop two practical algorithms based on this strategy: MoCAN for model-based scenarios and PeCAN for preference-based scenarios. Both algorithms leverage the dualization technique to reduce computational burden and improve training stability.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> A broad range of experiments demonstrate the effectiveness of the proposed MoCAN and PeCAN algorithms in improving computational efficiency and training stability for safety alignment in LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19323v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Are Large Language Models Chameleons?</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding the biases and variability in LLM responses is crucial to ensure robust and fair applications of these models in modeling decisions or collective behaviors.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers conducted over one million simulations where LLMs were asked to answer subjective questions. They compared these answers to real data obtained from the European Social Survey (ESS). To measure the differences and biases between LLMs and survey data, they employed methods such as calculating weighted means and introduced a new measure inspired by the Jaccard similarity index. This approach allowed them to quantitatively analyze and highlight the cultural, age, and gender biases in the models&#39; responses. Additionally, the study emphasized the importance of evaluating the robustness and variability of prompts used for LLMs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The study found significant cultural, age, and gender biases in LLMs&#39; responses compared to real-world survey data. These findings underline the necessity of careful prompt analysis and bias mitigation to ensure LLMs&#39; outputs are reliable and trustworthy for modeling individual or collective behaviors.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19320v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> To address the challenge of incorporating uncertainty estimation in the reward function for reinforcement learning (RL) from human feedback (RLHF), making it suitable for large language models (LLMs).</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers introduced Value-Incentivized Preference Optimization (VPO), a unified approach for both online and offline RLHF. VPO works by regularizing the maximum likelihood estimate of the reward function using the corresponding value function. This regularization is modulated by a sign indicating whether optimism or pessimism is chosen. The core of VPO is its implicit reward modeling, which ensures a simpler RLHF pipeline akin to direct preference optimization. Theoretical guarantees are provided for VPO in both online and offline scenarios, aligning with rates of standard RL methods. The method was validated through empirical experiments in text summarization and dialog tasks, demonstrating its practical effectiveness.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The method showed practicality and effectiveness in empirical experiments, specifically in text summarization and dialog tasks.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00062v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study</b></a><a class="link" href="http://arxiv.org/abs/2406.00062v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Examines LLMs for anonymizing clinical texts, introducing six new evaluation metrics to ensure privacy while maintaining data utility.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19326v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Reasoning3D – Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models</b></a> - Introduces Reasoning3D, a method using LLMs for 3D object segmentation based on textual queries, with rapid deployment in various applications.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b36407ab-b3ba-44b3-a80e-246b9f3bce27/image.png?t=1720815466"/></div><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19255v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Towards Next-Generation Urban Decision Support Systems through AI-Powered Generation of Scientific Ontology using Large Language Models – A Case in Optimizing Intermodal Freight Transportation</b></a><a class="link" href="http://arxiv.org/abs/2405.19255v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Details an AI workflow using the ChatGPT API for generating scientific ontologies, enhancing decision-making in urban and environmental management.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19164v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Learning from Litigation: Graphs and LLMs for Retrieval and Reasoning in eDiscovery</b></a><a class="link" href="http://arxiv.org/abs/2405.19164v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Explores a hybrid approach combining graph-based and LLM-based methods for improving eDiscovery processes in legal contexts.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19119v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Can Graph Learning Improve Task Planning?</b></a><a class="link" href="http://arxiv.org/abs/2405.19119v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"> </a>- Examines the integration of Graph Neural Networks with LLMs for enhancing task planning, demonstrating improvements in decision-making on complex tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.18937v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow"><b>Kestrel: Point Grounding Multimodal LLM for Part-Aware 3D Vision-Language Understanding</b></a> - Introduces Kestrel, a multimodal LLM enhancing 3D vision-language tasks by focusing on part-level understanding and grounded language comprehension.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.19291v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-29th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Grasp as You Say: Language-guided Dexterous Grasp Generation</a> - Introduces DexGYSNet, a dataset enabling robots to perform dexterous grasping from human language commands, enhancing human-robot interaction. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ed0960f4-b61f-4dea-88a1-8f6debcf5741/image.png?t=1720815412"/></div><p class="paragraph" style="text-align:left;">Validates the effectiveness of the DexGYSGrasp framework through extensive experiments, proving its capability in generating intent-aligned and high-quality dexterous grasps.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=1974227e-b8f6-44d4-93cf-0425693ca3b5&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Summary of LLMs related research papers published on 27th May, 2024</title>
  <description>Detailed categorization of research trend to stay informed about latest research in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a3bc5e10-7501-4595-9631-f7d1cbab7375/27th_May.png" length="354398" type="image/png"/>
  <link>https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-27th-may-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-27th-may-2024</guid>
  <pubDate>Sat, 22 Jun 2024 19:04:53 +0000</pubDate>
  <atom:published>2024-06-22T19:04:53Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17658v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Generative Query Reformulation Using Ensemble Prompting, Document Fusion, and Relevance Feedback</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2bc8a16f-28b2-4c69-9428-c33ae3431722/image.png?t=1719081005"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Paper improves the query reformulation</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The authors propose two ensemble-based prompting techniques, GenQREnsemble and GenQRFusion, which generate multiple sets of keywords to enhance retrieval performance. They also introduce post-retrieval variants that incorporate relevance feedback from various sources, including an oracle simulating a human user and a &#39;critic&#39; LLM. The research assesses the impact of ensemble query reformulations, feedback documents, domain-specific instructions, filtered reformulations, and fluent reformulations.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> An ensemble of query reformulations led to an improvement in retrieval effectiveness by up to 18% on nDCG@10 in pre-retrieval settings and 9% on post-retrieval settings on multiple benchmarks, outperforming all previously reported state-of-the-art results.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.17430v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Matryoshka Multimodal Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/05786579-c259-472c-a854-f7e3dad8050e/image.png?t=1719081140"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Large Multimodal Models (LMMs) can face inefficiency in handling dense visual scenarios like high-resolution images and videos. The existing token pruning/merging methods lack the flexibility to adjust information density versus efficiency. Therefore, there is a need to develop a more flexible model for representing visual content in LMMs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The Matryoshka Multimodal Models (M3) proposed in this research learn to represent visual content as nested sets of visual tokens across multiple coarse-to-fine granularities. This allows explicit control over visual granularity per instance during inference, enabling adjustment of the number of tokens based on content complexity. The study analyzes the granularity required for existing datasets and explores the trade-off between performance and visual token length at the sample level.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17428v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important because it introduces the NV-Embed model with enhanced techniques for training LLMs to serve as versatile embedding models, outperforming existing models in text embedding tasks.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces the NV-Embed model with a latent attention layer for pooled embeddings, removes the causal attention mask of LLMs during contrastive training, and implements a two-stage contrastive instruction-tuning method, utilizing curated hard negatives and a blend of retrieval and non-retrieval datasets for training.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17381v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Efficient language modeling with constant speed for various sequence lengths is crucial for improving model training and deployment.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces Lightning Attention, which splits attention calculation into intra-blocks and inter-blocks using different strategies. Intra-blocks use conventional attention computation, while inter-blocks utilize linear attention kernel tricks to eliminate the need for cumsum. Moreover, a tiling technique is applied during forward and backward procedures to maximize GPU hardware utilization. A new architecture called TransNormerLLM (TNL) is also introduced to enhance accuracy. Rigorous testing on diverse datasets with different model sizes and sequence lengths is conducted.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17233v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f70fe29b-382c-470d-9e73-2462aa38d7d9/image.png?t=1719081386"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Parameter quantization for LLMs is essential to reduce memory costs and enhance computational efficiency.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a Column-Level Adaptive weight Quantization (CLAQ) framework that utilizes K-Means clustering, outlier-guided adaptive precision search strategy, and dynamic outlier reservation scheme for LLM quantization. These strategies enable dynamic generation of quantization centroids, assign varying bit-widths to different columns, and retain some parameters in their original float point precision. Experiments were conducted on various mainstream open source LLMs to validate the effectiveness of the proposed methods.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17216v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Autoformalizing Euclidean Geometry</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a8688e41-b985-410b-86af-004dff125aed/image.png?t=1719083032"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Autoformalization is important for translating informal math into formal theorems and proofs that are machine-verifiable.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a neuro-symbolic framework that combines domain knowledge, SMT solvers, and LLMs for autoformalizing Euclidean geometry. The framework leverages theorem provers to automatically fill in diagrammatic information from informal proofs, making the formalization process easier for LLMs. Automatic semantic evaluation is also provided for autoformalized theorem statements. A benchmark called LeanEuclid is constructed with problems from Euclid&#39;s Elements and the UniGeo dataset formalized in the Lean proof assistant. Experiments with GPT-4 and GPT-4V demonstrate the capabilities and limitations of state-of-the-art LLMs in autoformalizing geometry problems.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.17067v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization</a></p><p class="paragraph" style="text-align:left;"><b>Why?: </b>LLMs have limitations in accurately understanding inputs and generating responses due to tokenization. Improving tokenization can enhance LLMs capabilities.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research constructs an adversarial dataset, ADT (Adversarial Dataset for Tokenizer), drawing on various LLMs vocabularies to challenge LLMs tokenization. It comprises ADT-Human and ADT-Auto subsets. The study evaluates the effectiveness of ADT on leading LLMs like GPT-4o, Llama-3, Qwen2.5-max, etc., thereby degrading their performance. Automatic data generation proves efficient and applicable to diverse LLMs, highlighting tokenization flaws.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.17009v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Position: Foundation Agents as the Paradigm Shift for Decision Making</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Decision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches face challenges such as low sample efficiency and poor generalization. Foundation agents can offer rapid adaptation to diverse tasks, making them crucial for transforming the learning paradigm.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research proposes the construction of foundation agents by formulating their fundamental characteristics and challenges inspired by the success of large language models. The roadmap includes steps from data collection to pretraining and adaptation, aligning knowledge and values with LLMs. Critical research questions are identified, and trends for foundation agents in real-world use cases are outlined.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.16964v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Exploring the LLM Journey from Cognition to Expression with Linear Representations</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> To investigate the evolution and interaction of cognitive and expressive capabilities in LLMs and understand the development patterns of these abilities.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research examines Baichuan-7B and Baichuan-33B, bilingual LLMs, to define and explore cognitive and expressive capabilities using linear representations in three phases: Pretraining, Supervised Fine-Tuning (SFT), and Reinforcement Learning from Human Feedback (RLHF). It involves statistical analyses to establish correlations between cognitive and expressive capabilities, assesses how cognitive capacity may restrict expressive potential, delves into the theoretical aspects of these developmental trajectories, evaluates optimization strategies like few-shot learning and repeated sampling, and investigates the connection between the hidden space and output space.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.16908v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> It is important to determine if large language models can effectively express their intrinsic uncertainty in natural language to improve their trustworthiness and reliability in providing answers.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research formalized faithful response uncertainty by measuring the gap between the model&#39;s confidence in its answers and the decisiveness with which they are conveyed. This metric penalizes both excessive and insufficient hedging, indicating how well the uncertainty is reflected. Various aligned LLMs were evaluated on knowledge-intensive question answering tasks to assess their ability to communicate uncertainty.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.16884v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Entity matching is crucial for entity resolution, and utilizing LLMs for this task has shown promise. However, existing LLM-based approaches often overlook global consistency between records. This research aims to investigate different methodologies to enhance LLM-based entity matching.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research compares three main strategies - matching, comparing, and selecting - for LLM-based entity matching. It explores how these strategies incorporate interactions among records from various perspectives. A compositional entity matching (ComEM) framework is designed, leveraging a mix of strategies and LLMs to capitalize on their respective strengths. This framework aims to improve both effectiveness and efficiency of entity matching.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Experimental results showcase significant performance enhancements and cost reductions in LLM-based entity matching by using the ComEM framework.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16802v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">AutoCV: Empowering Reasoning with Automated Process Labeling via Confidence Variation</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a4ce41b4-8c68-4521-9bad-1e7e4b75d1a7/image.png?t=1719082658"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research is important as it aims to enhance the reasoning capabilities of large language models by automatically annotating the reasoning steps, which can improve the accuracy of these models in selecting the correct answers.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The approach involves training a verification model on final answer correctness to generate automatic process annotations by assigning confidence scores to each reasoning step. It detects relative changes in these confidence scores to automatically annotate the reasoning process, eliminating the need for manual annotations and reducing computational costs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Substantial improvements in accuracy across five datasets in mathematics and common sense reasoning.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17402v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">THREAD: Thinking Deeper with Recursive Spawning</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> LLMs often struggle with understanding complex and lengthy contexts. The research introduces the THREAD framework to address this challenge by allowing the model to dynamically spawn new threads based on the context, enabling adaptive problem-solving.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research conducted involves proposing the THREAD framework that treats model generation as a thread of execution. This thread can either run to completion or spawn new threads dynamically, offloading work to child threads. The model decomposes complex tasks or questions into simpler sub-problems for separate child threads to solve. THREAD is implemented using a few-shot learning approach and tested on various benchmarks for agent tasks and question answering.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> THREAD achieves state-of-the-art performance with GPT-4 and GPT-3.5 on different benchmarks, including ALFWorld, TextCraft, and WebShop. It also outperforms existing frameworks by 10% to 50% absolute points with smaller models like Llama-3-8b and CodeLlama-7b.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.17703v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Mechanistic Interpretability of Binary and Ternary Transformers</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Can binary and ternary transformer networks offer a more interpretable alternative in LLMs while maintaining efficiency.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research compares how binary and ternary transformer networks performs compared to full-precision transformer networks.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.02575v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Cross-Modal Safety Alignment: Is textual unlearning all you need?</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Integrating new modalities into LLMs increases the vulnerability to adversarial attacks, bypassing traditional safety training methods.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research investigates whether unlearning solely in the textual domain is sufficient for cross-modality safety alignment. The study evaluates this approach across six datasets to analyze its effectiveness in reducing the Attack Success Rate (ASR) for text-based and vision-text-based attacks while maintaining utility. The experiments compare unlearning in VLMs with and without a multi-modal dataset to understand its impact on safety alignment.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.20774v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Exploring Backdoor Attacks against Large Language Model-based Decision Making</a></p><p class="paragraph" style="text-align:left;">The research proposes a framework for Backdoor Attacks against LLM-enabled Decision-making systems (BALD) by introducing attacks during the fine-tuning phase using word injection, scenario manipulation, and knowledge injection mechanisms. These attacks target different components in the LLM-based decision-making pipeline. Extensive experiments are conducted with three popular LLMs (GPT-3.5, LLaMA2, PaLM2) on two datasets (HighwayEnv, nuScenes) to demonstrate the effectiveness and stealthiness of the proposed backdoor triggers and mechanisms.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16918v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">The Uncanny Valley: Exploring Adversarial Robustness from a Flatness Perspective</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding the relationship between flatness of the loss surface and adversarial robustness is crucial for improving the robustness of machine learning models, including LLMs.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research empirically analyzes the relation between adversarial examples and relative flatness with respect to the parameters of one layer. It observes a peculiar property of adversarial examples during iterative first-order white-box attacks. The study is conducted across various model architectures and datasets, extending the results to large language models. The theoretical connection between relative flatness and adversarial robustness is established by bounding the third derivative of the loss surface.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.16856v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer</a></p><p class="paragraph" style="text-align:left;">Paper introduces a knowledge transfer (KT) method using detailed reasoning paths to transfer knowledge from larger LLMs to smaller ones. This helps fine-tune smaller models for producing accurate predictions with calibrated confidence levels. Experimental evaluation includes multiple-choice questions and sentiment analysis on various datasets to compare the KT method with vanilla and question-answer pair (QA) fine-tuning methods.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16833v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Fine-tuning LLMs requires significant hardware resources, which can be impractical for typical users. As fine-tuning can increase safety risks to LLMs, there is a need to develop methods to reduce these risks while maintaining performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research proposes Safe LoRA, a simple one-liner patch to the LoRA fine-tuning method. Safe LoRA introduces the projection of LoRA weights from selected layers to a safety-aligned subspace, reducing safety risks in LLM fine-tuning while maintaining utility. It is a training-free and data-free approach, requiring only the knowledge of weights from the base and aligned LLMs. Extensive experiments were conducted to demonstrate the effectiveness of Safe LoRA.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17706v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Video Enriched Retrieval Augmented Generation Using Aligned Video Captions</a> - Utilize visual captions from videos to add more context in LLMs input</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.17633v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs</a><br>Understand the relationship between narrative style and empathy in personal stories </p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research empirically examines the relationship between style and empathy using LLMs and large-scale crowdsourcing studies. It introduces the HEART taxonomy to delineate narrative style elements influencing empathy, evaluates LLMs in extracting these elements, and collects a dataset of empathy judgments via crowdsourcing.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17427v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model</a> - A novel LLM specifically designed for comprehensive 3D understanding.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f0b2c26d-7d94-4914-9ea9-fabd1a2e8bb3/image.png?t=1719082742"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> Reason3D takes point cloud data and text prompts as input to produce textual responses and segmentation masks. It uses a hierarchical mask decoder for locating small objects in expansive scenes, generating coarse estimates followed by detailed segmentation. The approach enhances the precision of object identification and segmentation.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17410v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">The Peripatetic Hater: Predicting Movement Among Hate Subreddits</a> - LLMs used to understand how users move between different hate subreddits can help in predicting radicalization and identifying the associated risks on social media platforms.</p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2406.00041v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">QUB-Cirdan at &#39;Discharge Me!&#39;: Zero shot discharge letter generation by open-source LLM</a> - LLMs to automate the generation of critical sections of patient discharge letters in clinics.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The approach utilized the Llama3 8B quantized model with a zero-shot method combined with Retrieval-Augmented Generation (RAG) to generate the &#39;Brief Hospital Course&#39; and &#39;Discharge Instructions&#39; sections. A curated template-based approach was developed to ensure reliability and consistency, and RAG was integrated for word count prediction. Several unsuccessful experiments were also conducted to provide insights for the competition.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17238v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">LLM-Assisted Static Analysis for Detecting Security Vulnerabilities</a> - LLMs based static analysis to identify security vulnerabilities</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/68ddb661-30d3-4ffc-9c68-d4e155819e86/image.png?t=1719082781"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers propose IRIS, a method that integrates LLMs with static analysis to conduct whole-repository reasoning for identifying security vulnerabilities. They create a new dataset, CWE-Bench-Java, containing 120 manually validated security vulnerabilities in real-world Java projects. The projects are complex, averaging 300,000 lines of code and reaching up to 7 million. IRIS utilizes GPT-4 to detect 69 out of 120 vulnerabilities, outperforming the state-of-the-art static analysis tool that only identifies 27. Moreover, IRIS substantially decreases false alarms by over 80% in the best scenario.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17104v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding</a> - Visual grounding links text queries with specific regions in an image. However, existing models struggle with complex queries.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> First, LLM-Optic utilizes LLMs as a text grounder to interpret complex queries and identify intended objects accurately. Then, a pre-trained visual grounding model generates candidate bounding boxes based on the refined query. Subsequently, LLM-Optic annotates these bounding boxes with numerical marks to link text with image regions. Finally, a Large Multimodal Model acts as a Visual Grounder to select objects that best match the original query.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16887v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">A Large Language Model-based multi-agent manufacturing system for intelligent shopfloor</a></p><p class="paragraph" style="text-align:left;">The research proposes a LLM-based multi-agent manufacturing system that defines different agents with specific roles like Machine Server Agent, Bid Inviter Agent, Bidder Agent, Thinking Agent, and Decision Agent. The LLMs support the Thinking Agent and Decision Agent in analyzing shopfloor conditions intelligently and choosing suitable machines rather than following predefined rules. The system involves negotiation steps among agents like BAs and BIA to finalize the distribution of orders and connect manufacturing resources, with Machine Server Agents coordinating physical shopfloor operations.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://arxiv.org/abs/2405.16803v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">TIE: Revolutionizing Text-based Image Editing for Complex-Prompt Following and High-Fidelity Editing</a> - Interpreting complex prompts and maintaining image consistency during editing are crucial in the image generation.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cb729c6b-68a2-49e2-82a5-7e54c65985e1/image.png?t=1719082894"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers design a Chain-of-Thought (CoT) process involving instruction decomposition, region localization, and detailed description. They fine-tune the LISA model, a lightweight multimodal LLM, using the CoT process and the mask of the edited image. By enhancing the diffusion models with prompt knowledge and image mask, they enable superior image generation with improved prompt understanding.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="https://html.onlineviewer.net/[&#39;http://arxiv.org/abs/2405.16792v1&#39;,%20&#39;http://arxiv.org/pdf/2405.16792v1&#39;]" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Laurel: Generating Dafny Assertions Using Large Language Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Automating the generation of helper assertions in Dafny programs can reduce the burden on proof engineers and improve the efficiency of program verification.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The tool, Laurel, uses LLMs to automatically generate helper assertions by designing two domain-specific prompting techniques. First, it helps the LLM determine the location of the missing assertion by analyzing the verifier&#39;s error message and inserting an assertion placeholder. Second, it provides example assertions from the same codebase based on a new lemma similarity metric. The techniques were evaluated on a dataset of helper assertions extracted from real-world Dafny codebases.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16755v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">CHESS: Contextual Harnessing for Efficient SQL Synthesis</a> - convert user context to SQL queries</p><p class="paragraph" style="text-align:left;">Paper utilizes hierarchical retrieval methods, model-generated keywords, locality-sensitive hashing indexing, vector databases, and an adaptive schema pruning technique based on problem complexity and contextual size. The approach is tested on both proprietary models like GPT-4 and open-source models such as Llama-3-70B, demonstrating its effectiveness through ablation studies.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17665v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">Enhanced Robot Arm at the Edge with NLP and Vision Systems</a> - LLMs and vision systems were used together to interpret and execute complex commands conveyed through natural language to improve the accessibility and adaptability of assistive robotic systems.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://html.onlineviewer.net/[&#39;http://arxiv.org/abs/2405.17424v1&#39;,%20&#39;http://arxiv.org/pdf/2405.17424v1&#39;]" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence</a> - The research introduces the Large Auto-Regressive Model (LARM) that utilizes text and multi-view images to predict subsequent actions in an auto-regressive manner. It incorporates a novel data format called auto-regressive node transmission structure and a corresponding dataset. LARM is trained using a two-phase training approach to harvest enchanted equipment in Minecraft, demonstrating more complex decision-making chains than previous methods. The training also enhances the speed of LARM by 6.8 times.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.17013v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-27th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2b7a78">MotionLLM: Multimodal Motion-Language Learning with Large Language Models</a> - single-human, multi-human motion generation, and motion captioning through fine-tuning of pre-trained LLMs, showcasing scalability and flexibility.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e8964b78-679b-4f76-abb9-75353e736ef3/image.png?t=1719082811"/></div><p class="paragraph" style="text-align:left;"><b>How?:</b> MotionLLM encodes and quantizes motions into discrete tokens understandable by LLMs, creating a unified vocabulary of motion and text tokens. By training LLMs with adapters using only 1-3% of their parameters, the framework achieves comparable results in single-human motion generation compared to diffusion models and transformer-based models. It can be extended to multi-human motion generation through autoregressive generation of single-human motions.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=0decef8b-d69e-47eb-a359-348aea49017c&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Summary of LLMs related research papers published on 25th May, 2024</title>
  <description>Detailed categorization of research trend to stay informed about latest research in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6700ea00-e38a-4df8-96e8-7c3246a5f059/25th_may.png" length="246926" type="image/png"/>
  <link>https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-25th-may-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/summary-llms-related-research-papers-published-25th-may-2024</guid>
  <pubDate>Sat, 15 Jun 2024 00:59:37 +0000</pubDate>
  <atom:published>2024-06-15T00:59:37Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16325v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/69921703-8b89-4634-acc2-78adc78d6feb/image.png?t=1718332445"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The paper addresses the challenge of enhancing the accuracy of sparse LLMs while also speeding up their pretraining and inference processes and decreasing their memory footprint.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The approach involves a new method termed SLoPe, which includes sparse plus lazy low-rank adapter pretraining. Key strategies are the addition of low-rank adapters in the final 1% of pretraining iterations, a double-pruned backward pass using N:M sparsity, and efficiency tests on large-scale LLMs (OPT-33B and OPT-66B).</p><p class="paragraph" style="text-align:left;"><b>Results:</b> SLoPe achieved training and inference acceleration up to 1.14x and 1.34x respectively for OPT-33B and OPT-66B, alongside reduced memory usage by up to 0.77x and 0.51x for training and inference, respectively.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16276v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Mechanism Design for LLM Fine-tuning with Multiple Reward Models</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research is crucial as it develops a structured mechanism to incentivize agents to report their preferences truthfully during the fine-tuning of large language models, essential for effective and equitable model training.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper treats the incentive issue as a multi-parameter mechanism design challenge, where training and payment rules are structured to fulfill specific objectives. It introduces an &#39;affine maximizer payment scheme&#39; aligned with &#39;social welfare maximizing training rules&#39;, ensuring compliance with dominant-strategy incentive compatibility and individual rationality. It also shows adaptability of this scheme in real-world scenarios where reporting biases exist.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16236v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">A statistical framework for weak-to-strong generalization</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research addresses the significant challenge of using less capable human feedback to train and align more capable LLMs without degrading their performance. It tackles the issue of weak-to-strong generalization, which is crucial for advancing LLM capabilities beyond current human limits.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The approach conceptualizes the problem as a transfer learning challenge, aiming to transfer latent knowledge from weaker models to enhance a stronger, pre-trained model. By identifying issues with simple fine-tuning methods and proposing a refinement-based strategy, the researchers develop a method that effectively incorporates feedback while maintaining the superior capabilities of LLMs. This involves systematic adjustments using a structured refinement process rather than straightforward fine-tuning.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The practical applicability of the refined approach was demonstrated through successful alignment tasks on three LLMs, showing that it can maintain the advanced capabilities of LLMs while effectively utilizing human-level feedback.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16178v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/419db3c8-48b0-4906-9dd3-3ae55f55f963/image.png?t=1718332588"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> To enhance the efficiency and response quality of LLMs augmented with retrieval capabilities by reducing latency and computational overhead caused by processing large volumes of retrieved documents.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The proposed Sparse RAG model introduces a mechanism for parallel encoding of retrieved documents to eliminate the latency generated by long-range attention mechanisms. It further employs a selection strategy where LLMs use special control tokens to identify and focus only on the most relevant documents during the auto-regressive generation phase. This not only curtails the amount of data processed but also helps in generating more relevant and focused responses.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Evaluation on two datasets demonstrated that Sparse RAG achieves an optimal balance between generation quality and computational efficiency, suggesting its effectiveness and generalizability across different types of generation tasks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16122v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e78a34b9-bd81-4106-b480-aa7af6643f95/image.png?t=1718332688"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the need for effective prompt optimization in LLMs, focusing on the automated, efficient selection and ordering of exemplars to enhance in-context learning without the need for fine-tuning or increasing computational demands at test-time.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study introduces a novel method named EASE that uses hidden embeddings from a pre-trained language model to represent ordered sets of exemplars. EASE employs a neural bandit algorithm to dynamically optimize the selection and sequencing of exemplars based on their potential impact on task performance. It considers the often overlooked influence of exemplar order within the prompt and enables efficient selection of optimal exemplar sets for all test instances of a task, thus eliminating extra computational needs during test-time. Additionally, the method is designed to concurrently optimize both the exemplars and the instructions in the prompt.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Through extensive empirical evaluations (including on novel tasks), EASE demonstrated superior performance over existing methods, providing insights into the impact of exemplar selection on in-context learning effectiveness.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16064v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Keypoint-based Progressive Chain-of-Thought Distillation for LLMs</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/43a1ed3d-bb96-4269-beb0-ff5d563569f5/image.png?t=1718332746"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the inadequacies in traditional chain-of-thought distillation methods for LLMs, particularly in accurately mimicking critical reasoning steps and managing the learning sequence which typically does not align with human cognitive patterns.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The study introduces a novel framework named KPOD which focuses on two main enhancements. First, a token weighting module is developed, utilizing mask learning to prioritize and accurately replicate keystone tokens during distillation. This component helps overcome the challenge where tokens of varying importance are treated equally, which can lead to errors in logical reasoning in the student models. Second, the research proposes an &#39;in-rationale progressive distillation&#39; strategy. This strategy involves initially training the smaller, student model to generate the simpler, final steps of a rationale. As training progresses, the scope expands to gradually include the entire set of reasoning steps. This mimics cognitive development by starting with easier tasks before advancing to more complex problems. To effectively implement this strategy, a weighted token generation loss is introduced to gauge the difficulty of each step in the rationale. Additionally, a value function is created to dynamically schedule the distillation process, taking into account both the difficulty of reasoning steps and the diversity of questions. The approach, unlike previous methods, aligns more closely with natural learning progressions and is designed to enhance the transfer of reasoning abilities more effectively.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> Extensive experiments on four reasoning benchmarks illustrate that KPOD outperforms previous methods by a large margin.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16057v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/acf57ce3-bd87-427f-9a0b-f845f1cea2c0/image.png?t=1718332798"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> This research is important because it tackles the challenge of efficiently fine-tuning and deploying large language models (LLMs) that have been reduced in size through pruning without sacrificing their performance.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper introduces SPP, a Sparsity-Preserved Parameter-efficient fine-tuning method that uses lightweight learnable column and row matrices. This method optimizes sparse LLM weights by maintaining the structural and sparsity integrity of pruned models through element-wise multiplication and residual addition. This allows the models to retain their sparsity patterns and ratios throughout training and merging processes. The effectiveness of SPP was demonstrated on the LLaMA and LLaMA-2 model families using various sparsity patterns (unstructured and N:M), particularly focusing on high sparsity ratios like 75%.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> The application of SPP significantly improved the models&#39; performance across different sparsity patterns, especially in models with higher sparsity ratios, proving it as a viable solution for efficient fine-tuning of sparse LLMs.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"> <span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16295v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Comparative Analysis of Open-Source Language Models in Summarizing Medical Text Data</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> The study addresses a significant gap by methodically evaluating the performance of LLMs specifically for the summarization of medical text data, an area that is under-researched but critical for enhancing digital health applications.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research involves the use of open-source LLMs, such as Llama2 and Mistral, to perform summarization tasks on medical text data. GPT-4 is utilized as an assessor to evaluate the summaries generated by these models. The approach includes both quantitative metrics and qualitative assessments by subject-matter experts to ensure a thorough evaluation of performance.</p><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16282v4?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2c81e5">Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models</a></p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16334v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Devil&#39;s Advocate: Anticipatory Reflection for LLM Agents</a></p><p class="paragraph" style="text-align:left;"><b>Why?: </b>Improves the reliability and efficiency of LLMs</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The method incorporates three phases of introspective interventions: </p><ol start="1"><li><p class="paragraph" style="text-align:left;">Anticipatory reflection to assess potential failures and solutions before actions, </p></li><li><p class="paragraph" style="text-align:left;">Post-action alignment to evaluate and adjust ongoing actions according to task objectives, and </p></li><li><p class="paragraph" style="text-align:left;">A comprehensive review post-completion to refine future strategies. </p><p class="paragraph" style="text-align:left;">This innovative approach was implemented in a zero-shot setup on a platform called WebArena, dedicated to simulating practical tasks in web environments.</p></li></ol><p class="paragraph" style="text-align:left;"><b>Results:</b> The experiments demonstrated that this introspective-driven methodology significantly enhances the performance of LLMs by reducing the number of trials and necessary revisions, enabling more effective navigation through unexpected challenges and improving the overall efficiency in completing tasks compared to existing zero-shot methods.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16281v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">ConStat: Performance-Based Contamination Detection in Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a0f599ca-c66c-4be1-9d1c-1aafcdeb1bb7/image.png?t=1718412850"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The paper addresses the critical issue of data contamination in benchmarking LLMs, which can lead to misleading performance metrics and unreliable model comparisons.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> Introducing a novel definition of contamination, the study defines it as performance that artificially inflates and does not generalize across different scenarios like rephrased samples, synthetic samples, or varied benchmarks. They developed &#39;ConStat&#39;, a statistical method that detects and quantifies contamination by assessing performance discrepancies between a primary and a reference benchmark relative to a set of reference models.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> ConStat successfully detected significant levels of contamination in several well-known models, thereby demonstrating its effectiveness in identifying and quantifying performance inflation in LLM benchmarks.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16241v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/11c558a9-4035-4809-90d0-67757a15e408/image.png?t=1718412909"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> The research addresses the privacy issues associated with user queries in LLMs by focusing on the computational and communication inefficiencies of using homomorphic encryption for private embedding table queries.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The researchers developed FastQuery, a framework designed to optimize private embedding table queries for LLMs. FastQuery employs a communication-aware embedding table quantization algorithm and a one-hot-aware dense packing algorithm to reduce both the computation and communication costs.</p><p class="paragraph" style="text-align:left;"><b>Results:</b> FastQuery achieves more than 4.3 times, 2.7 times, 1.3 times latency reduction, and more than 75.7 times, 60.2 times, 20.2 times communication reduction compared to existing HE-based frameworks on both LLAMA-7B and LLAMA-30B models.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16229v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Understanding the mechanisms of attacks on LLMs safety alignment is crucial to improve defenses against malicious manipulations.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research breaks down the safeguarding process into three stages: recognizing harmful instructions, initiating a refusing response, and completing the refusal. Techniques such as logit lens and activation patching are employed to analyze the influence of different attack strategies on these stages. The &#39;logit cohort&#39; technique is aimed at observing the prediction confidence of a model on various outputs, while &#39;activation cohort&#39; involves modifying certain neurons to observe changes in behavior. Moreover, cross-model probing is utilized to examine shifts in model representations before and after attacks. Two primary attack types were investigated: Explicit Harmful Attack (EHA) and Identity-Shifting Attack (ISA). Analysis of how these attacks affect each safeguarding stage provided insights into the divergent ways these attacks compromise model safety.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16133v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Detection of code generated by LLMs</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduces a novel approach to detect LLM-generated code using a zero-shot mechanism based on code rewriting. Their hypothesis suggests that synthetic code differs less from its rewritten variants compared to human-written code. The authors developed a self-supervised contrastive learning model to train on code samples and their rewritten forms. The contrastive learning framework essentially learns to distinguish whether two code snippets are similar (original and rewritten versions of a human-written code) or dissimilar (original and rewritten versions of synthetic code). To generate rewritten code versions, the researchers utilized another LLM trained specifically to alter code while maintaining functional equivalence. The assessment of the model&#39;s efficacy was conducted using two synthetic code detection benchmarks: the APPS benchmark and the MBPP benchmark.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16363v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">LLMs for User Interest Exploration in Large-scale Recommendation Systems</a></p><p class="paragraph" style="text-align:left;">Proposed system integrates LLMs with traditional recommendation techniques using a hierarchical approach. At the high level, &#39;interest clusters&#39; are defined with controllable granularity. These clusters use language descriptions to encapsulate various user interests. A fine-tuned LLM interprets these clusters to generate descriptions of potentially novel interests, which then guide the recommendation process at the lower level. Here, traditional machine learning models, like a transformer-based recommender system, are constrained to suggest only items that align with the newly identified interest clusters.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16273v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">M</a><sup><a class="link" href="http://arxiv.org/abs/2405.16273v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">3</a></sup><a class="link" href="http://arxiv.org/abs/2405.16273v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Generating human motions</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The M<sup>3</sup> GPT framework integrates multiple modalities—text, music, motion/dance—using a novel approach called <b>discrete vector quantization. </b>This technique allows the model to handle different input types within a single vocabulary, facilitating smoother and more efficient processing. The model generates motions directly in the raw motion space rather than using a discrete tokenizer, which helps in retaining more information and generating more detailed motion dynamics. A key aspect of the framework is its multitask capacity, where text is used as a connecting thread between different motion tasks. This not only simplifies the modeling process but also enhances the model’s learning efficiency through cross-modal reinforcement. The framework was trained and tested on several datasets, involving complex tasks such as zero-shot motion generation, to demonstrate its robustness and versatility.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16205v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases</a></p><p class="paragraph" style="text-align:left;">GeneAgent, a novel LLM designed to reduce the hallucinations in gene set knowledge discovery. GeneAgent incorporates a self-verification mechanism allowing it to autonomously interact with various biological databases to verify and incorporate relevant data into its outputs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16136v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">C3LLM: Conditional Multimodal Content Generation Using Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/57dd1271-e46b-4785-85d1-da488d21c1b8/image.png?t=1718412256"/></div><p class="paragraph" style="text-align:left;">C3LLM integrates three different tasks (video-to-audio, audio-to-text, text-to-audio) with LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16127v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Finetuning Large Language Model for Personalized Ranking</a></p><p class="paragraph" style="text-align:left;"><b>Why?:</b> Recommendation system</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The research introduced Direct Multi-Preference Optimization (DMPO), a method designed to make LLMs more suitable for recommendation tasks by addressing the differences in data nature between LLM pre-training and recommendation demands. DMPO functions by optimizing LLMs to increase the likelihood of selecting relevant (positive) items while reducing the chances of picking irrelevant (negative) ones. The researchers fine-tuned an existing LLM using DMPO by simultaneously maximizing the probability of positive sample outcomes and minimizing the probabilities of various negative samples. </p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16089v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">COLT: Towards Completeness-Oriented Tool Retrieval for Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3cfd6626-1909-460d-8f1d-4c9ed86dd11a/image.png?t=1718412390"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Integrating external tools with LLMs</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The COLT model involves a two-stage process: initial PLM-based models are fine-tuned for semantic understanding between user queries and tools, followed by a dual-view graph collaborative learning framework that constructs bipartite graphs linking queries, scenes, and tools for capturing diverse tool relationships and reducing redundancy.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16011v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Semantic Importance-Aware Communications with Semantic Correction Using Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ef52218e-2521-4936-bf96-074552745758/image.png?t=1718412508"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Improving semantic communications between agents and humans or agents and agents by enhancing understanding and minimizing semantic loss in transmitting visual data.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The paper introduces a method for understanding-level semantic communications (ULSC), employing an image caption neural network (ICNN) to convert visual data into natural language descriptions, which are then refined using a pre-trained large language model (LLM) for semantic importance quantification and error correction. It involves adaptive strategies for minimizing semantic loss while managing transmission delays and also uses LLM at the receiving end for further error corrections.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.16009v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=summary-of-llms-related-research-papers-published-on-25th-may-2024" target="_blank" rel="noopener noreferrer nofollow" style="color: #2C81E5">Streaming Long Video Understanding with Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/39deedb6-ae5e-4616-a43f-feee5d345c16/image.png?t=1718412552"/></div><p class="paragraph" style="text-align:left;"><b>Why?:</b> Efficient and precise long video understanding is crucial for enhancing multimedia content accessibility and interaction but is challenged by high computational costs and data redundancy.</p><p class="paragraph" style="text-align:left;"><b>How?:</b> The proposed VideoStreaming system addresses long video understanding through two main techniques: Memory-Propagated Streaming Encoding and Adaptive Memory Selection. Initially, long videos are segmented into manageable short clips. Each clip is then sequentially encoded while integrating a propagated memory from previously processed clips, which helps in retaining historical content relevance and reducing redundancy. This encoding process continues clip-by-clip, cumulatively updating the memory that captures essential video content elements. After encoding, the Adaptive Memory Selection mechanism comes into play, which dynamically selects a fixed number of relevant memories based on the specific questions posed about the video. This allows the system to focus only on pertinent information, thereby optimizing processing time and memory usage while ensuring that the answers generated by the LLM are contextually appropriate and detailed. This disentangled approach of video extraction and memory utilization enables the LLM to generate responses without reprocessing the entire video, allowing for efficient querying and reduced computational demands.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=d4c84a87-ad77-4015-8583-86e41106f6d6&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>LLMs related research papers published on 23rd May, 2024</title>
  <description>This edition covers ~100 research papers published today. It analyze these papers and provide insights to stay aware about the latest research in LLMs</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0d5d57e3-5a94-497b-8b81-3b3cbbeeaa97/23rd_may.png" length="429570" type="image/png"/>
  <link>https://llm.beehiiv.com/p/llms-related-research-papers-published-23rd-may-2024</link>
  <guid isPermaLink="true">https://llm.beehiiv.com/p/llms-related-research-papers-published-23rd-may-2024</guid>
  <pubDate>Sat, 25 May 2024 15:00:00 +0000</pubDate>
  <atom:published>2024-05-25T15:00:00Z</atom:published>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Core research improving LLMs!</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.15052v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training</a> - <a class="link" href="https://github.com/apple/axlearn?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">project page</a></p><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the problem of accurately measuring and comparing the performance of Mixture-of-Experts (MoE) models and dense models. Prior work has used FLOPs or activated parameters as a measure of model complexity, but this does not take into account the communication overhead in sparse layers, leading to a larger actual training budget for MoE. This setting favors MoE and does not give an accurate representation of its performance compared to dense models.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>To solve this problem, the research paper proposes to use <b>step time</b> as a more accurate measure of model complexity, and <b>to determine the total compute budget under the Chinchilla compute-optimal settings. </b>This ensures a fair comparison between MoE and dense models. Additionally, the paper introduces a <b>3D sharding method for running MoE efficiently on modern accelerators</b>, keeping the dense-to-MoE step time increase within a healthy range.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.15025v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">OAC: Output-adaptive Calibration for Accurate Post-training Quantization</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/56ce5732-5af0-4510-80c6-83027000c724/image.png?t=1717588235"/></div><p class="paragraph" style="text-align:start;">The research paper proposes a solution to reduce the computational cost of LLMs by introducing a technique called Output-adaptive Calibration (OAC), which incorporates the model output in the calibration process. This is done by formulating the quantization error based on the distortion of the output cross-entropy loss. OAC also utilizes output-adaptive Hessians, which are approximated for each layer under reasonable assumptions to reduce the computational complexity. These Hessians are used to update the weight matrices and detect the most salient weights, in order to maintain the model output. This approach outperforms existing methods such as SpQR and BiLLM, especially at extreme low-precision quantization (2-bit and binary).</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper achieved significant performance improvements compared to state-of-the-art baselines such as SpQR and BiLLM, particularly at extreme low-precision quantization. This demonstrates the effectiveness of the proposed OAC technique in reducing the memory footprint, latency, and energy required for inference, without sacrificing accuracy.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.15007v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">RE-Adapt: Reverse Engineered Adaptation of Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/43c1a70d-d712-4de9-b23b-892c653d2b9c/image.png?t=1717588363"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the problem of fine-tuning LLMs on new domains without degrading any pre-existing instruction-tuning. This is important because fine-tuning is often necessary to adapt the model to a specific domain, but it can also lead to overfitting and degrade the performance on the original task.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a solution called RE-Adapt, which stands for Reverse-Engineering Adaptation. This approach involves reverse engineering an adapter, which isolates the knowledge that an instruction-tuned model has learned beyond its corresponding pretrained base model. This adapter is then used to fine-tune the base model on a new domain and readapt it to instruction following. This process requires no additional data or training, making it efficient and effective.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper demonstrates that RE-Adapt and a low-rank variant called LoRE-Adapt outperform other methods of fine-tuning on multiple popular LLMs and datasets. This improvement is seen even when the models are used in conjunction with RAG, which is a more challenging task. This shows the effectiveness of the proposed approach in maintaining the performance of the model on the original task while improving its performance on</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14992v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Linking In-context Learning in Transformers to Human Episodic Memory</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/827f0302-b264-4947-ad51-ff1e1f81bd25/image.png?t=1717588483"/></div><p class="paragraph" style="text-align:start;">The research paper examines the relationship between attention heads in transformer models and human episodic memory. The researchers focus on induction heads, which contribute to in-context learning, and compare them to the contextual maintenance and retrieval (CMR) model of human episodic memory. By analyzing pre-trained LLMs, the researchers demonstrate that CMR-like heads often emerge in intermediate model layers and exhibit similar behavioural and functional characteristics to human memory biases.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14953v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Mallows-DPO: Fine-Tune Your LLM with Preference Dispersions</a></p><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the problem of improving reinforcement learning with human feedback (RLHF) through Direct Preference Optimization (DPO) and its limitations in characterizing the diversity of human preferences.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a new approach called Mallows-DPO, inspired by Mallows&#39; theory of preference ranking. This approach introduces <b>a dispersion index to capture the diversity of human preferences towards prompts</b>. It works by incorporating this dispersion index into existing DPO models, thus enhancing their performance in various benchmark tasks such as synthetic bandit selection, controllable generations, and dialogues.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14852v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6bc52628-12e7-4111-9590-e3b7751f9749/image.png?t=1717590033"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>LLMs compression technique</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a representation-agnostic framework called PV-Tuning, which improves upon existing fine-tuning strategies and provides convergence guarantees in restricted cases. It works by fine-tuning the compressed parameters over a limited amount of calibration data, but instead of using straight-through estimators (STE) which may not perform well in this setting, PV-Tuning uses a more optimized approach.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper achieved better performance in terms of accuracy-vs-bit-width trade-off compared to existing techniques for highly-performant language models such as Llama and Mistral. Specifically, using PV-Tuning, the research paper achieved the first Pareto-optimal quantization for Llama 2 family models at 2 bits per parameter.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.14831v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models</a> - <a class="link" href="https://github.com/OSU-NLP-Group/HippoRAG?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bedecdcd-a71d-43a1-9e99-c7df8c4b66b4/image.png?t=1717593562"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the problem of efficiently and effectively integrating a large amount of new experiences into LLMs after pre-training.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a retrieval framework called HippoRAG, inspired by the <b>hippocampal indexing theory of human long-term memory.</b> It works by synergistically orchestrating LLMs, knowledge graphs, and the Personalized PageRank algorithm to mimic the different roles of the neocortex and hippocampus in human memory. This allows for deeper and more efficient knowledge integration over new experiences.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper shows that HippoRAG outperforms existing retrieval-augmented generation (RAG) methods on multi-hop question answering tasks, achieving up to a <b>20%</b> improvement. It also performs comparably or better than iterative retrieval methods while being significantly cheaper and faster. Additionally, integrating HippoRAG into existing methods leads to further substantial gains.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14768v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models</a> - <a class="link" href="https://github.com/zjunlp/EasyEdit?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">project page</a> 🔥🔥🔥</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d86c9f70-c00e-4fab-a944-cef9f384e686/image.png?t=1717594627"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the problem of updating knowledge in LLMs to ensure accurate responses and facilitate lifelong model editing. Specifically, the paper looks at the challenge of where to store and access this updated knowledge, as well as how to maintain reliability, generalization, and locality in the editing process.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a solution called WISE <b>(Where to Insert updated knowledge for lifelong model Editing)</b>, which involves <b>a dual parametric memory scheme</b>. This includes <b>a main memory for pretrained knowledge </b>and<b> a side memory for edited knowledge.</b> The paper also suggests using a <b>router</b> <b>to decide which memory to access for a given query. </b>Additionally, the paper introduces a<b> knowledge-sharding mechanism to handle continual editing</b>, where different sets of edits are stored in separate subspaces of parameters and then merged into a shared memory without conflicts.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper reports significant performance improvements over previous model editing methods, as demonstrated through extensive experiments on various LLM architectures such as GPT, LLaMA, and Mistral.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14917v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2833b929-3432-4e81-825c-8ad0038631aa/image.png?t=1717595125"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>Existing post-training quantization (PTQ) methods are not ideal for accurately quantizing LLMs to low bit-widths (below 4 bits), which leads to a decrease in performance and requires substantial computation and memory resources.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a <b>Salience-Driven Mixed-Precision Quantization </b>scheme, called SliM-LLM which utilizes the salience distribution of weights<b> to determine the optimal bit-width and quantizers for accurate LLM quantization.</b> It also aligns the bit-width partition to groups for compact memory usage and fast integer inference. The proposed SliM-LLM mainly relies on two techniques: (1) Salience-Determined Bit Allocation, which uses the clustering characteristics of salience distribution to allocate bit-widths for each group, and (2) Salience-Weighted Quantizer Calibration, which optimizes the quantizer parameters by considering the element-wise salience within the group. This helps balance the maintenance of salient information and minimization of errors.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14660v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Implicit In-context Learning</a> 🔥<br><a class="link" href="https://github.com/LzVv123456/I2CL?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6f652d0d-5421-4a8e-bece-9714cd21622c/image.png?t=1717595299"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>In-context Learning (<b>ICL</b>) is a powerful approach that allows LLMs to adapt to new tasks during inference by providing demonstration examples. However, it incurs significant computational and memory costs and is susceptible to the selection and order of demonstration examples.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a new paradigm called <b>Implicit In-context Learning (I2CL)</b> to address the limitations of traditional ICL. I2CL <b>absorbs the demonstration examples within the activation space instead of prefixing them to the test queries. </b>It first generates a condensed vector representation, called a context vector, from the demonstration examples. Then, during inference, it integrates the context vector by injecting a linear combination of the context vector and query activations into the model&#39;s residual streams. This allows I2CL to achieve few-shot performance with zero-shot cost and also exhibits robustness against the variation of demonstration examples. Additionally, I2CL facilitates a novel representation of &quot;task-ids&quot;, which enhances task similarity detection and enables effective transfer learning.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14655v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Multi-turn Reinforcement Learning from Preference Human Feedback</a></p><p class="paragraph" style="text-align:start;">This research paper proposes novel methods for Reinforcement Learning (RL) from preference feedback between two full multi-turn conversations. In the tabular setting, it presents a mirror-descent-based policy optimization algorithm that is specifically designed for the general multi-turn preference-based RL problem. It works by continuously updating the policy based on the feedback received from the human, and it has been proven to converge to Nash equilibrium.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper shows that a deep RL variant of their algorithm outperforms existing RLHF baselines in their newly created environment, Education Dialogue, where a teacher agent guides a student in learning a random topic. This indicates a significant performance improvement in aligning LLMs with human preferences. Additionally, the paper also demonstrates that their algorithm can achieve the same performance as a reward-based RL baseline in an environment with explicit rewards, despite solely relying on a weaker preference signal.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14636v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bf17fe2c-a5c2-404e-b8de-db735f3eabae/image.png?t=1717595840"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the issue of processing LLMs services in real-time on bandwidth-constrained cloud servers, which has become difficult due to the rapid growth in the number of LLM users.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes PerLLM, a personalized inference scheduling framework with edge-cloud collaboration to improve the efficiency of processing diverse LLM services. It integrates the upper confidence bound algorithm based on the constraint satisfaction mechanism to optimize service scheduling and resource allocation within the edge-cloud infrastructure. This allows for meeting processing time requirements while minimizing energy costs.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> PerLLM achieved significant performance improvements compared to other methods. It achieved 2.2x, 2.1x, and 1.6x higher throughput and reduced energy costs by more than 50% in experimental results from different model deployments.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14591v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Base of RoPE Bounds Context Length</a></p><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>Long context capability in LLMs.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a novel property of long-term decay in order to better understand the role of <b>Rotary position embedding (RoPE)</b> in LLMs. It suggests that there is an <b>absolute lower bound for the base value of RoPE</b> in order to obtain certain context length capability in LLMs. This can help shed light on future long context training and improve the overall performance of LLMs.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14589v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Top-Down Partitioning for Efficient List-Wise Ranking</a></p><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>Paper try to overcomes the <b>limited document ranking capacity in LLMs</b>, which hinders their application for list-wise re-ranking.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The paper proposes a novel algorithm that <b>partitions the ranking to a specific depth (k) and processes documents in a top-down manner. </b>This algorithm is inherently <b>parallelizable</b>, as it uses a pivot element that allows for concurrent comparison of documents at any depth. This reduces the number of expected inference calls by 33% when ranking at depth 100, while maintaining the performance of previous approaches.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14438v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">LoRA-Ensemble: Efficient Uncertainty Modelling for Self-attention Networks</a> </p><p class="paragraph" style="text-align:start;">The paper proposes a <b>parameter-efficient deep ensemble method called </b><span style="text-decoration:underline;"><b>LoRA-Ensemble</b></span>, which is based on Low-Rank Adaptation (LoRA). It works by <b>training a single pre-trained self-attention network with shared weights for all ensemble members,</b> while also training member-specific low-rank matrices for attention projections. This approach reduces the computational cost and memory requirements of traditional explicit ensemble methods.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14428v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs</a> - <a class="link" href="https://github.com/onnoo/activation-spikes?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">GitHub</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/691da33e-465f-492b-a559-0d81be4f78bf/image.png?t=1717596635"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the challenge of activation quantization in GLU variants, which are commonly used in feed-forward networks of modern large language models (LLMs). These activation quantization errors, caused by excessive magnitudes of activation, significantly degrade the performance of quantized LLMs.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes two empirical methods, Quantization-free Module (QFeM) and Quantization-free Prefix (QFeP), to isolate the activation spikes during quantization. These methods work by identifying and removing the activation spikes, which are dedicated to a couple of tokens rather than being shared across a sequence. This helps to reduce the local quantization errors and improve the performance of the quantized LLMs.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper has achieved significant performance improvements in the activation quantization of various modern LLMs with GLU variants, including LLaMA-2/3, Mistral, Mixtral, SOLAR, and Gemma. These methods have shown to be more effective than current techniques like SmoothQuant, which fail to control the activation spikes. The code for these methods is available on GitHub.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14371v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">EdgeShard: Efficient LLM Inference via Collaborative Edge Computing</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9230357e-fee1-4a8f-94e7-bf6e7a8e0f82/image.png?t=1717597445"/></div><p class="paragraph" style="text-align:start;">Paper proposes to use collaborative edge computing, where the LLM model is partitioned into shards and deployed on distributed devices. This allows for more efficient and secure LLM inference by bringing the computation closer to the data sources. An adaptive joint device selection and model partition problem is formulated, and an efficient dynamic programming algorithm is designed to optimize the inference latency and throughput.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> EdgeShard achieves <b>up to 50% latency reduction and 2x throughput improvement </b>over baseline methods.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14366v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">MiniCache: KV Cache Compression in Depth Dimension for Large Language Models</a> 🔥</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5bf20d07-74e7-489d-9c75-053c377bb377/image.png?t=1717597501"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper tries to reduce latency in autoregressive generation by compressing the Key-Value (KV) cache, which stores previously generated tokens.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes a solution called MiniCache, which compresses the KV cache across layers from a depth perspective. This is done by disentangling the states into magnitude and direction components, and then interpolating the directions while preserving their lengths. Additionally, a token retention strategy is introduced to keep highly distinct state pairs unmerged. This approach is training-free and can be applied to existing KV cache compression strategies such as quantization and sparsity.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Let’s make LLMs safe!!</b></span></p><p class="paragraph" style="text-align:start;">This research paper <a class="link" href="http://arxiv.org/abs/2405.14577v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Representation noising effectively prevents harmful fine-tuning on LLMs</a> tries to mitigate the dual-use risk in releasing open-source LLMs, where bad actors can easily fine-tune these models for harmful purposes. It proposes a defence mechanism called Representation Noising (RepNoise) which removes information about harmful representations in the LLM, making it difficult for attackers to recover them during fine-tuning. This defence is effective even when attackers have access to the weights and the defender no longer has any control. RepNoise also has the ability to generalize across different subsets of harm, without degrading the overall capability of the LLM or hindering its ability to train on harmless tasks.</p><p class="paragraph" style="text-align:start;">The research paper <a class="link" href="http://arxiv.org/abs/2405.14490v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Impact of Non-Standard Unicode Characters on Security and Comprehension in Large Language Models</a> analyze fifteen distinct LLMs to expose their vulnerabilities. This analysis involves a standardized test with 38 queries and three key metrics: <b>jailbreaks, hallucinations, and comprehension errors.</b> The models are then assessed based on the total occurrences of these metrics. The paper also investigates the impact of non-standard Unicode characters on language models and their safeguarding mechanisms, such as Reinforcement Learning Human Feedback (RLHF). The study suggests that incorporating non-standard Unicode text in training data can enhance the capabilities of language models.</p><p class="paragraph" style="text-align:start;">The research paper <a class="link" href="http://arxiv.org/abs/2405.14488v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability</a> proposes the MoGU framework as a solution to enhance the safety of LLMs while preserving their usability. The framework transforms the base LLM into two variants: the usable LLM and the safe LLM, and uses dynamic routing to balance their contribution. This means that when faced with malicious instructions, the router will assign a higher weight to the safe LLM to ensure harmless responses, while for benign instructions, the router prioritizes the usable LLM to generate helpful responses.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>Creative ways to use LLMs!!</b></span></p><p class="paragraph" style="text-align:left;">Research paper <a class="link" href="http://arxiv.org/abs/2405.14982v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">In-context Time Series Predictor</a> proposes a new method that reformulates TSF tasks as input tokens, by constructing a series of (lookback, future) pairs within the tokens. This aligns more closely with the in-context mechanisms of LLMs and is more parameter-efficient, without the need for pre-trained LLM parameters. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/25ad86d8-ff57-44f3-ab6a-fa06e8ed7198/image.png?t=1717533795"/></div><p class="paragraph" style="text-align:left;">This method also addresses issues such as overfitting in existing Transformer-based TSF models.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14918v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">AnalogCoder: Analog Circuit Design via Training-Free Code Generation</a> presents AnalogCoder which is a training-free LLM agent for designing analog circuits through Python code generation. It works by incorporating a feedback-enhanced flow with tailored domain-specific prompts, which allows for automated and self-correcting design of analog circuits with a high success rate. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7274f2d3-145a-4eb7-a53d-068f74b67a6b/image.png?t=1717534025"/></div><p class="paragraph" style="text-align:start;">Additionally, AnalogCoder also proposes a circuit tool library to archive successful designs as reusable modular sub-circuits, making it easier to create composite circuits.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14767v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">FinRobot: An Open-Source AI Agent Platform for Financial Applications using Large Language Models</a> is an <a class="link" href="https://github.com/AI4Finance-Foundation/FinRobot?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">open source</a> platform that supports multiple financially specialized AI agents powered by LLMs. </p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/38b368ee-3324-4fdb-a14b-b92b6965d03e/image.png?t=1717534429"/></div><p class="paragraph" style="text-align:start;">FinRobot consists of four layers: Financial AI Agents, Financial LLM Algorithms, LLMOps and DataOps, and Multi-source LLM Foundation Models. These layers work together to break down complex financial problems, configure appropriate model application strategies, produce accurate models, and integrate various LLMs for direct access. FinRobot aims to democratize access to specialized LLM-based toolchains and promote wider adoption of AI in financial analysis.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14755v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Large language models can be zero-shot anomaly detectors for time series?</a> uses LLMs for time series anomaly detection. Paper proposes a framework called <b>sigllm</b> for time series anomaly detection. It includes a <b>time-series-to-text conversion</b> module and end-to-end pipelines that prompt language models to perform the detection task. The framework leverages the flexibility of LLMs to handle time series data and the ability to identify anomalies in the input sequence. This is achieved by using two paradigms - a prompt-based detection method that directly asks the LLM to indicate anomalies in the input, and a forecasting method that uses the LLMs forecasting capability to guide the anomaly detection process.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14751v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">AGILE: A Novel Framework of LLM Agents</a> proposes a new framework called <b>AGILE (AGent that Interacts and Learns from Environments)</b> that utilizes LLMs, memory, tools, and interactions with experts to enable agents to perform complex conversational tasks. The key idea is to formulate the construction of such an LLM agent as a reinforcement learning problem, where the LLM serves as the policy model. This means that the agent learns and improves its performance through trial and error, using labeled data of actions and the PPO algorithm for fine-tuning. Additionally, the agent is equipped with the ability to reflect, utilize tools, and consult with experts, allowing it to continuously improve its performance and handle challenging questions.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14748v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">MultiCast: Zero-Shot Multivariate Time Series Forecasting Using LLMs</a> tries to predict multivariate time series using a proposed solution called MultiCast, which is a zero-shot approach that uses large language models (LLMs) to handle multivariate time series data. This is achieved through three novel token multiplexing solutions that reduce dimensionality while retaining key repetitive patterns. Additionally, a quantization scheme is introduced to allow LLMs to better learn these patterns while minimizing the number of tokens used.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="color:#222222;font-family:Helvetica, Arial, sans-serif;font-size:1.5rem;"><b>LLMs for robotics & VLLMs</b></span></p><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.14314v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration</a> - <a class="link" href="https://read-llm.github.io/?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">project page</a></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/37dbd4c8-a3f1-4710-94f1-2095f14dc42a/ezgif-6-5cd7123b2b.gif?t=1717555417"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The paper addresses the challenge of grounding the reasoning ability of LLMs for embodied tasks, specifically in the context of multi-agent collaboration. This problem is caused by the complexity of the physical world and the need for effective coordination between agents.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The paper proposes a framework called Reinforced Advantage feedback (ReAd) to address this problem. This framework involves using critic regression to learn a sequential advantage function from LLM-planned data, and then treating the LLM planner as an optimizer to generate actions that maximize the advantage function. This allows the LLM to have the foresight to determine whether an action will contribute to accomplishing the final task. The paper provides theoretical analysis and extends advantage-weighted regression in reinforcement learning to multi-agent systems.</p><hr class="content_break"><p class="paragraph" style="text-align:left;"><a class="link" href="http://arxiv.org/abs/2405.15019v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Agentic Skill Discovery</a> propose a novel skill discovery approach for robots using LLMs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/381eac75-f10e-4343-bcc0-579aee201adb/image.png?t=1717554014"/></div><p class="paragraph" style="text-align:start;">The framework generates task proposals based on the scene description and robot&#39;s configurations and then uses reinforcement learning to develop corresponding policies. The reliability and trustworthiness of learned behaviors are ensured by an independent vision-language model. This approach allows for incremental skill acquisition, starting from zero skills, and leads to the emergence and expansion of a diverse and reliable skill library for the robot.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14691v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">CityGPT: Towards Urban IoT Learning, Analysis and Interaction with Multi-Agent System</a> proposes a framework called CityGPT, which utilizes three agents to facilitate the learning and analysis of IoT time series data. </p><p class="paragraph" style="text-align:start;">The requirement agent allows users to input natural language queries, which are then decomposed into temporal and spatial analysis processes. These processes are completed by corresponding data analysis agents, and the spatiotemporal fusion agent visualizes the results and provides textual descriptions based on user demands. The framework is agentized and facilitated by a large language model (LLM) to increase data comprehensibility.</p><hr class="content_break"><p class="paragraph" style="text-align:start;"><a class="link" href="http://arxiv.org/abs/2405.14622v2?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Calibrated Self-Rewarding Vision Language Models</a></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e906e27f-af4d-467e-8792-5f82fa2c791e/image.png?t=1717554332"/></div><p class="paragraph" style="text-align:start;">💡<b>Why?: </b>The research paper addresses the issue of hallucination in Large Vision-Language Models (LVLMs). This refers to the phenomenon where the generated text responses from the model appear linguistically plausible but contradict the input image, indicating a misalignment between image and text pairs.</p><p class="paragraph" style="text-align:start;">💻<b>How?: </b>The research paper proposes the Calibrated Self-Rewarding (CSR) approach to address this issue. This approach enables the model to self-improve by iteratively generating candidate responses, evaluating the reward for each response, and curating preference data for fine-tuning. The reward modeling incorporates visual constraints into the self-rewarding process to place greater emphasis on visual input, thus reducing hallucinations.</p><p class="paragraph" style="text-align:start;">📊<b>Results:</b> The research paper achieved a substantial improvement of 7.62% over existing methods across ten benchmarks and tasks. </p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"> <span style="font-size:1.5rem;"><b>LLMs evaluations </b></span></p><p class="paragraph" style="text-align:left;">This research paper <a class="link" href="http://arxiv.org/abs/2405.15092v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Dissociation of Faithful and Unfaithful Reasoning in LLMs</a> tries to benchmark Chain of Thought output. It analyzes LLM error recovery behaviours. Through this analysis, the researchers identified factors that influence LLM recovery behavior, such as the difficulty of the error and the amount of evidence for the correct answer in the context. The paper also distinguishes between faithful and unfaithful error recoveries, suggesting that there are distinct mechanisms driving each type of recovery.</p><p class="paragraph" style="text-align:start;">This research paper <a class="link" href="http://arxiv.org/abs/2405.14863v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns</a> tries to evaluate the conceptualization and reasoning abilities of LLMs by adapting a cross-domain alignment task from cognitive science. To do that, paper conducts a behavioural study in which several LLMs are prompted with a cross-domain mapping task and their responses are analyzed and compared at both the population and individual levels. This task is designed to investigate how well the models represent abstract and concrete concepts and how they reason about these concepts through their mappings between categories.</p><p class="paragraph" style="text-align:start;">This research paper <a class="link" href="http://arxiv.org/abs/2405.14804v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Can LLMs Solve longer Math Word Problems Better?</a> evaluates the ability of LLMs to solve long and complex Math Word Problems (MWPs), which is crucial for their applications in real-world scenarios. Paper published an extended grade-school math (E-GSM), a collection of MWPs with lengthy narratives, and two novel metrics to assess the efficacy and resilience of LLMs in solving these problems. For proprietary LLMs, a new instructional prompt is proposed to mitigate the influence of long context, while for open-source LLMs, a new data augmentation task is developed to improve their Context Length Generalizability (CoLeG).</p><p class="paragraph" style="text-align:start;">This research paper <a class="link" href="http://arxiv.org/abs/2405.14555v3?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models</a> evaluates the impact of bias in LLMs. Paper introduces two novel metrics - the Representative Bias Score (RBS) and the Affinity Bias Score (ABS) - to measure the biases within LLMs. These metrics are then applied in the Creativity-Oriented Generation Suite (CoGS), a collection of open-ended tasks designed to detect these biases. The CoGS uses customized rubrics to analyze the outputs of LLMs and identify any biases present.</p></div><div class="section" style="background-color:#f6f2f2;border-color:#2C81E5;border-radius:15px;border-style:solid;border-width:4px;margin:5.0px 5.0px 5.0px 5.0px;padding:15.0px 15.0px 15.0px 15.0px;"><p class="paragraph" style="text-align:center;"><span style="font-size:1.5rem;"><b>Survey papers</b></span></p><p class="paragraph" style="text-align:left;">This survey paper <a class="link" href="http://arxiv.org/abs/2405.14487v1?utm_source=llm.beehiiv.com&utm_medium=newsletter&utm_campaign=llms-related-research-papers-published-on-23rd-may-2024" target="_blank" rel="noopener noreferrer nofollow">A Comprehensive Overview of Large Language Models (LLMs) for Cyber Defences: Opportunities and Directions</a> discusses several approaches to identify anomalies of cyber threats, enhance incident response, and automate routine security operations. This is achieved by training LLMs on massive textual datasets, which allows them to encode context and provide powerful comprehension for downstream tasks.</p></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2Fa15a9679-18ec-43e6-8e45-6c2b064f95d6%2Fnew_logo.png%3Fv%3D1778403113&publication_name=LLMs+Research&utm_campaign=f35dda10-0786-4961-8f5e-2f0593ceb93c&utm_medium=post_rss&utm_source=llms_research">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
