<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Weaviate Blog</title>
        <link>https://weaviate.io/papers</link>
        <description>Weaviate Blog</description>
        <lastBuildDate>Sun, 01 Sep 2024 00:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <item>
            <title><![CDATA[Distillation Experiments]]></title>
            <link>https://weaviate.io/papers/distillation2</link>
            <guid>https://weaviate.io/papers/distillation2</guid>
            <pubDate>Sun, 01 Sep 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Experiments comparing Distillation to Finetuning]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-5ff5c2f45c270ea2c7022077c1af9dbb.png" width="1000" height="654" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="is-it-better-to-distill-or-finetune-language-models">Is it better to distill or finetune language models?<a href="https://weaviate.io/papers/distillation2#is-it-better-to-distill-or-finetune-language-models" class="hash-link" aria-label="Direct link to Is it better to distill or finetune language models?" title="Direct link to Is it better to distill or finetune language models?" translate="no">​</a></h2>
<p>Arcee did a bunch of experiments comparing model distillation to finetuning, base vs instruct model distillation and more.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="main-takeaways">Main Takeaways:<a href="https://weaviate.io/papers/distillation2#main-takeaways" class="hash-link" aria-label="Direct link to Main Takeaways:" title="Direct link to Main Takeaways:" translate="no">​</a></h2>
<p>♦️ Both logit-based and hidden states-based distillation methods consistently outperform standard SFT across various benchmarks.</p>
<p>♦️ General-Purpose Performance Gains: Significant improvements across datasets like OpenHermes, WebInstruct-Sub, and FineTome, particularly in MMLU and MMLU-Pro benchmarks, indicating enhanced knowledge absorption.</p>
<p>♦️ Domain-Specific Performance Gains: Distilling models for domain-specific tasks, especially when using the same training data as the teacher model leads to performance improvements.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="experiments">Experiments<a href="https://weaviate.io/papers/distillation2#experiments" class="hash-link" aria-label="Direct link to Experiments" title="Direct link to Experiments" translate="no">​</a></h2>
<p>♦️Experiment 1: What's better Supervised-Finetune(SFT) or Distill+SFT?</p>
<p>Three models—Hermes-Distilled (logit-based), Hermes-Hidden-States, and Hermes Vanilla (SFT-only)—were evaluated, all distilled from Arcee Spark using a subset of the Teknium's OpenHermes-2.5 dataset (200k examples). Both distillation methods were better than SFT-only model across major benchmarks such as BBH, MUSR, and MMLU-PRO. The logit-based approach was better then the hidden-states-based distillation.</p>
<p>♦️Experiment 2: Effectiveness of Logit-based Distillation in a Generic Domain</p>
<p>The 1.5B Distilled model, trained on a 200k subset of WebInstruct-Sub, demonstrated performance improvements over the baseline Qwen2-1.5B-Instruct model across all metrics. Its performance was also comparable to the teacher model, Arcee Spark, particularly on MUSR and GPQA benchmarks.</p>
<p>♦️Experiment 3: Distillation on Instruct vs. Base Student Models</p>
<p>The 1.5B-Instruct-Distilled model (logit-based), trained on WebInstruct-Sub, showed performance improvements over the vanilla Qwen2-1.5B-Instruct model on the MMLU benchmark, showing benefits of distillation for enhancing knowledge retrieval.</p>
<p>♦️Experiment 4: Effectiveness of Domain-specific Distillation</p>
<p>Distilling Arcee Agent, a 7B parameter model specialized in function calling, into Qwen2-1.5B-Instruct using the same dataset that trained the teacher model resulted in performance gains. This approach underscores the potential of using the same training data for both teacher and student models to achieve even greater improvements, particularly in domain-specific tasks.</p>
<p>More resources:
<a href="https://blog.arcee.ai/announcing-distillkit/" target="_blank" rel="noopener noreferrer" class="">DistillKit by Arcee AI</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2402.13116" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2402.13116" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/distillation2#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fdistillation2&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Language Model Distillation]]></title>
            <link>https://weaviate.io/papers/distillation</link>
            <guid>https://weaviate.io/papers/distillation</guid>
            <pubDate>Thu, 08 Aug 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Distilling Large Language models in Small Language Models!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-0d2e9892961dbaa75b52800733619f38.png" width="931" height="522" class="img_ev3q"></p>
<!-- -->
<p><strong>Distillation has become popular recently due to its ability to efficiently compress the knowledge of larger LLMs into smaller ones. Here’s how it works, why it’s useful, and examples of how you can perform distillation</strong></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-is-distillation">What is distillation?<a href="https://weaviate.io/papers/distillation#what-is-distillation" class="hash-link" aria-label="Direct link to What is distillation?" title="Direct link to What is distillation?" translate="no">​</a></h2>
<p>Distillation is a model compression technique where a smaller "student" model is trained to mimic the behavior of a larger "teacher" model. This is achieved by transferring knowledge from the teacher to the student, usually through methods like logit-based or hidden states-based distillation. These methods are designed to help the student model replicate the teacher's output distribution or internal representations, often leading to a more efficient model with comparable performance.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="when-would-we-use-this">When would we use this?<a href="https://weaviate.io/papers/distillation#when-would-we-use-this" class="hash-link" aria-label="Direct link to When would we use this?" title="Direct link to When would we use this?" translate="no">​</a></h2>
<p>Distillation is commonly used when deploying large models is impractical due to resource constraints, such as in real-time applications or edge devices. For instance, a smaller student model can be distilled from a powerful teacher model like Llama3.1 405B, retaining much of the original model’s capability but with significantly lower computational demands. Distillation is also useful when adapting models to specific tasks or domains, as seen in domain-specific distillation cases like "function calling," where specialized knowledge from a teacher model is transferred to a smaller model for specific use cases.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="whats-the-benefit">What’s the benefit?<a href="https://weaviate.io/papers/distillation#whats-the-benefit" class="hash-link" aria-label="Direct link to What’s the benefit?" title="Direct link to What’s the benefit?" translate="no">​</a></h2>
<p>Distillation offers a significant reduction in model size and computational requirements while maintaining a high level of performance. This is especially valuable in scenarios where memory and processing power are limited. Moreover, distillation allows for flexibility in model architecture choices; for example, distilling knowledge from a Llama-3.1-70B model into a much smaller StableLM-2-1.6B model. Distillation methods like those provided in Arcee-AI's DistillKit, including logit-based and hidden states-based distillation, can lead to substantial performance gains over traditional training routines without requiring additional data.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="examples-of-distillation-techniques">Examples of Distillation Techniques:<a href="https://weaviate.io/papers/distillation#examples-of-distillation-techniques" class="hash-link" aria-label="Direct link to Examples of Distillation Techniques:" title="Direct link to Examples of Distillation Techniques:" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-logit-based-distillation">(1) Logit-based Distillation:<a href="https://weaviate.io/papers/distillation#1-logit-based-distillation" class="hash-link" aria-label="Direct link to (1) Logit-based Distillation:" title="Direct link to (1) Logit-based Distillation:" translate="no">​</a></h3>
<p>This method involves transferring knowledge by using both the hard targets (actual labels) and soft targets (teacher logits) to guide the student model. The student is trained to minimize the difference between its output distribution and the teacher’s output, typically using Kullback-Leibler (KL) divergence. This method is particularly effective for maintaining performance close to the teacher model while improving the student’s generalization abilities.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-hidden-states-based-distillation">(2) Hidden States-based Distillation:<a href="https://weaviate.io/papers/distillation#2-hidden-states-based-distillation" class="hash-link" aria-label="Direct link to (2) Hidden States-based Distillation:" title="Direct link to (2) Hidden States-based Distillation:" translate="no">​</a></h3>
<p>Here, the focus is on aligning the intermediate layer representations of the student with those of the teacher. This layer-wise guidance helps the student model capture similar features and improves its performance and generalization. This method also allows for cross-architecture distillation, enabling knowledge transfer between different model architectures, such as distilling from a Llama-3.1-70B model into a StableLM-2-1.6B model.</p>
<p>More resources:
<a href="https://arcee-ai-distillkit.my.canva.site/" target="_blank" rel="noopener noreferrer" class="">DistillKit by Arcee AI</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2402.13116" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2402.13116" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/distillation#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fdistillation&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[LoRA: Low-Rank Adaptation of Large Language Models]]></title>
            <link>https://weaviate.io/papers/lora</link>
            <guid>https://weaviate.io/papers/lora</guid>
            <pubDate>Sun, 28 Jul 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Tuning LLMs by only learning a fraction of weight updates!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-69fac3dcafdb12cfed9c6c763f12198f.png" width="4144" height="1753" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-problem-lora-solves">The Problem LoRA Solves:<a href="https://weaviate.io/papers/lora#the-problem-lora-solves" class="hash-link" aria-label="Direct link to The Problem LoRA Solves:" title="Direct link to The Problem LoRA Solves:" translate="no">​</a></h2>
<ul>
<li class="">In early 2021, Microsoft partnered with OpenAI to explore the commercial viability of GPT-3.</li>
<li class="">They found that prompting was insufficient for production tasks like natural language to code generation.</li>
<li class="">Fine-tuning was necessary but prohibitively expensive due to the large size of model checkpoints.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How It Works:<a href="https://weaviate.io/papers/lora#how-it-works" class="hash-link" aria-label="Direct link to How It Works:" title="Direct link to How It Works:" translate="no">​</a></h2>
<ul>
<li class="">LoRA generalizes full fine-tuning(updating every single parameter) by asking two questions:<!-- -->
<ol>
<li class="">Do we need to fine-tune all parameters?</li>
<li class="">For the weight matrices we fine-tune, how expressive should the updates be in terms of matrix rank?</li>
</ol>
</li>
<li class="">These questions define a 2D plane where full fine-tuning is the top-right corner(full rank and full parameter updates) and the origin represents the original model.</li>
<li class="">Any point in this plane is a valid LoRA configuration.</li>
</ul>
<p>The chosen rank of the update matrix controls the expressivity of the finetuning process.</p>
<ul>
<li class="">A d x d matrix can represent any linear transformation in a d-dimensional vector space.</li>
<li class="">By first transforming the input to a lower-dimensional space and then back to the original space, we can restrict the kind of linear transformations that can be represented.</li>
<li class="">This reduces the number of parameters that need to be stored from (dxd) to (dxr + dxr) where r &lt;&lt; d.</li>
<li class="">A point near the origin often performs as well as full fine-tuning. - because often Neural Networks are over-parametrized and thus the weight matrices are full of linearly dependent</li>
<li class="">This suggests that we can start with a low-rank configuration and gradually increase the rank if needed.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="common-practices-when-using-lora">Common practices when using LoRA:<a href="https://weaviate.io/papers/lora#common-practices-when-using-lora" class="hash-link" aria-label="Direct link to Common practices when using LoRA:" title="Direct link to Common practices when using LoRA:" translate="no">​</a></h2>
<ul>
<li class="">How to choose the rank R of the update matrix: Start with a low rank and increase it if needed.</li>
<li class="">When to use full fine-tuning?: When finetuning on data that is completely new and absent from the pretraining of the base model (for example if you are tuning an English model on Martian then full fine-tuning may be necessary).</li>
<li class="">Can I use LoRA for any model architecture?: As long as the model uses matrix multiplication, LoRA can be applied. So basically pretty much every model architecture can use LoRA!</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="benefits-of-lora">Benefits of LoRA:<a href="https://weaviate.io/papers/lora#benefits-of-lora" class="hash-link" aria-label="Direct link to Benefits of LoRA:" title="Direct link to Benefits of LoRA:" translate="no">​</a></h2>
<ul>
<li class="">Reduced checkpoint sizes: On GPT-3, checkpoint size was reduced from 1TB to 25MB.</li>
<li class="">No additional inference latency: LoRA updates can be merged with the original parameters during inference. W_new = W_old + AxB</li>
<li class="">Ability to quickly switch between tasks: LoRA modules can be loaded and unloaded efficiently.(A_frenchxB_french),(A_germanxB_german),(A_spanishxB_spanish)</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="some-interesting-engineering-ideas-enabled-by-lora">Some interesting engineering ideas enabled by LoRA:<a href="https://weaviate.io/papers/lora#some-interesting-engineering-ideas-enabled-by-lora" class="hash-link" aria-label="Direct link to Some interesting engineering ideas enabled by LoRA:" title="Direct link to Some interesting engineering ideas enabled by LoRA:" translate="no">​</a></h2>
<ul>
<li class="">Caching LoRA modules in RAM for faster model switching and routing between different finetunes.</li>
<li class="">Training multiple LoRA modules in parallel on different batches of the training set.</li>
<li class="">Creating a tree of adaptive models where each node is a LoRA module.</li>
</ul>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2106.09685" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2106.09685" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/lora#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Flora&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Prover-verifier Games Improve Legibility of LLM Outputs]]></title>
            <link>https://weaviate.io/papers/prover</link>
            <guid>https://weaviate.io/papers/prover</guid>
            <pubDate>Mon, 22 Jul 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Optimizing to improve legibility of an LLMs output!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-6464c0c35b338b8162ccf32df9d41317.png" width="1194" height="759" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="training-llms-to-write-solutions-such-that-smaller-models-can-better-check-them-this-makes-them-easier-to-check-for-humans-too">Training LLMs to write solutions such that smaller models can better check them. This makes them easier to check for humans, too.<a href="https://weaviate.io/papers/prover#training-llms-to-write-solutions-such-that-smaller-models-can-better-check-them-this-makes-them-easier-to-check-for-humans-too" class="hash-link" aria-label="Direct link to Training LLMs to write solutions such that smaller models can better check them. This makes them easier to check for humans, too." title="Direct link to Training LLMs to write solutions such that smaller models can better check them. This makes them easier to check for humans, too." translate="no">​</a></h2>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="key-findings"><strong>Key Findings</strong><a href="https://weaviate.io/papers/prover#key-findings" class="hash-link" aria-label="Direct link to key-findings" title="Direct link to key-findings" translate="no">​</a></h2>
<ol>
<li class="">
<p><strong>Performance vs. Legibility Trade-Off</strong>:</p>
<ul>
<li class="">A model with only supervised fine-tuning on human-written derivations achieved high legibility but poor accuracy.</li>
<li class="">Optimizing solely for correctness led to high accuracy but poor legibility.</li>
<li class="">They propose method balanced both, retaining high legibility with moderate accuracy.</li>
</ul>
</li>
<li class="">
<p><strong>They Train Three Models</strong>:</p>
<ul>
<li class=""><strong>Robust Verifier</strong>: Accurately distinguishes correct from incorrect solutions, even when misleading.</li>
<li class=""><strong>Helpful Prover</strong>: Generates human-legible solutions, reducing human error.</li>
<li class=""><strong>Sneaky Prover</strong>: Produces subtly incorrect solutions, initially confusing human evaluators and highlighting improvement areas.</li>
</ul>
</li>
</ol>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works"><strong>How It Works</strong><a href="https://weaviate.io/papers/prover#how-it-works" class="hash-link" aria-label="Direct link to how-it-works" title="Direct link to how-it-works" translate="no">​</a></h2>
<p>To train the prover models to produce human-legible text, they use a scheme where the prover's solutions must be easily verifiable by a weaker model.</p>
<p>They use models from the GPT-4 family with significant differences in pretraining compute, they train using a GAN like(generator and discriminator) alternating training process as follows:</p>
<ol>
<li class="">Generate multiple solutions to math problems from the prover, some of which are naturally incorrect.</li>
<li class="">Alternate between two training phases:<!-- -->
<ul>
<li class="">Verifier Training: Train the weak verifier to predict the correctness of the prover’s solutions.</li>
<li class="">Prover Training: Instruct the prover to be either "helpful" or "sneaky," rewarding correct solutions in the former role and incorrect solutions that fool the verifier in the latter.</li>
</ul>
</li>
</ol>
<p>Initially, the sneaky prover easily deceives the verifier. However, as incorrect solutions are added to verifier training, the prover must continuously develop new strategies. Training concludes when the sneaky prover can no longer deceive the verifier.</p>
<p><strong>Training Framework</strong>: Multi-agent setup with honest prover (correct solutions) and sneaky prover (deceptive solutions), both checked by a smaller model(robust verifier).</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="details">Details:<a href="https://weaviate.io/papers/prover#details" class="hash-link" aria-label="Direct link to Details:" title="Direct link to Details:" translate="no">​</a></h2>
<ul>
<li class="">Unique, effective loss function.</li>
<li class="">Honest prover's solutions become detailed and reasoned over time.</li>
<li class="">Sneaky prover's solutions develop subtler flaws.</li>
<li class="">Legibility transfer: Honest prover’s solutions become easier for humans to verify with training.</li>
<li class="">Humans perform better (speed and accuracy) with more trained models in time-constrained tasks.</li>
<li class="">Traditional reinforcement learning for correctness leads to less legible solutions.</li>
<li class="">Approach helps in training models to explain actions comprehensibly, reducing deception.</li>
<li class="">"Legibility tax" affects large model performance. - as legibility increases performance decreases!</li>
</ul>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://cdn.openai.com/prover-verifier-games-improve-legibility-of-llm-outputs/legibility.pdf" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://cdn.openai.com/prover-verifier-games-improve-legibility-of-llm-outputs/legibility.pdf" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/prover#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fprover&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[RouteLLM: Learning to Route LLMs with Preference Data]]></title>
            <link>https://weaviate.io/papers/routellm</link>
            <guid>https://weaviate.io/papers/routellm</guid>
            <pubDate>Sun, 14 Jul 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Route between LLMs to optimize cost and quality!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-ad579ee91c1c05c3964ad05400e6599f.png" width="4391" height="1874" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="you-dont-need-a-2-trillion-parameter-model-to-tell-you-the-capitol-of-france-is-paris">You don't need a 2 trillion parameter model to tell you the capitol of France is Paris.<a href="https://weaviate.io/papers/routellm#you-dont-need-a-2-trillion-parameter-model-to-tell-you-the-capitol-of-france-is-paris" class="hash-link" aria-label="Direct link to You don't need a 2 trillion parameter model to tell you the capitol of France is Paris." title="Direct link to You don't need a 2 trillion parameter model to tell you the capitol of France is Paris." translate="no">​</a></h2>
<p>Be smart and route between a panel of models according to query difficulty and model specialty!</p>
<p>New paper proposes a framework to train a router that routes queries to the appropriate LLM to optimize the trade-off b/w cost vs. performance.</p>
<p>Model inference cost varies significantly: Per one million output tokens: Llama-3-70b ($1) vs. GPT-4-0613 ($60), Haiku ($1.25) vs. Opus ($75)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="overview">Overview:<a href="https://weaviate.io/papers/routellm#overview" class="hash-link" aria-label="Direct link to Overview:" title="Direct link to Overview:" translate="no">​</a></h2>
<p>The RouteLLM paper propose a router training framework based on human preference data and augmentation techniques, demonstrating over 2x cost saving on widely used benchmarks.</p>
<p>They define the problem as having to choose between two classes of models:
(1) strong models - produce high quality responses but at a high cost (GPT-4o, Claude3.5)</p>
<p>(2) weak models - relatively lower quality and lower cost (Mixtral8x7B, Llama3-8b)</p>
<p>A good router requires a deep understanding of the question’s complexity as well as the strengths and weaknesses of the available LLMs.</p>
<p>Explore different routing approaches:</p>
<ul>
<li class="">Similarity-weighted (SW) ranking</li>
<li class="">Matrix factorization</li>
<li class="">BERT query classifier</li>
<li class="">Causal LLM query classifier</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="neat-ideas-to-build-from">Neat Ideas to Build From:<a href="https://weaviate.io/papers/routellm#neat-ideas-to-build-from" class="hash-link" aria-label="Direct link to Neat Ideas to Build From:" title="Direct link to Neat Ideas to Build From:" translate="no">​</a></h2>
<ul>
<li class="">
<p>Users can collect a small amount of in-domain data to improve performance for their specific use cases via dataset augmentation.</p>
</li>
<li class="">
<p>Can expand this problem from routing between a strong and weak LLM to a multiclass model routing approach where we have specialist models(language vision model, function calling model etc.)</p>
</li>
<li class="">
<p>Larger framework controlled by a router - imagine a system of 15-20 tuned small models and the router as the n+1'th model responsible for picking the LLM that will handle a particular query at inference time.</p>
</li>
<li class="">
<p><strong>MoA architectures:</strong> Routing to different architectures of a Mixture of Agents would be a cool idea as well. Depending on the query you decide how many proposers there should be, how many layers in the mixture, what the aggregate models should be etc.</p>
</li>
<li class="">
<p><strong>Route based caching:</strong> If you get redundant queries that are slightly different then route the query+previous answer to a small model to light rewriting instead of regenerating the answer</p>
</li>
</ul>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2406.18665" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2406.18665" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/routellm#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Froutellm&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-Encoders]]></title>
            <link>https://weaviate.io/papers/axn</link>
            <guid>https://weaviate.io/papers/axn</guid>
            <pubDate>Sun, 07 Jul 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[The quality of a reranked retreiver and the speed of a bi-encoder retreiver!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-10a25bd18fd5bbe296569a48669d223a.png" width="3981" height="1664" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-do-you-get-the-retrieval-quality-of-a-cross-encoderre-ranker-and-the-efficiency-of-a-bi-encoder">How do you get the retrieval quality of a cross-encoder/re-ranker and the efficiency of a bi-encoder?<a href="https://weaviate.io/papers/axn#how-do-you-get-the-retrieval-quality-of-a-cross-encoderre-ranker-and-the-efficiency-of-a-bi-encoder" class="hash-link" aria-label="Direct link to How do you get the retrieval quality of a cross-encoder/re-ranker and the efficiency of a bi-encoder?" title="Direct link to How do you get the retrieval quality of a cross-encoder/re-ranker and the efficiency of a bi-encoder?" translate="no">​</a></h2>
<p>Typically people choose to do this with the trusty old retrieve-and-re-rank approach.</p>
<p>This new paper from DeepMind proposes Adaptive Cross-Encoder Nearest Neighbor Search, an alternative which approximates the re-ranker query-document similarities while still using a bi-encoder setup.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="high-level-overview">High-level Overview:<a href="https://weaviate.io/papers/axn#high-level-overview" class="hash-link" aria-label="Direct link to High-level Overview:" title="Direct link to High-level Overview:" translate="no">​</a></h2>
<p>You can think of this as an efficient way to train an adaptor for the query vector that transforms the query vector in such a way that makes the similarity scores b/w query-documents more like the cross-encoder similarity scores.</p>
<ul>
<li class="">Once you pass the query vector through the trained adopter then you can simply use cosine similarity with the document embeddings</li>
<li class="">Can use existing bi-encoder models to initialize the item and query embeddings</li>
<li class="">In an offline indexing step -&gt; compute query/item embeddings to index a given set of items from a target domain making sure the similarity scores are similar to cross encoder scores</li>
<li class="">At test time -&gt; compute the test query embedding to approximate cross-encoder scores of the given test query for a small set of adaptively-chosen items</li>
<li class="">Perform retrieval with the approximate cross-encoder scores</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="details">Details:<a href="https://weaviate.io/papers/axn#details" class="hash-link" aria-label="Direct link to Details:" title="Direct link to Details:" translate="no">​</a></h2>
<p>At test time, they keep item embeddings fixed and perform retrieval over multiple rounds, alternating between:</p>
<blockquote>
<blockquote>
<p>estimating the test query embedding by minimizing error in approximating CE scores of items retrieved thus far, and</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>using the updated test query embedding for retrieving more items in the next round.</p>
</blockquote>
</blockquote>
<p>In the last round once the test query embedding is fully refined, this test query embedding can now be used to retrieve items using cosine similarity.</p>
<p>Proposed k-NN search method can achieve up to 5% and 54% improvement in k-NN recall for k = 1 and 100 respectively over the widely-used DE-based retrieve-and-rerank approach.</p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2405.03651" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2405.03651" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/axn#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Faxn&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Many-Shot In-Context Learning]]></title>
            <link>https://weaviate.io/papers/manyshoticl</link>
            <guid>https://weaviate.io/papers/manyshoticl</guid>
            <pubDate>Fri, 05 Jul 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[All the details around teaching LLMs by giving examples!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-6efa327f047e43183e4bf4a14fa88108.png" width="1315" height="509" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="should-you-finetune-your-llm-or-just-give-relevant-examples-in-the-prompt-how-many-examples-should-you-give-for-best-performance-if-you-give-more-will-it-hurt-perf-does-order-of-the-examples-matter">Should you finetune your LLM or just give relevant examples in the prompt? How many examples should you give for best performance?? If you give more will it hurt perf?? Does order of the examples matter!??<a href="https://weaviate.io/papers/manyshoticl#should-you-finetune-your-llm-or-just-give-relevant-examples-in-the-prompt-how-many-examples-should-you-give-for-best-performance-if-you-give-more-will-it-hurt-perf-does-order-of-the-examples-matter" class="hash-link" aria-label="Direct link to Should you finetune your LLM or just give relevant examples in the prompt? How many examples should you give for best performance?? If you give more will it hurt perf?? Does order of the examples matter!??" title="Direct link to Should you finetune your LLM or just give relevant examples in the prompt? How many examples should you give for best performance?? If you give more will it hurt perf?? Does order of the examples matter!??" translate="no">​</a></h2>
<p>New paper from Deepmind answers all these questions and more, so much to take away from this one, lets dig in!</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="main-takeaways">Main Takeaways:<a href="https://weaviate.io/papers/manyshoticl#main-takeaways" class="hash-link" aria-label="Direct link to Main Takeaways:" title="Direct link to Main Takeaways:" translate="no">​</a></h2>
<ul>
<li class="">Large performance jumps when going from providing very few(1-5) examples(few-shot incontext learning(ICL) to providing many(100s-1000s) examples(many-shot ICL) - The harder the task the more it benefits from more examples in the prompt!</li>
<li class="">Propose using synthetically generated examples(as opposed to human labelled ones) and find that works quite well</li>
<li class="">Propose providing only questions and no answers, in the examples, and find this also works quite well!!</li>
<li class="">Show that many-shot ICL can overcome pre-training biases, perform comparably to supervised fine-tuning, and learn non-NLP prediction tasks.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="juicy-details">Juicy Details:<a href="https://weaviate.io/papers/manyshoticl#juicy-details" class="hash-link" aria-label="Direct link to Juicy Details:" title="Direct link to Juicy Details:" translate="no">​</a></h2>
<ul>
<li class="">
<p>Full supervised/instruction fine-tuning only slightly outperforms many-shot ICL for many tasks</p>
</li>
<li class="">
<p>They mainly test Gemini 1.5 but also try GPT4 and Claude 3.5 and show that different LLMs have varying degrees of success when using many-shot ICL - not a model agnostic trick</p>
</li>
<li class="">
<p>They show that if you provide encough examples in the prompt it can adapt to unseen non-lingual tasks and even in domains that might be misaligned with an LLM’s training data</p>
</li>
<li class="">
<p>Surprisingly, the order of examples in the prompt also influences many-shot performance - would be interesting to see how optimization systems like DSPy can help with this</p>
</li>
<li class="">
<p>Adding more examples, then optimal, in the prompt can also sometimes degrade performance for certain tasks - <strong>weird finding</strong> - opportunity for DSPy to do its thing here aswell</p>
</li>
<li class="">
<p>Many-shot ICL achieves comparable or superior performance when using only problems compared to using problems with solutions - signals that providing solutions with many-shot ICL might just be redundant</p>
</li>
<li class="">
<p>Many shot ICL also shows an improvement in out-of-distribution general problem-solving abilities from in-context learning - Math tasks and etc.</p>
</li>
<li class="">
<p>Biases instilled in the model during pre-training can also be overcome with many shot ICL - a small number of shots leads to the model giving into the bias but with enough examples this eventually diminishes as task learning takes effect in the many shot regime.</p>
</li>
</ul>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2404.11018" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2404.11018" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/manyshoticl#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fmanyshoticl&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Token Pooling to Scale Multi-Vector Retrieval Systems]]></title>
            <link>https://weaviate.io/papers/colbertpooling</link>
            <guid>https://weaviate.io/papers/colbertpooling</guid>
            <pubDate>Sun, 30 Jun 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Clustering tokens to make ColBERT more efficient and usable with Vector DBs!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-a5b08b51ed431f8240c02023b9622097.png" width="1070" height="636" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="multi-vector-retrieval-approaches-like-colbert-have-great-retrieval-quality-but-vector-count-can-balloon-answerai-propose-a-solution">🏹Multi-vector retrieval approaches, like ColBERT, have great retrieval quality but vector count can balloon, AnswerAI propose a solution!<a href="https://weaviate.io/papers/colbertpooling#multi-vector-retrieval-approaches-like-colbert-have-great-retrieval-quality-but-vector-count-can-balloon-answerai-propose-a-solution" class="hash-link" aria-label="Direct link to 🏹Multi-vector retrieval approaches, like ColBERT, have great retrieval quality but vector count can balloon, AnswerAI propose a solution!" title="Direct link to 🏹Multi-vector retrieval approaches, like ColBERT, have great retrieval quality but vector count can balloon, AnswerAI propose a solution!" translate="no">​</a></h2>
<p>Below is an explanation of how ColBERT works and AnswerAI's proposed modification!</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="breakdown-of-different-types-of-encoders">Breakdown of different types of encoders:<a href="https://weaviate.io/papers/colbertpooling#breakdown-of-different-types-of-encoders" class="hash-link" aria-label="Direct link to Breakdown of different types of encoders:" title="Direct link to Breakdown of different types of encoders:" translate="no">​</a></h2>
<p><strong>Cross-encoders:</strong></p>
<ul>
<li class="">Document text &amp; query text strings concatenated and passed into a cross-encoder which then outputs a rank/score.</li>
</ul>
<p><strong>Bi-encoders:</strong></p>
<ul>
<li class="">Document text passed into an encoder and generates a document embedding</li>
<li class="">Query text separately passed into an encoder and generates a query embedding</li>
<li class="">Similarity of query and doc embedding calculated</li>
<li class="">Retrieval performance can suffer especially on Out-Of-Domain data</li>
</ul>
<p><strong>Multi-vector bi-encoder - (eg. ColBERT):</strong></p>
<ul>
<li class="">Functions as a bi-encoder: all documents representations are pre-computed in isolation</li>
<li class="">Similarity computation occurs between individual query and document token vectors, as opposed to the full document.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="main-weakness-of-multi-vector-approaches">Main weakness of multi-vector approaches:<a href="https://weaviate.io/papers/colbertpooling#main-weakness-of-multi-vector-approaches" class="hash-link" aria-label="Direct link to Main weakness of multi-vector approaches:" title="Direct link to Main weakness of multi-vector approaches:" translate="no">​</a></h2>
<ol>
<li class="">
<p>Storage and memory usage balloons up, each token in a document requires a vector(as opposed to one document = one vector)</p>
</li>
<li class="">
<p>Complicated to efficiently search through multiple vectors</p>
</li>
</ol>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="answerais-proposed-token-pooling-solution">AnswerAI's Proposed Token Pooling Solution:<a href="https://weaviate.io/papers/colbertpooling#answerais-proposed-token-pooling-solution" class="hash-link" aria-label="Direct link to AnswerAI's Proposed Token Pooling Solution:" title="Direct link to AnswerAI's Proposed Token Pooling Solution:" translate="no">​</a></h2>
<ul>
<li class="">Main Idea: a lot of the tokens are likely to carry somewhat redundant semantic information, we can semantically cluster them!</li>
<li class="">Requires no model modification whatsoever, nor any complex processing, while&nbsp;greatly improving the scalability of easily updatable indexing methods - like HNSW, which are typically harder to use with ColBERT.</li>
</ul>
<p><strong>Approach:</strong></p>
<ul>
<li class="">Token pooling by clustering similar tokens within a given document and averaging (mean pooling) their representation.</li>
<li class="">After being pooled, the vectors are then quantised to 2-bits using the ColBERTv2 quantisation approach</li>
<li class="">Each cluster is represented by a single token by averaging the values of contained tokens</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="results">Results:<a href="https://weaviate.io/papers/colbertpooling#results" class="hash-link" aria-label="Direct link to Results:" title="Direct link to Results:" translate="no">​</a></h2>
<ul>
<li class="">All results compared to non-pooled ColBERT vector approach</li>
<li class="">Pooling by a factor 2 achieves a&nbsp;50%&nbsp;vector count reduction and 100.6% retrieval performance on average.</li>
<li class="">Pool factor = 3 achieves 66%&nbsp;reduction while reaching&nbsp;99% performance.</li>
<li class="">Pool factor = 4 achieves 75%&nbsp;reduction while reaching&nbsp;97% performance.</li>
</ul>
<p><a href="https://github.com/stanford-futuredata/ColBERT" target="_blank" rel="noopener noreferrer" class="">Code</a></p>
<p><a href="https://www.answer.ai/posts/colbert-pooling.html" target="_blank" rel="noopener noreferrer" class="">Blog</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://www.answer.ai/posts/colbert-pooling.html" download="">🔗 Blog Link</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/colbertpooling#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fcolbertpooling&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Mixture-of-Agents Enhances Large Language Model Capabilities]]></title>
            <link>https://weaviate.io/papers/moa</link>
            <guid>https://weaviate.io/papers/moa</guid>
            <pubDate>Sat, 29 Jun 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Combining small LLMs to outperform larger monolithic LLMs!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-a226000952df08ea3d493335b94d1825.png" width="15472" height="5061" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="can-multiple-smaller-open-source-llms-be-combined-to-outperform-larger-monolithic-llms">🤖Can multiple smaller open-source LLMs be combined to outperform larger monolithic LLMs?<a href="https://weaviate.io/papers/moa#can-multiple-smaller-open-source-llms-be-combined-to-outperform-larger-monolithic-llms" class="hash-link" aria-label="Direct link to 🤖Can multiple smaller open-source LLMs be combined to outperform larger monolithic LLMs?" title="Direct link to 🤖Can multiple smaller open-source LLMs be combined to outperform larger monolithic LLMs?" translate="no">​</a></h2>
<p>New paper shows that LLMs tend to generate better responses when presented with outputs from other models, even if less capable.</p>
<p>They use this to build a Mixture of Agents(MoA) architecture where multiple LLMs are used to iteratively enhance the generation quality.</p>
<p>LLMs in deeper layers are presented responses from LLMs in earlier layers and iteratively improve the response; mitigates individual model deficiencies and enhance overall response.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="moa-architecture">MoA Architecture<a href="https://weaviate.io/papers/moa#moa-architecture" class="hash-link" aria-label="Direct link to MoA Architecture" title="Direct link to MoA Architecture" translate="no">​</a></h2>
<p>The complete architecture consists of LLM agents playing one of two roles:</p>
<ol>
<li class="">
<p>Proposers: These models generate initial reference responses.</p>
</li>
<li class="">
<p>Aggregators: These models synthesize the different responses from the proposers into one high-quality response.</p>
</li>
</ol>
<ul>
<li class="">
<p>Models used: <code>Qwen1.5-110B-Chat</code>, <code>Qwen1.572B-Chat</code>, <code>WizardLM-8x22B</code>, <code>LLaMA-3-70B-Instruct</code>, <code>Mixtral-8x22B-v0.1</code>, <code>dbrx-instruct</code></p>
</li>
<li class="">
<p>3 MoA layers and use the same set of models in each MoA layer</p>
</li>
<li class="">
<p><code>Qwen1.5-110BChat</code> as the aggregator in the last layer</p>
</li>
<li class="">
<p>Some models work better as proposers and others as both proposers and aggregators.</p>
</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-do-you-choose-which-models-to-include-in-the-moa">How do you choose which models to include in the MoA??<a href="https://weaviate.io/papers/moa#how-do-you-choose-which-models-to-include-in-the-moa" class="hash-link" aria-label="Direct link to How do you choose which models to include in the MoA??" title="Direct link to How do you choose which models to include in the MoA??" translate="no">​</a></h2>
<p>Two metrics are used to assess which models are included in the mixture:</p>
<ul>
<li class="">
<p>Performance: The average win rate of models in layer i decides if they are included in layer i + 1.</p>
</li>
<li class="">
<p>Diversity: The diversity of model outputs is important - using heterogeneous models across layers is better then using the same model</p>
</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="details">Details:<a href="https://weaviate.io/papers/moa#details" class="hash-link" aria-label="Direct link to Details:" title="Direct link to Details:" translate="no">​</a></h2>
<ul>
<li class="">
<p>MoA achieves a new SOTA win rate of 65.8% on AlpacaEval 2.0 compared to the previous best of 57.5% achieved by GPT-4 Omni.</p>
</li>
<li class="">
<p>Overall performance comparable to GPT-4 Turbo while being 2× more cost-effective.</p>
</li>
<li class="">
<p>No finetuning required only utilizes the interface of prompting and generation of LLMs.</p>
</li>
<li class="">
<p>Extends the MoE concept to the model level by operating at the model level rather than at the activation level.</p>
</li>
<li class="">
<p>You can swap the final aggregator to any LLM of your choice (Gemini, GPT-4o, Claude3.5) and this improves performance!</p>
</li>
</ul>
<p><a href="https://github.com/togethercomputer/MoA#interactive-cli-demo" target="_blank" rel="noopener noreferrer" class="">Demo</a></p>
<p><a href="https://github.com/togethercomputer/MoA" target="_blank" rel="noopener noreferrer" class="">Code</a></p>
<p><a href="https://www.together.ai/blog/together-moa" target="_blank" rel="noopener noreferrer" class="">Blog</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2406.04692" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2406.04692" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/moa#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fmoa&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer]]></title>
            <link>https://weaviate.io/papers/gliner</link>
            <guid>https://weaviate.io/papers/gliner</guid>
            <pubDate>Sat, 22 Jun 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Using Metadata Filters to Improve Recall in RAG!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-8325172277453f676dc67575e72af221.png" width="1146" height="476" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="using-metadata-filters-to-improve-recall-in-rag">Using Metadata Filters to Improve Recall in RAG<a href="https://weaviate.io/papers/gliner#using-metadata-filters-to-improve-recall-in-rag" class="hash-link" aria-label="Direct link to Using Metadata Filters to Improve Recall in RAG" title="Direct link to Using Metadata Filters to Improve Recall in RAG" translate="no">​</a></h2>
<p>Filtered search using metadata filtering is a simple technique that can significantly improve retrieval quality in a RAG pipeline, but how do you extract metadata from chunks if your data doesn't already come with it??</p>
<p>GLiNER is a powerful model that allows you to extract arbitrary entities such as names, times, places, etc. from any text chunk. It outperforms decoder models like ChatGPT and others at zero-shot identification of names entities in text chunks.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How it Works:<a href="https://weaviate.io/papers/gliner#how-it-works" class="hash-link" aria-label="Direct link to How it Works:" title="Direct link to How it Works:" translate="no">​</a></h2>
<p>GLiNER operates by taking in text chunks and entity labels, that you want to identify in the chunks. Both inputs are concatenated, encoded, and projected into the same latent space and fed into a classifier that predicts the entity labels per word in the text input.</p>
<p>This method allows the model to generalize across different NER tasks and labels passed in at query time.</p>
<p>The fact that entity vectors and the text chunk is concatenated allows the entity labels to attend to the text chunks and vice versa in the encoder step which allows GLiNER to work very well OOD.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="architechture">Architechture:<a href="https://weaviate.io/papers/gliner#architechture" class="hash-link" aria-label="Direct link to Architechture:" title="Direct link to Architechture:" translate="no">​</a></h2>
<p>GLiNER consists of three joined components:</p>
<ol>
<li class="">
<p>An encoder backbone(DeBERTa) that generates token-level representations of the entity labels and text tokens.</p>
</li>
<li class="">
<p>A simple feedforward network that takes in entity token representations from the encoder and embeds them into vectors.</p>
</li>
<li class="">
<p>A Span layer that embeds groups of words (ie. "McGill University" -&gt; vector) from the text into vectors</p>
</li>
</ol>
<p>The entity vectors and span vectors are then combined and used to train a classifier that identifies which text spans correctly paired with entity labels.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="effectiveness">Effectiveness:<a href="https://weaviate.io/papers/gliner#effectiveness" class="hash-link" aria-label="Direct link to Effectiveness:" title="Direct link to Effectiveness:" translate="no">​</a></h2>
<ul>
<li class="">
<p>Generalization: The model demonstrates SoTA performance across various NER benchmarks, outperforming traditional task-specific models.</p>
</li>
<li class="">
<p>Adaptability: GLiNER is super easy to use, pass in any text and any labels you want to extract and it simply works making it a flexible solution to add to your RAG pipeline.</p>
</li>
<li class="">
<p>Scalability: The unified approach simplifies the deployment process, as a single model can handle multiple NER tasks.</p>
</li>
</ul>
<p><a href="https://huggingface.co/spaces/tomaarsen/gliner_medium-v2.1" target="_blank" rel="noopener noreferrer" class="">Demo</a>
<a href="https://github.com/urchade/GLiNER" target="_blank" rel="noopener noreferrer" class="">Code</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2311.08526" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2311.08526" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/gliner#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fgliner&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs]]></title>
            <link>https://weaviate.io/papers/goldfish</link>
            <guid>https://weaviate.io/papers/goldfish</guid>
            <pubDate>Tue, 18 Jun 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Training LLMs without making them memorize!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-b09484d9e6765c944852cbb0958c9ecd.png" width="1704" height="733" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-do-you-train-a-large-language-model-without-it-memorizing-training-data">How do you train a Large Language Model without it memorizing training data?<a href="https://weaviate.io/papers/goldfish#how-do-you-train-a-large-language-model-without-it-memorizing-training-data" class="hash-link" aria-label="Direct link to How do you train a Large Language Model without it memorizing training data?" title="Direct link to How do you train a Large Language Model without it memorizing training data?" translate="no">​</a></h2>
<p>This paper proposes a technique called Goldfish Loss that is now used to mitigate the risk of LLMs memorizing copyrighted or private training data.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short:<a href="https://weaviate.io/papers/goldfish#in-short" class="hash-link" aria-label="Direct link to In Short:" title="Direct link to In Short:" translate="no">​</a></h3>
<p>The paper introduces Goldfish Loss, a method where the model does not compute the loss on every token but excludes (e.g.) 1 in 4 tokens from loss computation. This makes it difficult for the model to memorize the training data.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How it Works:<a href="https://weaviate.io/papers/goldfish#how-it-works" class="hash-link" aria-label="Direct link to How it Works:" title="Direct link to How it Works:" translate="no">​</a></h3>
<p>Goldfish Loss works by omitting a portion of tokens from loss computation during training. When the model encounters these excluded tokens at test time, it has to guess, reducing its ability to reproduce training samples exactly.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="effectiveness">Effectiveness:<a href="https://weaviate.io/papers/goldfish#effectiveness" class="hash-link" aria-label="Direct link to Effectiveness:" title="Direct link to Effectiveness:" translate="no">​</a></h3>
<p>In standard training on Wikipedia articles, about 85% of them get perfectly memorized after 100 updates. With Goldfish Loss, the model usually diverges from the training data within the first 5 tokens it generates.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="trade-off">Trade-off:<a href="https://weaviate.io/papers/goldfish#trade-off" class="hash-link" aria-label="Direct link to Trade-off:" title="Direct link to Trade-off:" translate="no">​</a></h3>
<p>The model learns slower because it does not get credit for the dropped tokens. Training on N tokens with Goldfish Loss is equivalent to standard training on 0.75N tokens.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="benefits">Benefits:<a href="https://weaviate.io/papers/goldfish#benefits" class="hash-link" aria-label="Direct link to Benefits:" title="Direct link to Benefits:" translate="no">​</a></h3>
<p>Goldfish training is scalable and helps avoid the need for unlearning methods, which are often not scalable. This makes it possible to prevent the memorization of copyrighted text/code during training.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="results">Results:<a href="https://weaviate.io/papers/goldfish#results" class="hash-link" aria-label="Direct link to Results:" title="Direct link to Results:" translate="no">​</a></h3>
<p>The paper validates Goldfish Loss by pre-training a model for 200B tokens, showing that it effectively prevents memorization without significantly compromising the learning rate.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="details-in-the-paper">Details in the Paper:<a href="https://weaviate.io/papers/goldfish#details-in-the-paper" class="hash-link" aria-label="Direct link to Details in the Paper:" title="Direct link to Details in the Paper:" translate="no">​</a></h3>
<ul>
<li class="">Explanation of the Goldfish Loss technique</li>
<li class="">Comparison of memorization rates with standard training</li>
<li class="">Analysis of the trade-offs between learning rate and memorization prevention</li>
<li class="">Validation experiments and results</li>
</ul>
<p><a href="https://github.com/ahans30/goldfish-loss" target="_blank" rel="noopener noreferrer" class="">Code</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2406.10209" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2406.10209" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/goldfish#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fgoldfish&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Visual Instruction Tuning]]></title>
            <link>https://weaviate.io/papers/vit</link>
            <guid>https://weaviate.io/papers/vit</guid>
            <pubDate>Sun, 28 Apr 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Training a LLM to understand images!]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-d615b462063a413adb6e7edc20a8e6fb.png" width="1150" height="922" class="img_ev3q"></p>
<!-- -->
<p>How do you teach a Large Language Model to see? Here's a breakdown!</p>
<p>This paper proposes a technique called Visual Instruction Tuning that is now used by many of the language vision models we see in the field such as LLaVA, GPT4-Vision and Gemini etc.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short:<a href="https://weaviate.io/papers/vit#in-short" class="hash-link" aria-label="Direct link to In Short:" title="Direct link to In Short:" translate="no">​</a></h3>
<p>The paper introduces a method to generate multimodal language-image instruction-following data using a language-only GPT-4 model. This data is then used to train LLaVA, a model that combines a vision encoder and a large language model (LLM) for general-purpose visual and language understanding.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How it Works:<a href="https://weaviate.io/papers/vit#how-it-works" class="hash-link" aria-label="Direct link to How it Works:" title="Direct link to How it Works:" translate="no">​</a></h3>
<p>VIT works by using GPT-4 to generate instructions for corresponding images and captions. This dataset is used to train LLaVA to learn to follow instructions and understand images. A vision encoder (CLIP ViT 40) is combined with an LLM (Vicuna) to process text and images and generate text.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="architecture">Architecture:<a href="https://weaviate.io/papers/vit#architecture" class="hash-link" aria-label="Direct link to Architecture:" title="Direct link to Architecture:" translate="no">​</a></h3>
<p>LLaVA consists of two main components:</p>
<ol>
<li class="">
<p>Vision Encoder (VE): A pre-trained vision encoder (e.g. CLIP) that takes an image as input and generates a visual embedding.</p>
</li>
<li class="">
<p>Large Language Model (LLM): A pre-trained LLM (Vicuna) that takes a text prompt as input and generates a language embedding.</p>
</li>
</ol>
<p>The VE and LLM are combined through a series of layers and mechanisms to enable multimodal understanding and generation:</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="benefits">Benefits:<a href="https://weaviate.io/papers/vit#benefits" class="hash-link" aria-label="Direct link to Benefits:" title="Direct link to Benefits:" translate="no">​</a></h3>
<p>The combination of the VE and LLM enables LLaVA to understand and generate text and images in a unified framework, leverage the strengths of both visual and language models, generalize to unseen images and text prompts.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="results">Results:<a href="https://weaviate.io/papers/vit#results" class="hash-link" aria-label="Direct link to Results:" title="Direct link to Results:" translate="no">​</a></h3>
<p>LLaVA achieves a 85.1% relative score compared to GPT-4 on a synthetic multimodal instruction-following dataset</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="details-in-the-paper">Details in the Paper:<a href="https://weaviate.io/papers/vit#details-in-the-paper" class="hash-link" aria-label="Direct link to Details in the Paper:" title="Direct link to Details in the Paper:" translate="no">​</a></h3>
<ul>
<li class="">
<p>The VIT generation process, including prompt engineering and filtering</p>
</li>
<li class="">
<p>The LLaVA architecture, including the vision encoder and LLM components</p>
</li>
<li class="">
<p>Experimental results, including comparisons to GPT-4 and other baselines</p>
</li>
<li class="">
<p>Ablation studies and analysis of the effectiveness of different components and training strategies</p>
</li>
</ul>
<p><a href="https://huggingface.co/spaces/liuhaotian/LLaVA-1.6" target="_blank" rel="noopener noreferrer" class="">Demo</a></p>
<p><a href="https://llava-vl.github.io/" target="_blank" rel="noopener noreferrer" class="">Code</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2304.08485" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2304.08485" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/vit#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fvit&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Retrieval-Augmented Dual Instruction Tuning (RA-DIT)]]></title>
            <link>https://weaviate.io/papers/radit</link>
            <guid>https://weaviate.io/papers/radit</guid>
            <pubDate>Thu, 25 Apr 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Fine-Tuning for Better Retrieval-Augmented Generation]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-65d4ebba7d2a2fe2a62b13d570c887bb.png" width="1125" height="510" class="img_ev3q"></p>
<!-- -->
<p>Can we finetune our LLM and retriever together to improve RAG performance?
This paper proposes a technique to do exactly that!</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="rag-basics">RAG Basics:<a href="https://weaviate.io/papers/radit#rag-basics" class="hash-link" aria-label="Direct link to RAG Basics:" title="Direct link to RAG Basics:" translate="no">​</a></h3>
<p>When you prompt an LLM, RAG supplies relevant documents. A separate retrieval model computes the probability of each text chunk being relevant and provides the top chunks to the LLM. The LLM generates tokens based on the chunks, prompt, and previous tokens.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short:<a href="https://weaviate.io/papers/radit#in-short" class="hash-link" aria-label="Direct link to In Short:" title="Direct link to In Short:" translate="no">​</a></h3>
<p>Fine-tuning LLMs and retrieval models together improves performance without extensive data processing, enabling better retrieval-augmented generation.
LLMs aren't exposed to retrieval-augmented inputs during pretraining, limiting their ability to use retrieved text effectively. Fine-tuning the LLM and retrieval model together can improve performance without requiring extensive data processing.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How it Works:<a href="https://weaviate.io/papers/radit#how-it-works" class="hash-link" aria-label="Direct link to How it Works:" title="Direct link to How it Works:" translate="no">​</a></h3>
<p>Authors from Meta fine-tuned Llama 2 (65B parameters) and DRAGON+, a retriever, to create RA-DIT 65B. They fine-tuned Llama 2 on prompts with retrieved text and questions, and fine-tuned DRAGON+ to retrieve more relevant chunks. Fine-tuning was supervised for tasks like question-answering and self-supervised for text chunk completion.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="results">Results:<a href="https://weaviate.io/papers/radit#results" class="hash-link" aria-label="Direct link to Results:" title="Direct link to Results:" translate="no">​</a></h3>
<p>RA-DIT 65B achieved 49.1% accuracy on average across four question datasets, outperforming LLaMA 2 65B with DRAGON+ (45.1%) and LLaMA 2 65B alone (32.9%). With five example inputs, RA-DIT 65B reached 51.8% accuracy.
RA-DIT offers an efficient way to enhance LLM performance with RAG, making it a valuable technique for developers.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="details">Details:<a href="https://weaviate.io/papers/radit#details" class="hash-link" aria-label="Direct link to Details:" title="Direct link to Details:" translate="no">​</a></h3>
<p>RA-DIT fine-tunes Llama 2 and DRAGON+ to work together effectively, leveraging the strengths of both models to generate better output. By fine-tuning the LLM to better use retrieved knowledge and the retrieval model to select more relevant text, RA-DIT achieves improved performance without requiring extensive data processing.</p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2310.01352" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2310.01352" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/radit#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fradit&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Spotting LLMs With Binoculars: Zero-Shot Detection Of Machine-Generated Text]]></title>
            <link>https://weaviate.io/papers/paper24</link>
            <guid>https://weaviate.io/papers/paper24</guid>
            <pubDate>Mon, 19 Feb 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Zero-shot detection of LLM generated content.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-02497c7ab1cc48f68bfd12a8743404b6.png" width="1808" height="928" class="img_ev3q"></p>
<!-- -->
<p>Can you reliably tell apart fake, LLM-generated, text from human-written text?🤖⚖️👱</p>
<p>Binoculars is a technique that requires no training and can 0-shot detect 90% of LLM-generated content at a 0.01% false positive rate.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short⏩:<a href="https://weaviate.io/papers/paper24#in-short" class="hash-link" aria-label="Direct link to In Short⏩:" title="Direct link to In Short⏩:" translate="no">​</a></h3>
<p>Human tokens are, on average more surprising to LLMs than other LLM tokens. They use this insight to identify a classification threshold.</p>
<p>Given two LLMs, M1 and M2. Their main insight is that human text should diverge from M1 more than M2 diverges from M1, provided the LLMs M1 and M2 are more similar to each other than they are to a human.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="details">Details🔎:<a href="https://weaviate.io/papers/paper24#details" class="hash-link" aria-label="Direct link to Details🔎:" title="Direct link to Details🔎:" translate="no">​</a></h3>
<p>They look at the text in question through the lenses of two different LMs and calculate two perplexity scores:</p>
<ol>
<li class="">
<p>Perplexity of the text using an "observer" LLM(M1).</p>
</li>
<li class="">
<p>Compute all the next-token predictions that a "performer" LLM(M2) would make at each position in the string, and compute their perplexity according to the "observer" LLM(M1).</p>
</li>
</ol>
<p>Then, they take the ratio of the two PPL scores: PPL1/PPL2.</p>
<p>If the string is written by a machine, we should expect these two perplexities to be similar. If it is written by a human they should be different.</p>
<blockquote>
<blockquote>
<p>They find that if PPL1/PPL2 &gt; 0.9 then text is human generated; otherwise it's LLM generated.</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>Works on detecting fake multilingual text as well.</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>They think of PPL as how surprising the next token is - human tokens are, on average more surprising to LLMs than LLM tokens.</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>They use PPL2, what they call cross perplexity, to account for the increase in perplexity due to the prompt; normalizing the observed perplexity by the expected perplexity of a machine acting on the same text, we can arrive at a detection metric that is fairly invariant to the prompt</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>They use Falcon-7b model (M1) and the Falcon-7b-instruct (M2)</p>
</blockquote>
</blockquote>
<p><a href="https://github.com/ahans30/Binoculars/tree/main" target="_blank" rel="noopener noreferrer" class="">🧑‍💻Code</a></p>
<p><a href="https://huggingface.co/spaces/tomg-group-umd/Binoculars" target="_blank" rel="noopener noreferrer" class="">🤗HuggingFace Demo</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2401.12070" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2401.12070" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper24#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper24&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Retrieval-Augmented Generation for Large Language Models: A Survey]]></title>
            <link>https://weaviate.io/papers/paper22</link>
            <guid>https://weaviate.io/papers/paper22</guid>
            <pubDate>Tue, 13 Feb 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Overview of the different RAG techniques.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-8b73e1958d2720389e73df7809a52f2a.png" width="514" height="512" class="img_ev3q"></p>
<!-- -->
<p>A recent survey on Retrieval-Augmented Generation (RAG) mentions an evolving paradigm:
Modular RAG.</p>
<p>Modular RAG is comprised of various functional modules. Thus, modular RAG is not standalone. Instead, different RAG patterns are composed of different modules.</p>
<p>For example, the following animation shows:
🥚 The original naive RAG paradigm consists of the “Retrieval”, "Augmentation," and "Generation" modules.</p>
<p>🐣 After naive RAG has shown some limitations, advanced RAG has emerged as a new paradigm. A typical pattern of Advanced RAG builds upon the foundation of Naive RAG by adding “Rewrite” and “Rerank” modules.</p>
<p>🐓 Different RAG patterns, such as DSP, can be composed of entirely different modules.</p>
<p>The modular RAG paradigm is slowly becoming the norm in the RAG domain due to its versatility and flexibility, allowing:</p>
<ul>
<li class="">the adaption of modules within the RAG process to suit your specific problem,</li>
<li class="">for a serialized pipeline or an end-to-end training approach across multiple modules.</li>
</ul>
<p>I definitely recommend checking out the full survey if you want to catch up on recent advancements in the RAG domain.</p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2312.10997" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2312.10997" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper22#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper22&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Matryoshka Representation Learning]]></title>
            <link>https://weaviate.io/papers/paper21</link>
            <guid>https://weaviate.io/papers/paper21</guid>
            <pubDate>Mon, 29 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Overview of OpenAI's New Truncatable - Matryoshka Embeddings]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-237ed4b707a303e4ad3353daaf4edab8.jpeg" width="1125" height="933" class="img_ev3q"></p>
<!-- -->
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="an-overview-of-openais-new-truncatable---matryoshka-embeddings">An Overview of OpenAI's New Truncatable - Matryoshka Embeddings🪆<a href="https://weaviate.io/papers/paper21#an-overview-of-openais-new-truncatable---matryoshka-embeddings" class="hash-link" aria-label="Direct link to An Overview of OpenAI's New Truncatable - Matryoshka Embeddings🪆" title="Direct link to An Overview of OpenAI's New Truncatable - Matryoshka Embeddings🪆" translate="no">​</a></h3>
<p>OpenAI recently announced embeddings that you can simply use chunks of (say the first 8, 16, 32, 64, 128 or 256 ... dimensions of the total 2048d vector) they use Matryoshka representation learning(MRL).</p>
<p>This is how they work, In Short⏩:</p>
<ul>
<li class="">
<p>MLR allows you to use a subset of the dimensions of the embedding vector - earlier dimensions store more information than dimensions later on in the vector, which simply add more details</p>
</li>
<li class="">
<p>You can understand how this works by the analogy of trying to classify an image at multiple resolutions - the lower res give high-level info and the higher res add details - Human perception of the natural world also has a naturally coarse-to-fine granularity</p>
</li>
<li class="">
<p>This is done by modifying the loss function which is optimized. If previously the loss function was L, for MRL we break down the Loss function into the sum of the losses on individual vector dimension ranges: Loss_Total =  L(upto 8d) + L(upto 16d) + L(upto 32d) + ... + L(upto 2048d) - Now there is incentive for the model to capture information in each sub-section of the vec.</p>
</li>
<li class="">
<p>After modifying the loss you get these truncatable vectors for free/no additional costs - this works on almost all loss functions and pre-existing models can be finetuned to output MRL vectors! - super easy-to-adopt technique</p>
</li>
<li class="">
<p>You can actually use any slice of dimensions, not just 8, 16,32 ... - b/c information is diffused in an interpolative fashion; so you can choose an arbitrary-sized chunk dimension that falls between the chosen granularity of the representations</p>
</li>
</ul>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2205.13147" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2205.13147" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper21#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper21&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[A Simple Overview of the LLM Training Steps 🔡]]></title>
            <link>https://weaviate.io/papers/paper20</link>
            <guid>https://weaviate.io/papers/paper20</guid>
            <pubDate>Wed, 24 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[A breakdown of the different training steps that go into creating a LLM.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-44342a5706c59b4b5f0df7dd5b061320.jpeg" width="1200" height="668" class="img_ev3q"></p>
<!-- -->
<p>A Simple Overview of the LLM Training Steps:🔡</p>
<ol>
<li class="">Unsupervised Pretraining:</li>
</ol>
<blockquote>
<blockquote>
<p>High quantity, low quality data
The model is trained to predict the next token for trillions of tokens.
Produces what is called the foundation or base model.</p>
</blockquote>
</blockquote>
<ol start="2">
<li class="">Supervised Finetuning:</li>
</ol>
<blockquote>
<blockquote>
<p>Low quantity, high quality {prompt, response}
Enables the model to be finetuned for dialogue - turning the base model into a chatbot
Often referred to as instruction tuning</p>
</blockquote>
</blockquote>
<ol start="3">
<li class="">Reinforcement Learning from Human Feedback (RLHF): - lots of innovation going on here (will cover DPO, PTO, and KTO soon)</li>
</ol>
<p>This is a two-step process:</p>
<p>a. Train a reward model to act as a scoring function:</p>
<blockquote>
<blockquote>
<p>This model will take in a prompt + response and provide a score of how good it is.
Human labelers are asked to pick good vs. bad responses and this data is used to train a model.</p>
</blockquote>
</blockquote>
<p>b. Optimize LLM to generate responses for which the reward model will give high scores.</p>
<blockquote>
<blockquote>
<p>Use an iterative procedure to update a part of the model such that:</p>
<ol>
<li class="">Produces outputs with higher score</li>
<li class="">Outputs that are not too far away from the SFT model from Step 2</li>
<li class="">Outputs that aren't getting worse a text completion</li>
</ol>
</blockquote>
</blockquote>
<p>Specifically for this phase it is better to think of this as learning an optimal strategy/policy for predicting a probability distribution over tokens and we want to tweak this distribution to produce higher quality text, here the:</p>
<blockquote>
<blockquote>
<p>The policy is a language model that takes in a prompt and returns a probability distribution over text.
The action space of this policy is all the tokens corresponding to the vocabulary of the language model (~50k tokens)
The observation space: distribution of possible input token sequences
The reward Model is a combination of the preference model(score higher) and a constraint on policy shift(don't change too much, get worse at text completion).</p>
</blockquote>
</blockquote>
<p>RLHF Learning Resources:</p>
<ol>
<li class="">
<p><a href="https://arxiv.org/pdf/2203.02155.pdf" target="_blank" rel="noopener noreferrer" class="">InstructGPT Paper</a></p>
</li>
<li class="">
<p><a href="https://arxiv.org/pdf/2204.05862.pdf" target="_blank" rel="noopener noreferrer" class="">RLHF Paper Anthropic</a></p>
</li>
<li class="">
<p><a href="https://openai.com/research/instruction-following" target="_blank" rel="noopener noreferrer" class="">OpenAI Blog</a></p>
</li>
<li class="">
<p><a href="https://huyenchip.com/2023/05/02/rlhf.html" target="_blank" rel="noopener noreferrer" class="">RLHF Blog Chip Huyen</a></p>
</li>
<li class="">
<p><a href="https://interconnects.ai/p/how-rlhf-works" target="_blank" rel="noopener noreferrer" class="">RLHF Nathan Lambert</a></p>
</li>
<li class="">
<p><a href="https://youtube.com/watch?v=bZQun8Y4L2A&amp;ab_channel=MicrosoftDeveloper" target="_blank" rel="noopener noreferrer" class="">Karpathy Talk</a></p>
</li>
</ol>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper20#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper20&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Using a 7B Model + RAG to Identify and Edit Word-level Hallucinations]]></title>
            <link>https://weaviate.io/papers/paper19</link>
            <guid>https://weaviate.io/papers/paper19</guid>
            <pubDate>Sat, 20 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Finetuning a 7B model to outperform GPT-4 for hallucination detection.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-37c5610019a34e8cd6b12a9a47e84826.jpeg" width="1200" height="1066" class="img_ev3q"></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="using-a-7b-model--rag-to-identify-and-edit-word-level-hallucinations-in-llms-better-then-gpt-4">Using a 7B Model + RAG to Identify and Edit Word-level Hallucinations in LLMs better then GPT-4:<a href="https://weaviate.io/papers/paper19#using-a-7b-model--rag-to-identify-and-edit-word-level-hallucinations-in-llms-better-then-gpt-4" class="hash-link" aria-label="Direct link to Using a 7B Model + RAG to Identify and Edit Word-level Hallucinations in LLMs better then GPT-4:" title="Direct link to Using a 7B Model + RAG to Identify and Edit Word-level Hallucinations in LLMs better then GPT-4:" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short⏩:<a href="https://weaviate.io/papers/paper19#in-short" class="hash-link" aria-label="Direct link to In Short⏩:" title="Direct link to In Short⏩:" translate="no">​</a></h3>
<blockquote>
<p>Train a model that consists of a Retreiver and a Language Model:</p>
</blockquote>
<blockquote>
<blockquote>
<p>The retriever, Mret, takes the original output you want to check to hallucination (y) and optionally input prompt (x) and retrieves top relevant documents (C). So C = Mret(x, y). This can be a vector database like Weaviate for example.</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>The detector and editor, Medit, takes in the context - (C), input - (x) and output - (y) and detects (and if possible also edits/corrects) factual errors in (y) given the retrieved context (C): y* = Medit(x, y, C).</p>
</blockquote>
</blockquote>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="️training">🏋️Training:<a href="https://weaviate.io/papers/paper19#%EF%B8%8Ftraining" class="hash-link" aria-label="Direct link to 🏋️Training:" title="Direct link to 🏋️Training:" translate="no">​</a></h3>
<blockquote>
<p>Create a synthetic hallucination dataset of 35k C = context, y=incorrect output, y*=annotated fixed output -&gt; (C, y, y*)</p>
</blockquote>
<blockquote>
<p>Magic Synthetic Dataset Creation:</p>
</blockquote>
<blockquote>
<blockquote>
<p>GPT-4 is few-shot prompted to add different types of errors to a passage</p>
</blockquote>
</blockquote>
<blockquote>
<blockquote>
<p>It is also instructed to mark phrases or sentences for deletion along with their error type and insert phrases and sentences along with insertion tags</p>
</blockquote>
</blockquote>
<blockquote>
<p>Start off with Llama2-Chat 7B to initialize Medit and then train on (C, y, y∗)</p>
</blockquote>
<blockquote>
<p>Medit takes in (C, y) as input and learns to predict the edited outputs with tags to represent error type y∗ using standard language modeling objective.</p>
</blockquote>
<blockquote>
<p>The model, once trained, can identify different types of hallucination and mark which words they come from - it also suggests edits to improve factuality.</p>
</blockquote>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="result-">Result 📈:<a href="https://weaviate.io/papers/paper19#result-" class="hash-link" aria-label="Direct link to Result 📈:" title="Direct link to Result 📈:" translate="no">​</a></h3>
<p>The model has a fine-grained hallucination detection accuracy 46.5% while it's binary acc.{hallucination, no hallucination} is 79%.</p>
<p>For comparison ChatGPT has a fine-grained hallucination detection acc. of 21.5% (59% binary acc) w/o RAG and 26%(68.5% binary hall detect) w/ RAG</p>
<p><a href="https://github.com/abhika-m/FAVA" target="_blank" rel="noopener noreferrer" class="">💻Code</a></p>
<p><a href="https://huggingface.co/datasets/fava-uw/fava-data" target="_blank" rel="noopener noreferrer" class="">🔷Data</a></p>
<p><a href="https://huggingface.co/fava-uw/fava-model" target="_blank" rel="noopener noreferrer" class="">🏗️Model</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2401.06855" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2401.06855" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper19#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper19&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Fine-grained Hallucination Detection and Editing for Language Models]]></title>
            <link>https://weaviate.io/papers/paper18</link>
            <guid>https://weaviate.io/papers/paper18</guid>
            <pubDate>Fri, 19 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Provides a taxonomy of different types of hallucinations.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-22ea2ae5101f8e873d45a26da3e4268c.jpeg" width="1200" height="813" class="img_ev3q"></p>
<!-- -->
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-breakdown-of-the-different-types-of-hallucinations-from-ai2">A breakdown of the different types of hallucinations from AI2:🍄<a href="https://weaviate.io/papers/paper18#a-breakdown-of-the-different-types-of-hallucinations-from-ai2" class="hash-link" aria-label="Direct link to A breakdown of the different types of hallucinations from AI2:🍄" title="Direct link to A breakdown of the different types of hallucinations from AI2:🍄" translate="no">​</a></h3>
<ol>
<li class="">Verifiably Factually Wrong ❌</li>
</ol>
<ul>
<li class="">
<p>Entity: an entity in a statement is incorrect (eg. Christmas falls on Nov. 25th)</p>
</li>
<li class="">
<p>Relation: semantic relationship in a statement is incorrect (eg. The mouse ate the cat.)</p>
</li>
<li class="">
<p>Contradictory: statements that entirely contradict relevant evidence from the web (eg. Raptors are yet to win an NBA final.)</p>
</li>
</ul>
<ol start="2">
<li class="">Unverifiable Types of Hallucinations ⁉️</li>
</ol>
<ul>
<li class="">
<p>Invented: statements of concepts that do not exist in world knowledge (eg. MJ created the sideways somersault)</p>
</li>
<li class="">
<p>Subjective: Statement that lacks universal validity - basically an opinion (eg. The Raptors are the best NBA team)</p>
</li>
<li class="">
<p>Unverifiable: potentially factual statement but cannot be grounded in world evidence(eg. Jensen sleeps in a leather jacket.)</p>
</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="word-vs-sentence-level">🔍Word vs. Sentence Level:<a href="https://weaviate.io/papers/paper18#word-vs-sentence-level" class="hash-link" aria-label="Direct link to 🔍Word vs. Sentence Level:" title="Direct link to 🔍Word vs. Sentence Level:" translate="no">​</a></h3>
<blockquote>
<p>Entity and Relation are usually word level, and so can be fixed with small edits if you know where they occur.</p>
</blockquote>
<blockquote>
<p>Contradictory, Invented, Subjective, and Unverifiable are often sentence level and thus need to be removed completely to fix the issue.</p>
</blockquote>
<p><a href="https://fine-grained-hallucination.github.io/" target="_blank" rel="noopener noreferrer" class="">💻Code</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2401.06855" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2401.06855" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper18#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper18&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Long-Context Retrieval Models with Monarch Mixer]]></title>
            <link>https://weaviate.io/papers/paper16</link>
            <guid>https://weaviate.io/papers/paper16</guid>
            <pubDate>Mon, 15 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[32k context length retreival models with sub-quadratic attention mechanism.]]></description>
            <content:encoded><![CDATA[<p><img decoding="async" loading="lazy" alt="A preview of the paper" src="https://weaviate.io/assets/images/hero-0b709c8abd18ca664ec3b01009b53e50.jpeg" width="1200" height="874" class="img_ev3q"></p>
<!-- -->
<p>A breakdown of the Long Context Retrieval Embedding Models from Stanford!💥</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="in-short">In Short⏩:<a href="https://weaviate.io/papers/paper16#in-short" class="hash-link" aria-label="Direct link to In Short⏩:" title="Direct link to In Short⏩:" translate="no">​</a></h3>
<ol>
<li class="">
<p>They release 3 long context(2k/8k/32k) BERT-like encoder embedding models on HuggingFace</p>
</li>
<li class="">
<p>The models are only 80M params and outperform MUCH larger models (4-85x larger)</p>
</li>
<li class="">
<p>Accessible via @togethercompute endpoints and integrated into @llama_index and @LangChainAI</p>
</li>
<li class="">
<p>They also release LoCo a long context retrieval benchmark.</p>
</li>
</ol>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="️architechtural-details">🏗️Architechtural Details:<a href="https://weaviate.io/papers/paper16#%EF%B8%8Farchitechtural-details" class="hash-link" aria-label="Direct link to 🏗️Architechtural Details:" title="Direct link to 🏗️Architechtural Details:" translate="no">​</a></h3>
<ol>
<li class="">
<p>They replace the Attention and MLP blocks in the transformer architecture with diagonal block matrix (Monarch Matrices -M2) operations which are hardware optimized and subquadratic in the sequence length - O(N^(1.5))</p>
</li>
<li class="">
<p>This enables scaling sequence length and model parameters better.</p>
</li>
</ol>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="training-details">🪃Training Details:<a href="https://weaviate.io/papers/paper16#training-details" class="hash-link" aria-label="Direct link to 🪃Training Details:" title="Direct link to 🪃Training Details:" translate="no">​</a></h3>
<ol>
<li class="">
<p>These M2 models are trained for long context retrieval on a mixture of long and short context tasks data - surprisingly only training on long context doesn't work.</p>
</li>
<li class="">
<p>Use a cosine similarity loss instead of the trusty supervised contrastive training loss.</p>
<blockquote>
<p>This loss function. can be computed independently per datapoint in a batch instead of needing to sum over all negative examples in a batch.</p>
</blockquote>
<blockquote>
<p>Thus training can be scaled for large batch sizes of long context inputs without OOM'ing</p>
</blockquote>
</li>
</ol>
<p><a href="https://hazyresearch.stanford.edu/blog/2024-01-11-m2-bert-retrieval" target="_blank" rel="noopener noreferrer" class="">📜Blog</a></p>
<p><a href="https://github.com/HazyResearch/m2" target="_blank" rel="noopener noreferrer" class="">🧑‍💻Code</a></p>
<p><a href="https://huggingface.co/togethercomputer/m2-bert-80M-32k-retrieval" target="_blank" rel="noopener noreferrer" class="">🔷Models</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/abs/2310.12109" download="">🔗 arXiv Link</a></p>
<p><a class="btn_VbJ1 btnMain_ywTD" href="https://arxiv.org/pdf/2310.12109" download="">📜 Download paper</a></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="ready-to-start-building">Ready to start building?<a href="https://weaviate.io/papers/paper16#ready-to-start-building" class="hash-link" aria-label="Direct link to Ready to start building?" title="Direct link to Ready to start building?" translate="no">​</a></h2>
<p>Check out the <a href="https://docs.weaviate.io/weaviate/quickstart" target="_blank" rel="noopener noreferrer" class="">Quickstart tutorial</a>, or <a href="https://console.weaviate.cloud/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog_signup&amp;utm_content=%2Fpapers%2Fpaper16&amp;utm_term=ready-to-start-building" target="_blank" rel="noopener noreferrer">sign up for a free Weaviate Cloud account</a>.</p>
<!-- -->
<div class="communityWrapper_ZpuS"><div class="container_sUl4"><div class="wrapper_FyvH"><div class="rightSide_UqS8"><div class="socialBox_W1XR"><a href="https://github.com/weaviate/weaviate" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="github_DEOB"></div><p class="text_g9NY">GitHub</p></a></div><div class="socialBox_W1XR"><a href="https://forum.weaviate.io/" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="forum_pUq6"></div><p class="text_g9NY">Forum</p></a></div><div class="socialBox_W1XR"><a href="https://twitter.com/weaviate_io" target="_blank" rel="noopener noreferrer" class="mobileSocialBox_UAY5"><div class="twitter_ewvw"></div><p class="text_g9NY">X (Twitter)</p></a></div></div><div class="leftSide_WlMC"><h2 class="communityHeader_jLni">Don't want to miss another blog post?</h2><span class="rightText_noBq"><p>Sign up for our bi-weekly newsletter to stay updated!</p> <br>By submitting, I agree to the<!-- --> <a href="https://weaviate.io/service">Terms of Service </a>and<!-- --> <a href="https://weaviate.io/privacy">Privacy Policy</a>.</span><div class="communityForm_pedn"><iframe src="https://embeds.beehiiv.com/15b21ebd-decd-433b-ada8-2d405e345f2e?slim=true" data-test-id="beehiiv-embed" frameborder="0" scrolling="no" style="margin:0;border-radius:0px;button-colour:#61BD73;background-color:transparent;width:100%"></iframe></div></div></div></div></div>
<!-- -->
]]></content:encoded>
        </item>
    </channel>
</rss>