
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Ashim Sharma</title>
      <link>https://ashimsharma10.github.io/blog</link>
      <description>Software engineer sharing projects, notes, and guides on ML infrastructure.</description>
      <language>en-us</language>
      <managingEditor>sharmaashim00@gmail.com (Ashim Sharma)</managingEditor>
      <webMaster>sharmaashim00@gmail.com (Ashim Sharma)</webMaster>
      <lastBuildDate>Sun, 06 Sep 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://ashimsharma10.github.io/tags/reasoning/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://ashimsharma10.github.io/blog/rlvr-and-the-experience-era</guid>
    <title>RLVR and the Experience Era of LLMs</title>
    <link>https://ashimsharma10.github.io/blog/rlvr-and-the-experience-era</link>
    <description>Post-training has moved from imitating human labels to learning from verifiable outcomes. This write-up covers the RLVR objective, GRPO and the choice of KL penalty, advantage collapse and the methods that recover the lost gradient, reward hacking and noisy verifiers, outcome versus process reward models, the production data flywheel that turns failures into weight updates, long chain-of-thought and latent reasoning, and the OpenRLHF stack that runs it all.</description>
    <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
    <author>sharmaashim00@gmail.com (Ashim Sharma)</author>
    <category>rlvr</category><category>reinforcement-learning</category><category>grpo</category><category>llm</category><category>continual-learning</category><category>reasoning</category>
  </item>

    </channel>
  </rss>
