
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Ashim Sharma</title>
      <link>https://ashimsharma10.github.io/blog</link>
      <description>Software engineer sharing projects, notes, and guides on ML infrastructure.</description>
      <language>en-us</language>
      <managingEditor>sharmaashim00@gmail.com (Ashim Sharma)</managingEditor>
      <webMaster>sharmaashim00@gmail.com (Ashim Sharma)</webMaster>
      <lastBuildDate>Sun, 06 Sep 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://ashimsharma10.github.io/tags/reinforcement-learning/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://ashimsharma10.github.io/blog/rlvr-and-the-experience-era</guid>
    <title>RLVR and the Experience Era of LLMs</title>
    <link>https://ashimsharma10.github.io/blog/rlvr-and-the-experience-era</link>
    <description>Post-training has moved from imitating human labels to learning from verifiable outcomes. This write-up covers the RLVR objective, GRPO and the choice of KL penalty, advantage collapse and the methods that recover the lost gradient, reward hacking and noisy verifiers, outcome versus process reward models, the production data flywheel that turns failures into weight updates, long chain-of-thought and latent reasoning, and the OpenRLHF stack that runs it all.</description>
    <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
    <author>sharmaashim00@gmail.com (Ashim Sharma)</author>
    <category>rlvr</category><category>reinforcement-learning</category><category>grpo</category><category>llm</category><category>continual-learning</category><category>reasoning</category>
  </item>

  <item>
    <guid>https://ashimsharma10.github.io/blog/world-models</guid>
    <title>World Models and the Path to Physical AI</title>
    <link>https://ashimsharma10.github.io/blog/world-models</link>
    <description>A world model is a learned simulator. Give it a state and an action and it predicts what happens next. This write-up follows the idea from Ha and Schmidhuber&#39;s V+M+C agent through Dreamer, MuZero and TD-MPC, to Genie, Cosmos and V-JEPA, the drift that comes with imagining too far ahead, LeCun&#39;s case for predicting representations instead of pixels, and why robots need all of this.</description>
    <pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate>
    <author>sharmaashim00@gmail.com (Ashim Sharma)</author>
    <category>world-models</category><category>reinforcement-learning</category><category>robotics</category><category>physical-ai</category><category>jepa</category><category>model-based-rl</category>
  </item>

    </channel>
  </rss>
