
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Ashim Sharma</title>
      <link>https://ashimsharma10.github.io/blog</link>
      <description>Software engineer sharing projects, notes, and guides on ML infrastructure.</description>
      <language>en-us</language>
      <managingEditor>sharmaashim00@gmail.com (Ashim Sharma)</managingEditor>
      <webMaster>sharmaashim00@gmail.com (Ashim Sharma)</webMaster>
      <lastBuildDate>Thu, 13 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://ashimsharma10.github.io/tags/vllm/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://ashimsharma10.github.io/blog/vllm-how-a-token-gets-served</guid>
    <title>vLLM: How a Token Actually Gets Served</title>
    <link>https://ashimsharma10.github.io/blog/vllm-how-a-token-gets-served</link>
    <description>What happens inside a serving engine: why the GPU fills up with something other than the model, what paging the cache buys you, why reading a prompt and writing an answer fight each other, and which knobs settle the fight.</description>
    <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
    <author>sharmaashim00@gmail.com (Ashim Sharma)</author>
    <category>vllm</category><category>llm</category><category>inference</category><category>gpu</category><category>kv-cache</category>
  </item>

    </channel>
  </rss>
