<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Selection Guide on 卓琪的开发笔记</title>
    <link>https://zhuoqidev.com/tags/selection-guide/</link>
    <description>Recent content in Selection Guide on 卓琪的开发笔记</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <copyright>© 2026 Liu ZhuoQi</copyright>
    <lastBuildDate>Sun, 19 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://zhuoqidev.com/tags/selection-guide/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>How to Choose an LLM Inference Engine — A 2026 Map from Local Single-GPU to PD Disaggregation</title>
      <link>https://zhuoqidev.com/en/posts/llm-inference-engine-selection/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      
      <guid>https://zhuoqidev.com/en/posts/llm-inference-engine-selection/</guid>
      <description>&lt;p&gt;Aliyun&amp;rsquo;s CAP has a piece on picking an inference engine that narrows the field to four: &lt;strong&gt;Ollama, vLLM, SGLang, and Hugging Face Pipeline&lt;/strong&gt;. In 2024, that framing was fine.&lt;/p&gt;&#xA;&lt;p&gt;By 2026, it&amp;rsquo;s missing half the map. NVIDIA&amp;rsquo;s TensorRT-LLM has completed its &amp;ldquo;PyTorch-ification,&amp;rdquo; SGLang became famous as the first open-source project to reproduce DeepSeek&amp;rsquo;s large-scale deployment, Hugging Face slapped a &amp;ldquo;maintenance mode&amp;rdquo; banner on TGI and told you to switch to vLLM — and the real throughline of the entire 2025 inference landscape can be summed up in one word: &lt;strong&gt;disaggregate&lt;/strong&gt;.&lt;/p&gt;</description>
      
    </item>
    
  </channel>
</rss>
