<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>KV 缓存 on GPT资讯  --  知识铺</title>
    <link>https://index.zshipu.com/gpt/tags/KV-%E7%BC%93%E5%AD%98/</link>
    <description>Recent content in KV 缓存 on GPT资讯  --  知识铺</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Sat, 12 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/gpt/tags/KV-%E7%BC%93%E5%AD%98/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Amazon SageMaker Inference 推出前缀感知路由 降低 LLM 延迟</title>
      <link>https://index.zshipu.com/gpt/post/20260912/Amazon-SageMaker-Inference-%E6%8E%A8%E5%87%BA%E5%89%8D%E7%BC%80%E6%84%9F%E7%9F%A5%E8%B7%AF%E7%94%B1-%E9%99%8D%E4%BD%8E-LLM-%E5%BB%B6%E8%BF%9F/</link>
      <pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/gpt/post/20260912/Amazon-SageMaker-Inference-%E6%8E%A8%E5%87%BA%E5%89%8D%E7%BC%80%E6%84%9F%E7%9F%A5%E8%B7%AF%E7%94%B1-%E9%99%8D%E4%BD%8E-LLM-%E5%BB%B6%E8%BF%9F/</guid>
      <description>Amazon SageMaker Inference 推出前缀感知路由 降低 LLM 延迟 Amazon SageMaker Inference 现在提供前缀感知路由这一新路由策略。它将共享相同提示前缀的请求发送到同一实例，从而让 KV 缓存保持温暖。在 Llama 3.1 70B 模型基准测试中，这一功能将 P50 时间到首 token 减少高达 77%，同时提高了 KV 缓存命中率。 SageMaker Inference 前缀感知路由功能发布 Amazon SageMaker Inference 正式推出前缀感知路由功能。</description>
    </item>
  </channel>
</rss>
