<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>推理成本优化 on 知识铺的博客</title>
    <link>https://index.zshipu.com/geek/tags/%E6%8E%A8%E7%90%86%E6%88%90%E6%9C%AC%E4%BC%98%E5%8C%96/</link>
    <description>Recent content in 推理成本优化 on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Mon, 31 Aug 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/geek/tags/%E6%8E%A8%E7%90%86%E6%88%90%E6%9C%AC%E4%BC%98%E5%8C%96/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Amazon EC2 上用 NVIDIA MPS 将 ASR 推理成本降低 75%</title>
      <link>https://index.zshipu.com/geek/post/20260831/Amazon-EC2-%E4%B8%8A%E7%94%A8-NVIDIA-MPS-%E5%B0%86-ASR-%E6%8E%A8%E7%90%86%E6%88%90%E6%9C%AC%E9%99%8D%E4%BD%8E-75/</link>
      <pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/geek/post/20260831/Amazon-EC2-%E4%B8%8A%E7%94%A8-NVIDIA-MPS-%E5%B0%86-ASR-%E6%8E%A8%E7%90%86%E6%88%90%E6%9C%AC%E9%99%8D%E4%BD%8E-75/</guid>
      <description>GPU 共享直接决定 ASR 生产推理的每美元吞吐量 生产环境里的 ASR 系统，真正要算的不是单次推理能跑多快，而是每个 GPU 在延迟不崩的情况下能处理多少请求。这直接决定了每美元能买到多少吞吐量。信号明确指出，当 ASR 流水线推到生产后，核心问题不再是速度，而是“能从每个 GPU 榨取多少吞吐量”。 GPU 共享技术正是解决这</description>
    </item>
  </channel>
</rss>
