<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>SageMaker on 知识铺的博客</title>
    <link>https://index.zshipu.com/ai002/tags/SageMaker/</link>
    <description>Recent content in SageMaker on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Thu, 10 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/ai002/tags/SageMaker/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Amazon SageMaker Inference 推出 prefix-aware routing 降低 LLM 延迟</title>
      <link>https://index.zshipu.com/ai002/post/20260910/Amazon-SageMaker-Inference-%E6%8E%A8%E5%87%BA-prefix-aware-routing-%E9%99%8D%E4%BD%8E-LLM-%E5%BB%B6%E8%BF%9F/</link>
      <pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai002/post/20260910/Amazon-SageMaker-Inference-%E6%8E%A8%E5%87%BA-prefix-aware-routing-%E9%99%8D%E4%BD%8E-LLM-%E5%BB%B6%E8%BF%9F/</guid>
      <description>Amazon SageMaker Inference 推出 prefix-aware routing 降低 LLM 延迟 Amazon SageMaker Inference 现已提供 prefix-aware routing 路由策略。该策略将共享相同 prompt prefix 的请求发送到同一实例，使 KV cache 保持温暖。在 Llama 3.1 70B 的基准测试中，这一方法将 P50 time-to-first-token 最多降低了 77%。 prefix-aware routing 的路由策略定义 prefix-aware routing 是一种新的路由策略。它依据请求中的 prompt prefix 来决定目标实例。所有携带相同 prefix 的推理请求会被定向到同一个</description>
    </item>
    <item>
      <title>AWS SageMaker AI 上小 LLM 推理基准测试：G7 与 G5、G6 实例对比</title>
      <link>https://index.zshipu.com/ai002/post/20260909/AWS-SageMaker-AI-%E4%B8%8A%E5%B0%8F-LLM-%E6%8E%A8%E7%90%86%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95G7-%E4%B8%8E-G5G6-%E5%AE%9E%E4%BE%8B%E5%AF%B9%E6%AF%94/</link>
      <pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai002/post/20260909/AWS-SageMaker-AI-%E4%B8%8A%E5%B0%8F-LLM-%E6%8E%A8%E7%90%86%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95G7-%E4%B8%8E-G5G6-%E5%AE%9E%E4%BE%8B%E5%AF%B9%E6%AF%94/</guid>
      <description>两个 30B MoE 模型的基准测试设置 AWS 在 SageMaker AI 上针对两个 30B Mixture-of-Experts 模型展开基准测试。这两个模型分别是 Qwen3-Coder-30B 和 NVIDIA Nemotron-3-Nano-30B。测试覆盖 G5、G6、G6e 以及 G7 GPU 实例，核心指标包括吞吐量、延迟和每 token 成本。 测试直接在 Amazon SageMaker AI 环境中运行，目的是让开发者了解不同 GPU 实例在实际小 LLM 推理场景下</description>
    </item>
    <item>
      <title>托管 MLflow 与 SageMaker AI Model Registry 同步扩展至跨账户治理</title>
      <link>https://index.zshipu.com/ai002/post/20260909/%E6%89%98%E7%AE%A1-MLflow-%E4%B8%8E-SageMaker-AI-Model-Registry-%E5%90%8C%E6%AD%A5%E6%89%A9%E5%B1%95%E8%87%B3%E8%B7%A8%E8%B4%A6%E6%88%B7%E6%B2%BB%E7%90%86/</link>
      <pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai002/post/20260909/%E6%89%98%E7%AE%A1-MLflow-%E4%B8%8E-SageMaker-AI-Model-Registry-%E5%90%8C%E6%AD%A5%E6%89%A9%E5%B1%95%E8%87%B3%E8%B7%A8%E8%B4%A6%E6%88%B7%E6%B2%BB%E7%90%86/</guid>
      <description>托管 MLflow 与 SageMaker AI Model Registry 同步扩展至跨账户治理 托管 MLflow 与 Amazon SageMaker AI Model Registry 的同步现已支持更丰富的模型元数据，包括训练指标、评估结果、推理规范和血缘关系，并实现生命周期阶段提升。在 Part 1 单账户治理基础上，Part 2 进一步将此同步扩展到跨账户场景，采用 hub-and-spoke 模式通过 AWS RAM 实现集中治理，以及另一种跨账户拓扑。 单账</description>
    </item>
  </channel>
</rss>
