<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>基准测试 on 知识铺的博客</title>
    <link>https://index.zshipu.com/geek/tags/%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/</link>
    <description>Recent content in 基准测试 on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Tue, 01 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/geek/tags/%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>AWS 发布 Aws-Bench，把云任务设为智能代理评估核心</title>
      <link>https://index.zshipu.com/geek/post/20260901/AWS-%E5%8F%91%E5%B8%83-Aws-Bench%E6%8A%8A%E4%BA%91%E4%BB%BB%E5%8A%A1%E8%AE%BE%E4%B8%BA%E6%99%BA%E8%83%BD%E4%BB%A3%E7%90%86%E8%AF%84%E4%BC%B0%E6%A0%B8%E5%BF%83/</link>
      <pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/geek/post/20260901/AWS-%E5%8F%91%E5%B8%83-Aws-Bench%E6%8A%8A%E4%BA%91%E4%BB%BB%E5%8A%A1%E8%AE%BE%E4%B8%BA%E6%99%BA%E8%83%BD%E4%BB%A3%E7%90%86%E8%AF%84%E4%BC%B0%E6%A0%B8%E5%BF%83/</guid>
      <description>AWS 发布了 Aws-Bench，用于评估云任务中的智能代理。这一发布把云任务直接设为评估对象，与此前通用 AI 基准形成明确区分。 AI Agent 评估长期缺少云任务真实场景 当前 AI Agent 评估面临的最大痛点在于脱离实际运行环境。多数现有基准把重点放在语言理解、数学推理或简单工具调用上，测试环境往往是沙盒模拟或</description>
    </item>
  </channel>
</rss>
