<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>大模型评测 on 知识铺的博客</title>
    <link>https://index.zshipu.com/ai001/tags/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/</link>
    <description>Recent content in 大模型评测 on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Sun, 30 Aug 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/ai001/tags/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>评测集不是一堆问题，而是可管理的测试资产</title>
      <link>https://index.zshipu.com/ai001/post/20260830/%E8%AF%84%E6%B5%8B%E9%9B%86%E4%B8%8D%E6%98%AF%E4%B8%80%E5%A0%86%E9%97%AE%E9%A2%98%E8%80%8C%E6%98%AF%E5%8F%AF%E7%AE%A1%E7%90%86%E7%9A%84%E6%B5%8B%E8%AF%95%E8%B5%84%E4%BA%A7/</link>
      <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai001/post/20260830/%E8%AF%84%E6%B5%8B%E9%9B%86%E4%B8%8D%E6%98%AF%E4%B8%80%E5%A0%86%E9%97%AE%E9%A2%98%E8%80%8C%E6%98%AF%E5%8F%AF%E7%AE%A1%E7%90%86%E7%9A%84%E6%B5%8B%E8%AF%95%E8%B5%84%E4%BA%A7/</guid>
      <description>E01暴露的三大评测缺陷 E01中研究者手工构造了10个问题用于测试大模型，很快发现这些问题存在明显缺陷。首先，它们完全是“随手想”的，没有任何系统化的分类或覆盖逻辑。这导致评测结果无法说明模型在不同能力维度上的真实表现，模型答对几个问题就很难推断其整体水平。 其次，这些问题严重脱离</description>
    </item>
  </channel>
</rss>
