<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>TDD on 知识铺的博客</title>
    <link>https://index.zshipu.com/ai001/tags/TDD/</link>
    <description>Recent content in TDD on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Tue, 01 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/ai001/tags/TDD/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Agent开发者口中的Eval，本质就是写测试用例</title>
      <link>https://index.zshipu.com/ai001/post/20260901/Agent%E5%BC%80%E5%8F%91%E8%80%85%E5%8F%A3%E4%B8%AD%E7%9A%84Eval%E6%9C%AC%E8%B4%A8%E5%B0%B1%E6%98%AF%E5%86%99%E6%B5%8B%E8%AF%95%E7%94%A8%E4%BE%8B/</link>
      <pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai001/post/20260901/Agent%E5%BC%80%E5%8F%91%E8%80%85%E5%8F%A3%E4%B8%AD%E7%9A%84Eval%E6%9C%AC%E8%B4%A8%E5%B0%B1%E6%98%AF%E5%86%99%E6%B5%8B%E8%AF%95%E7%94%A8%E4%BE%8B/</guid>
      <description>Agent 开发者口中的 Eval，本质就是写测试用例 Agent 开发者口中的 Eval，本质就是写测试用例。和传统单测一样，它要求先为模型输出设定验证标准，再检查实际结果是否符合预期。从 TDD 的测试先行，到 Sentry 的运行时监控，这套思路直接把软件工程的测试方法论搬到了大模型 Agent 场景。 在 LLM 和 AI Agent 的技术讨论中，Eva</description>
    </item>
  </channel>
</rss>
