<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>多代理系统 on 知识铺的博客</title>
    <link>https://index.zshipu.com/ai002/tags/%E5%A4%9A%E4%BB%A3%E7%90%86%E7%B3%BB%E7%BB%9F/</link>
    <description>Recent content in 多代理系统 on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Sat, 05 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/ai002/tags/%E5%A4%9A%E4%BB%A3%E7%90%86%E7%B3%BB%E7%BB%9F/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Claude代理用共享待办列表完成费马大定理形式化证明</title>
      <link>https://index.zshipu.com/ai002/post/20260905/Claude%E4%BB%A3%E7%90%86%E7%94%A8%E5%85%B1%E4%BA%AB%E5%BE%85%E5%8A%9E%E5%88%97%E8%A1%A8%E5%AE%8C%E6%88%90%E8%B4%B9%E9%A9%AC%E5%A4%A7%E5%AE%9A%E7%90%86%E5%BD%A2%E5%BC%8F%E5%8C%96%E8%AF%81%E6%98%8E/</link>
      <pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai002/post/20260905/Claude%E4%BB%A3%E7%90%86%E7%94%A8%E5%85%B1%E4%BA%AB%E5%BE%85%E5%8A%9E%E5%88%97%E8%A1%A8%E5%AE%8C%E6%88%90%E8%B4%B9%E9%A9%AC%E5%A4%A7%E5%AE%9A%E7%90%86%E5%BD%A2%E5%BC%8F%E5%8C%96%E8%AF%81%E6%98%8E/</guid>
      <description>Anthropic 的 Claude 代理团队用 11 天生成了 1300 万行 Lean 代码、近 3 万个中间定理和 60 亿输出 token，完成了费马大定理的完整机器验证证明。但在获得共享待办列表前，所有尝试均告失败，问题不在于模型能力。 费马大定理的形式化工作长期被视为数学形式化领域的硬骨头。Anthropic 这次直接让多个 Claude 代理自主协作，</description>
    </item>
    <item>
      <title>LLMPvP竞技场：AI代理通过MCP在围棋与象棋中实时对战排名</title>
      <link>https://index.zshipu.com/ai002/post/20260904/LLMPvP%E7%AB%9E%E6%8A%80%E5%9C%BAAI%E4%BB%A3%E7%90%86%E9%80%9A%E8%BF%87MCP%E5%9C%A8%E5%9B%B4%E6%A3%8B%E4%B8%8E%E8%B1%A1%E6%A3%8B%E4%B8%AD%E5%AE%9E%E6%97%B6%E5%AF%B9%E6%88%98%E6%8E%92%E5%90%8D/</link>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/ai002/post/20260904/LLMPvP%E7%AB%9E%E6%8A%80%E5%9C%BAAI%E4%BB%A3%E7%90%86%E9%80%9A%E8%BF%87MCP%E5%9C%A8%E5%9B%B4%E6%A3%8B%E4%B8%8E%E8%B1%A1%E6%A3%8B%E4%B8%AD%E5%AE%9E%E6%97%B6%E5%AF%B9%E6%88%98%E6%8E%92%E5%90%8D/</guid>
      <description>传统LLM基准仅凭一次提示打分，无法揭示模型能否在对手持续惩罚下维持跨多步计划。开发者因此搭建LLMPvP竞技场，让AI代理通过MCP在国际象棋和围棋中互相对战，每局结束后用Glicko-2更新各游戏类型排名，且API密钥从不进入平台服务器。 静态基准测试把模型能力简化成单次输入输</description>
    </item>
  </channel>
</rss>
