<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>FreeToken on 知识铺的博客</title>
    <link>https://index.zshipu.com/geek001/tags/FreeToken/</link>
    <description>Recent content in FreeToken on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Sat, 05 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/geek001/tags/FreeToken/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>RTX 4060 跑 35B 模型每秒 39 Token，FreeToken 如何做到</title>
      <link>https://index.zshipu.com/geek001/post/20260905/RTX-4060-%E8%B7%91-35B-%E6%A8%A1%E5%9E%8B%E6%AF%8F%E7%A7%92-39-TokenFreeToken-%E5%A6%82%E4%BD%95%E5%81%9A%E5%88%B0/</link>
      <pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/geek001/post/20260905/RTX-4060-%E8%B7%91-35B-%E6%A8%A1%E5%9E%8B%E6%AF%8F%E7%A7%92-39-TokenFreeToken-%E5%A6%82%E4%BD%95%E5%81%9A%E5%88%B0/</guid>
      <description>RTX 4060 显卡运行 35B 模型达到每秒 39 token，这一结果来自伯克利与 MIT 联合开源的 FreeToken 框架。InfoQ 报道显示，该项目针对消费级硬件的量化与内核优化，让本地大模型推理首次在主流显卡上实现实用速度。 FreeToken 项目把注意力放在消费级 GPU 上。测试显示，在 RTX 4060 上运行 35B 参数模型时，生成速度达到每秒 39 Token。</description>
    </item>
  </channel>
</rss>
