<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>低显存部署 on GPT资讯  --  知识铺</title>
    <link>https://index.zshipu.com/gpt/tags/%E4%BD%8E%E6%98%BE%E5%AD%98%E9%83%A8%E7%BD%B2/</link>
    <description>Recent content in 低显存部署 on GPT资讯  --  知识铺</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Fri, 04 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/gpt/tags/%E4%BD%8E%E6%98%BE%E5%AD%98%E9%83%A8%E7%BD%B2/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>AirLLM 用逐层流式加载让 4GB 显存跑通 Llama 3 70B</title>
      <link>https://index.zshipu.com/gpt/post/20260904/AirLLM-%E7%94%A8%E9%80%90%E5%B1%82%E6%B5%81%E5%BC%8F%E5%8A%A0%E8%BD%BD%E8%AE%A9-4GB-%E6%98%BE%E5%AD%98%E8%B7%91%E9%80%9A-Llama-3-70B/</link>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/gpt/post/20260904/AirLLM-%E7%94%A8%E9%80%90%E5%B1%82%E6%B5%81%E5%BC%8F%E5%8A%A0%E8%BD%BD%E8%AE%A9-4GB-%E6%98%BE%E5%AD%98%E8%B7%91%E9%80%9A-Llama-3-70B/</guid>
      <description>AirLLM 让 4GB 显存单卡直接运行 Llama 3 70B，8GB 跑 405B，12GB 跑 DeepSeek-V3，且无需量化、蒸馏或剪枝。项目核心是逐层流式加载技术，把模型参数按需从 CPU 或磁盘调入 GPU，推理完立即释放，显存占用被压到极低。 AirLLM 的出现直接改变了很多人对大模型硬件需求的认知。此前，70B 参数量级的</description>
    </item>
  </channel>
</rss>
