<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>PySpark on 知识铺的博客</title>
    <link>https://index.zshipu.com/geek001/tags/PySpark/</link>
    <description>Recent content in PySpark on 知识铺的博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Fri, 04 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://index.zshipu.com/geek001/tags/PySpark/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Pandas 与 Spark DataFrames：单机生产力还是分布式性能，该怎么选</title>
      <link>https://index.zshipu.com/geek001/post/20260904/Pandas-%E4%B8%8E-Spark-DataFrames%E5%8D%95%E6%9C%BA%E7%94%9F%E4%BA%A7%E5%8A%9B%E8%BF%98%E6%98%AF%E5%88%86%E5%B8%83%E5%BC%8F%E6%80%A7%E8%83%BD%E8%AF%A5%E6%80%8E%E4%B9%88%E9%80%89/</link>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://index.zshipu.com/geek001/post/20260904/Pandas-%E4%B8%8E-Spark-DataFrames%E5%8D%95%E6%9C%BA%E7%94%9F%E4%BA%A7%E5%8A%9B%E8%BF%98%E6%98%AF%E5%88%86%E5%B8%83%E5%BC%8F%E6%80%A7%E8%83%BD%E8%AF%A5%E6%80%8E%E4%B9%88%E9%80%89/</guid>
      <description>Pandas 和 Spark DataFrames 的 API 都支持 filter、group、join 和 transform，但前者只能在单机运行，后者原生分布式，这让选错工具直接付出性能或生产力代价。两者表面相似却面向完全不同的世界。 单机内存限制 vs 横向扩展能力直接决定数据规模上限 Pandas 在单台服务器上运行，所有数据必须装进内存。一旦数</description>
    </item>
  </channel>
</rss>
