pandas 数据处理 10 个高效技巧,让分析速度提升 5 倍
原创
已于 2026-08-24 23:12 发布
·
5 阅读
1. 用向量化替代 apply
Python
import pandas as pd
import numpy as np
df = pd.DataFrame({'value': np.random.randn(1_000_000)})
df['flag_apply'] = df['value'].apply(lambda x: 1 if x > 0 else 0)
df['flag_vec'] = np.where(df['value'] > 0, 1, 0)
2. 用 query 简化布尔索引
Python
result = df.query('value > 0 and category == "A"')
3. 用 category 类型省内存
Python
df['city'] = df['city'].astype('category')
print(df.memory_usage(deep=True))
4. 分块读取大文件
Python
chunks = []
for chunk in pd.read_csv('big.csv', chunksize=100_000):
chunk = chunk[chunk['amount'] > 0]
chunks.append(chunk.groupby('user_id')['amount'].sum())
result = pd.concat(chunks).groupby(level=0).sum()
5. 指定 dtype 减少内存
Python
df = pd.read_csv('data.csv', dtype={
'user_id': 'int32',
'amount': 'float32',
'status': 'category'
})
6. 用 map 做字典映射
Python
mapping = {1: '待付款', 2: '已付款', 3: '已发货'}
df['status_name'] = df['status'].map(mapping)
7. groupby + agg 一次聚合多列
Python
summary = df.groupby('category').agg(
total=('amount', 'sum'),
avg=('amount', 'mean'),
cnt=('order_id', 'count')
).reset_index()
8. 用 merge 时注意索引
Python
df1 = df1.set_index('user_id')
df2 = df2.set_index('user_id')
result = df1.join(df2, how='left')
9. 避免链式赋值
Python
df[df.age > 30]['score'] = 100
df.loc[df.age > 30, 'score'] = 100
10. 用 eval 加速复杂表达式
Python
df.eval('total = price * quantity * (1 - discount)', inplace=True)
性能实测对比
| 操作 | 普通写法 | 优化后 | 提升 |
| --- | --- | --- | --- |
| 条件标记 | 1.2s | 0.008s | 150x |
| 字符串列内存 | 64MB | 1MB | 64x |
| 大文件读取 | OOM | 3.2s | - |
| 复杂计算 | 0.9s | 0.4s | 2.2x |
发表于天涯网络_专业开发者社区 - 个人技术分享
评论(0)
还没有评论,来说两句吧