python - dataframe combine_first does not work as well as fillna
问题描述
the first dataframe is:
data_date cookie_type dau next_dau dau_7 dau_15
0 20181006 avg(0-d) 2288 NaN NaN NaN
1 20181006 avg(e-f) 2284 NaN NaN NaN
2 20181007 avg(e-f) 2296 100 NaN NaN
the second dataframe is :
data_date cookie_type next_dau
0 20181006 avg(e-f) 908
1 20181006 avg(0-d) 904
how to update the first dataframe's next_dau from the second one i have tried combine_first and fillna, they seem not support multi-index:
cols = ['data_date', 'cookie_type']
if (frame1 is not None and not frame1.empty):
frame1.set_index(cols)
print(frame1)
print(next_day_dau)
frame1.combine_first(next_day_dau.set_index(cols))
frame1.combine_first(dau_7.set_index(cols))
frame1.combine_first(dau_15.set_index(cols))
then i updated to:
frame1.index = frame1.data_date.astype(str) + frame1.cookie_type
next_day_dau.index = next_day_dau.data_date.astype(str) + next_day_dau.cookie_type
dau_7.index = dau_7.data_date.astype(str) + dau_7.cookie_type
dau_15.index = dau_15.data_date.astype(str) + dau_15.cookie_type
"""frame1.loc[next_day_dau.index, "next_dau"] = next_day_dau.next_dau
frame1.loc[dau_7.index, "dau_7"] = dau_7.dau_7
frame1.loc[dau_15.index, "dau_15"] = dau_15.dau_15"""
frame1.combine_first(next_day_dau)
frame1.combine_first(dau_7)
frame1.combine_first(dau_15)
print(frame1)
print(next_day_dau)
loc raise a error because of next_day_dau dose not contain all the indexes in frame1 then i tried combine-first and fillna with inplace=True ,all dont work.
{'data_date': {'20181007avg(0-d)': 20181007, '20181007avg(e-f)': 20181007, '20181006avg(0-d)': 20181006, '20181006avg(e-f)': 20181006}, 'cookie_type': {'20181007avg(0-d)': 'avg(0-d)', '20181007avg(e-f)': 'avg(e-f)', '20181006avg(0-d)': 'avg(0-d)', '20181006avg(e-f)': 'avg(e-f)'}, 'dau': {'20181007avg(0-d)': 2288, '20181007avg(e-f)': 2284, '20181006avg(0-d)': 2288, '20181006avg(e-f)': 2284}, 'next_dau': {'20181007avg(0-d)': nan, '20181007avg(e-f)': nan, '20181006avg(0-d)': nan, '20181006avg(e-f)': nan}, 'dau_7': {'20181007avg(0-d)': nan, '20181007avg(e-f)': nan, '20181006avg(0-d)': nan, '20181006avg(e-f)': nan}, 'dau_15': {'20181007avg(0-d)': nan, '20181007avg(e-f)': nan, '20181006avg(0-d)': nan, '20181006avg(e-f)': nan}}
{'data_date': {0: '20181007', 1: '20181007'}, 'cookie_type': {0: 'avg(e-f)', 1: 'avg(0-d)'}, 'next_dau': {0: 2284, 1: 2288}}
解决方案
您可以使用 pandasmerge
来解决您的用例。更多文档在这里:
https ://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.merge.html
print(t1)
cookie_type data_date next_dau
0 avg(0-d) 20181006 1
1 avg(e-f) 20181006 2
2 avg(e-f) 20181007 NaN
print(t2)
cookie_type data_date next_dau
0 avg(e-f) 20181006 908
1 avg(0-d) 20181006 904
2 avg(e-f) 20181007 905
result = pd.merge(t1, t2, on=['data_date', 'cookie_type'])
cookie_type data_date next_dau_x next_dau_y
0 avg(0-d) 20181006 1 904
1 avg(e-f) 20181006 2 908
2 avg(e-f) 20181007 NaN 905
现在,要仅更新Not Nan值,您可以使用where
子句。
result['col'] = result['next_dau_x'].where(result['next_dau_x'].notnull(), result['next_dau_y'])
现在,删除不需要的列。
result = result.drop(['next_dau_x','next_dau_y'], axis=1)
cookie_type data_date col
0 avg(0-d) 20181006 1
1 avg(e-f) 20181006 2
2 avg(e-f) 20181007 905
推荐阅读
- perl - Perl 错误 - 全局符号需要明确的包名
- ios - 搜索并获取 Realm 数据。无法解析格式字符串“”
- javascript - AngularJS +下拉菜单未分配其默认值“选择...”
- algorithm - 从给定数字到 1 所需的最小跳跃次数
- python - 坚持编写一个以字典形式返回大量名词的函数(python)
- java - java优先级队列没有根据hashmap的频率得到正确的顺序
- python - 有什么方法可以检查输入字段中的文本是使用 .set() 函数加载的还是由用户输入的?
- go - 依赖注入失去了结构参数的类型安全性
- python-3.x - 可以通过 Tensorflows 下载在 MNIST 中随机更改标签吗?
- python - 在 ipdb 中使用列表推导时未定义“x”