r - 根据 R 中其他列中的值随机更新列值
问题描述
我想添加一个新列SubCategory
,其中的值根据列的值随机填充Category
。这是详细信息:
Sub_Hair = c("Shampoo", "Conditioner", "Gel", "HairOil", "Dye")
Sub_Beauty = c("Face", "Eye", "Lips")
Sub_Nail= c("NailPolish", "NailPolishRemover", "NailArtKit", "ManiPadiKit")
Sub_Others = c("Electric", "NonElectric")
> product_data_1[1:10, c("Pcode", "Category", "MRP")]
Pcode Category MRP
1 16156L Beauty $8.88
2 16162M Others $21.27
3 16168M Others $2.98
4 16169E Nail $26.64
5 16207A Hair $6.38
6 17012B Beauty $33.03
7 17012C Beauty $20.58
8 17012F Beauty $36.29
9 17091A Nail $20.55
10 17107D Nail $28.20
我正在尝试下面的代码。但是,行正在更新,每个类别只有一个子类别。例如,所有具有“美容”类别的行,子类别都是“眼睛”,而不是从“面部、眼睛和嘴唇”中随机选择的值。这是代码和输出:
product_data_1 = within(product_data_1, SubCategory[Category == "Beauty"] <- sample(Sub_Beauty, 1))
product_data_1 = within(product_data_1, SubCategory[Category == "Hair"] <- sample(Sub_Hair, 1))
product_data_1 = within(product_data_1, SubCategory[Category == "Nail"] <- sample(Sub_Nail, 1))
product_data_1 = within(product_data_1, SubCategory[Category == "Others"] <- sample(Sub_Others, 1))
> product_data_1[1:10, c("Pcode", "Category", "MRP", "SubCategory")]
Pcode Category MRP SubCategory
1 16156L Beauty $8.88 Eye
2 16162M Others $21.27 Electric
3 16168M Others $2.98 Electric
4 16169E Nail $26.64 NailPolish
5 16207A Hair $6.38 Gel
6 17012B Beauty $33.03 Eye
7 17012C Beauty $20.58 Eye
8 17012F Beauty $36.29 Eye
9 17091A Nail $20.55 NailPolish
10 17107D Nail $28.20 NailPolish
解决方案
这是一个基本的 R 解决方案。它使用Hadley Wickham在这篇JSS 文章中解释的拆分/应用/组合策略。
我会将Sub_*
向量放在一个列表中,Sub_list
. 请注意,split
将对结果进行排序,Category
因此列表Sub_list
还必须按顺序排列向量。
Sub_list <- list(Sub_Beauty, Sub_Hair, Sub_Nail, Sub_Others)
sp <- split(product_data_1, product_data_1$Category)
set.seed(1234)
sp <- lapply(seq_along(sp), function(i){
sp[[i]]$SubCategory <- sample(Sub_list[[i]], nrow(sp[[i]]), replace = TRUE)
sp[[i]]
})
result <- do.call(rbind, sp)
result <- result[order(as.integer(row.names(result))), ]
result
# Pcode Category MRP SubCategory
#1 16156L Beauty $8.88 Eye
#2 16162M Others $21.27 NonElectric
#3 16168M Others $2.98 NonElectric
#4 16169E Nail $26.64 NailPolish
#5 16207A Hair $6.38 Shampoo
#6 17012B Beauty $33.03 Eye
#7 17012C Beauty $20.58 Face
#8 17012F Beauty $36.29 Lips
#9 17091A Nail $20.55 NailPolishRemover
#10 17107D Nail $28.20 ManiPadiKit
最后清理。
rm(Sub_list)
数据
product_data_1 <- read.table(text = "
Pcode Category MRP
1 16156L Beauty $8.88
2 16162M Others $21.27
3 16168M Others $2.98
4 16169E Nail $26.64
5 16207A Hair $6.38
6 17012B Beauty $33.03
7 17012C Beauty $20.58
8 17012F Beauty $36.29
9 17091A Nail $20.55
10 17107D Nail $28.20
", header = TRUE)
推荐阅读
- r - 在 Netlify 上部署后,Blogdown 网站未显示我的图片
- php - heroku 运行 php artisan 迁移
- tortoisesvn - 结帐时检查总和错误
- javascript - 在 ComponentWillMount 函数 / setState 不起作用之前反应渲染组件
- boolean-logic - 简化xnor的正确方法是什么
- c++11 - 获取最接近双精度的 std::vector 中的值的项目
- elasticsearch - 删除索引后是否可以从 Elasticsearch 恢复数据?
- reactjs - 使用 Lab Material UI Pickers 时出现导入错误
- java - 退出框架并创建新框架?
- kubernetes - k8s 权限被拒绝问题