r - R - 使用for循环或lapply对文件夹中的两个文件应用函数并将结果保存在一个数据帧中
问题描述
我在“数据”中有一个包含 20 个文件夹的数据集,它们的结构相同。文件夹级别的唯一区别是它们的名称(从“1”到“20”)。请看下面的图案。这些文件始终具有相同的文件名和相同的列结构。文件夹之间的文件中的列长度可能存在差异,但同一文件夹中的.csv
文件之间没有差异。.csv
数据框中没有缺失值。我想使用文件中的“均值”列。
数据结构
data
- 1 (folder)
- alpha (file)
- mean (column)
- .... (more columns)
- beta (file)
- mean (column)
- .... (more columns)
- ... (more files)
- 2 (folder)
- alpha (file)
- mean (column)
- .... (more columns)
- beta (file)
- mean (column)
- .... (more columns)
- ... (more files)
- ... (more folders with the same structure)
我想在一个文件夹中比较 alpha 的平均值和 beta 的平均值。然而,最后,我想要一个数据框,它是所有单个文件夹的所有结果的子集。所以我可以从这个数据框中创建分面箱线图和描述性统计数据。
我还是 R 新手,显然缺乏它的技能(也很抱歉复杂的代码和我的英语)。我可以为每个文件夹手动执行任务,但我不能将结果与 for 循环或 lapply 解决方案放在一起。
我发现许多线程需要合并数据帧,而无需事先从同一文件夹中的两个文件执行函数。我确实希望我制作了一个可行的最小示例,其中每个来自 2 个文件夹的 2 个数据帧。
library(plyr)
library(tidyverse)
alpha1 <- read_csv('data/1/alpha.csv')
beta1 <- read_csv('data/1/beta.csv')
alpha2 <- read_csv('data/2/alpha2.csv')
beta2 <- read_csv('data/2/beta2.csv')
文件夹 1
alpha1 <- structure(list(Name = c("A", "B", "C", "D", "E", "F", "G", "H",
"I", "J", "K"), mean = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11)), class = c("spec_tbl_df", "tbl_df", "tbl",
"data.frame"), row.names = c(NA, -11L), spec = structure(list(
cols = list(Name = structure(list(), class = c("collector_character",
"collector")), mean = structure(list(), class = c("collector_double",
"collector"))), default = structure(list(), class = c("collector_guess",
"collector")), skip = 1), class = "col_spec"))
beta1 <- structure(list(Name = c("A", "B", "C", "D", "E", "F", "G", "H",
"I", "J", "K"), mean = c(2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12)), class = c("spec_tbl_df", "tbl_df", "tbl",
"data.frame"), row.names = c(NA, -11L), spec = structure(list(
cols = list(Name = structure(list(), class = c("collector_character",
"collector")), mean = structure(list(), class = c("collector_double",
"collector"))), default = structure(list(), class = c("collector_guess",
"collector")), skip = 1), class = "col_spec"))
alpha_mean <- alpha1 %>% select(mean_alpha = mean)
alphabeta <- alpha_mean %>% add_column(mean_beta = beta1$mean)
alphabeta_table <- ddply(alphabeta, .(), transform, alphabeta = (mean_alpha/mean_beta))
alphabeta_table
.id mean_alpha mean_beta alphabeta
1 <NA> 1 2 0.5000000
2 <NA> 2 3 0.6666667
3 <NA> 3 4 0.7500000
4 <NA> 4 5 0.8000000
5 <NA> 5 6 0.8333333
6 <NA> 6 7 0.8571429
7 <NA> 7 8 0.8750000
8 <NA> 8 9 0.8888889
9 <NA> 9 10 0.9000000
10 <NA> 10 11 0.9090909
11 <NA> 11 12 0.9166667
文件夹 2
alpha2 <- structure(list(Name = c("A", "B", "C", "D", "E", "F", "G", "H",
"I", "J", "K", "L", "M"), mean = c(2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14)), class = c("spec_tbl_df",
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -13L), spec = structure(list(
cols = list(Name = structure(list(), class = c("collector_character",
"collector")), mean = structure(list(), class = c("collector_double",
"collector"))), default = structure(list(), class = c("collector_guess",
"collector")), skip = 1), class = "col_spec"))
beta2 <- structure(list(Name = c("A", "B", "C", "D", "E", "F", "G", "H",
"I", "J", "K", "L", "M"), mean = c(3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15)), class = c("spec_tbl_df",
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -13L), spec = structure(list(
cols = list(Name = structure(list(), class = c("collector_character",
"collector")), mean = structure(list(), class = c("collector_double",
"collector"))), default = structure(list(), class = c("collector_guess",
"collector")), skip = 1), class = "col_spec"))
alpha2_mean <- alpha2 %>% select(mean_alpha = mean)
alphabeta2 <- alpha2_mean %>% add_column(mean_beta = beta2$mean)
alphabeta2_table <- ddply(alphabeta2, .(), transform, alphabeta = (mean_alpha/ mean_beta))
alphabeta2_table
.id mean_alpha mean_beta alphabeta
1 <NA> 2 3 0.6666667
2 <NA> 3 4 0.7500000
3 <NA> 4 5 0.8000000
4 <NA> 5 6 0.8333333
5 <NA> 6 7 0.8571429
6 <NA> 7 8 0.8750000
7 <NA> 8 9 0.8888889
8 <NA> 9 10 0.9000000
9 <NA> 10 11 0.9090909
10 <NA> 11 12 0.9166667
11 <NA> 12 13 0.9230769
12 <NA> 13 14 0.9285714
13 <NA> 14 15 0.9333333
期望的输出
我想要的输出是:
.id mean_alpha mean_beta alphabeta
1 1 1 2 0.5000000
2 1 2 3 0.6666667
3 1 3 4 0.7500000
4 1 4 5 0.8000000
5 1 5 6 0.8333333
6 1 6 7 0.8571429
7 1 7 8 0.8750000
8 1 8 9 0.8888889
9 1 9 10 0.9000000
10 1 10 11 0.9090909
11 1 11 12 0.9166667
1 2 2 3 0.6666667
2 2 3 4 0.7500000
3 2 4 5 0.8000000
4 2 5 6 0.8333333
5 2 6 7 0.8571429
6 2 7 8 0.8750000
7 2 8 9 0.8888889
8 2 9 10 0.9000000
9 2 10 11 0.9090909
10 2 11 12 0.9166667
11 2 12 13 0.9230769
12 2 13 14 0.9285714
13 2 14 15 0.9333333
1 3 ... ... ...
2 3 ... ... ...
...
感谢您的任何帮助!
解决方案
试试这个解决方案:
使用 . 获取所有文件夹
list.dirs
。对于每个文件夹,读取“alpha”和“beta”文件并返回一个包含
alpha
,beta
和alphabeta
值的 3 列标题。将所有数据框与
id
列绑定,以了解每个值来自哪个文件夹。
all_folders <- list.dirs('Data/', recursive = FALSE, full.names = TRUE)
result <- purrr::map_df(all_folders, function(x) {
all_Files <- list.files(x, full.names = TRUE, pattern = 'alpha|beta')
df1 <- read.csv(all_Files[1])
df2 <- read.csv(all_Files[2])
tibble::tibble(alpha = df1$mean, beta = df2$mean, alphabeta = alpha/beta)
}, .id = "id")
推荐阅读
- c++ - 如何检查qt simple mediaplayer中是否已经存在文件
- django - ModuleNotFoundError:没有名为“weasyprint”的模块
- python - 如果成员没有定义怎么办?
- flutter - NextJs OAuth 标准流程
- json - React .map 不是函数和解决方案
- vue.js - 用于表格的Vue js简单组件不起作用
- javascript - 如何根据我运行的项目更改组件中的文本?
- java - 它在 import org.apache.catalina.authenticator.SpnegoAuthenticator.AuthenticateAction 中显示错误;
- visual-studio-code - 禁用“重构预览”确认选项卡
- html - 锚标签和 CSS Flexbox