r - 从 xml2 和 rvest 子集 data.frame 出错
问题描述
我正在从维基百科抓取数据并将其放入 data.frame 中。但是,我不能对生成的 data.frame 进行子集化。
如果我使用 dput 将数据重新加载到另一个变量中,那么子集就可以正常工作。我不清楚我是否做错了什么,或者 R 或我正在使用的某个包中是否存在错误。这是一个可重现的示例。
第 1 步:将数据加载到reps
library(rvest)
library(xml2)
url = "https://en.m.wikipedia.org/wiki/List_of_current_members_of_the_United_States_House_of_Representatives"
file = xml2::read_html(url)
tables = rvest::html_nodes(file, "table")
reps = rvest::html_table(tables[6])
reps = as.data.frame(reps)[1,1:3]
reps$District
# [1] "Alabama 1"
# I expected this line to return TRUE
reps$District == "Alabama 1"
# [1] FALSE
# Because the above line returns FALSE, this code returns an empty data.frame
reps[reps$District=="Alabama 1",]
# [1] District Member Party
# <0 rows> (or 0-length row.names)
特别奇怪的是,如果我使用 dput 写入和重新加载数据,子集可以正常工作:
dput(reps)
# structure(list(District = "Alabama 1", Member = "Bradley Byrne",
# Party = NA), row.names = 1L, class = "data.frame")
x=structure(list(District = "Alabama 1", Member = "Bradley Byrne",
Party = NA), row.names = 1L, class = "data.frame")
# now it's TRUE!
x$District=="Alabama 1"
# [1] TRUE
# and so the subset works
x[x$District == "Alabama 1", ]
# District Member Party
# 1 Alabama 1 Bradley Byrne NA
我相信我正在使用最新版本的 R 和所有软件包:
> sessionInfo()
R version 4.0.2 (2020-06-22)
Platform: x86_64-apple-darwin17.0 (64-bit)
Running under: macOS Catalina 10.15.5
Matrix products: default
BLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib
LAPACK: /Library/Frameworks/R.framework/Versions/4.0/Resources/lib/libRlapack.dylib
locale:
[1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
attached base packages:
[1] stats graphics grDevices utils datasets methods base
loaded via a namespace (and not attached):
[1] httr_1.4.2 compiler_4.0.2 selectr_0.4-2 magrittr_1.5 R6_2.4.1 tools_4.0.2
[7] curl_4.3 xml2_1.3.2 stringi_1.4.6 stringr_1.4.0 rvest_0.3.6
解决方案
\u00A0
将匹配
,因此您可以将所有这些替换为gsub
.
reps$District <- gsub("\u00A0", " ", reps$District, fixed = TRUE)
完整代码:
library(rvest)
library(xml2)
url = "https://en.m.wikipedia.org/wiki/List_of_current_members_of_the_United_States_House_of_Representatives"
file = xml2::read_html(url)
tables = rvest::html_nodes(file, "table")
reps = rvest::html_table(tables[6])
reps = as.data.frame(reps)[1,1:3]
reps$District <- gsub("\u00A0", " ", reps$District, fixed = TRUE)
reps$District == "Alabama 1"
# [1] TRUE