首页 > 解决方案 > 在 python 中调用 CNBC 的后端 API

问题描述

作为这个问题的后续,我如何找到用于从 CNBC 新闻上的后端 API 检索数据的 XHR 请求,以便能够抓取这个CNBC 搜索查询

最终目标是拥有一个包含以下内容的文档:标题、日期、全文和网址。

我发现了这个: https ://api.sail-personalize.com/v1/personalize/initialize?pageviews=1&isMobile=0&query=coronavirus&qsearchterm=coronavirus

这告诉我我没有访问权限。有没有办法访问信息?

标签: pythonapiseleniumweb-scrapingselenium-chromedriver

解决方案


实际上,我之前对您的回答是针对您有关XHR请求的问题:

但是在这里我们使用screenshot

在此处输入图像描述

import requests

params = {
    "queryly_key": "31a35d40a9a64ab3",
    "query": "coronavirus",
    "endindex": "0",
    "batchsize": "100",
    "callback": "",
    "showfaceted": "true",
    "timezoneoffset": "-120",
    "facetedfields": "formats",
    "facetedkey": "formats|",
    "facetedvalue":
    "!Press Release|",
    "needtoptickers": "1",
    "additionalindexes": "4cd6f71fbf22424d,937d600b0d0d4e23,3bfbe40caee7443e,626fdfcd96444f28"
}

goal = ["cn:title", "_pubDate", "cn:liveURL", "description"]


def main(url):
    with requests.Session() as req:
        for page, item in enumerate(range(0, 1100, 100)):
            print(f"Extracting Page# {page +1}")
            params["endindex"] = item
            r = req.get(url, params=params).json()
            for loop in r['results']:
                print([loop[x] for x in goal])


main("https://api.queryly.com/cnbc/json.aspx")

Pandas DataFrame版本:

import requests
import pandas as pd

params = {
    "queryly_key": "31a35d40a9a64ab3",
    "query": "coronavirus",
    "endindex": "0",
    "batchsize": "100",
    "callback": "",
    "showfaceted": "true",
    "timezoneoffset": "-120",
    "facetedfields": "formats",
    "facetedkey": "formats|",
    "facetedvalue":
    "!Press Release|",
    "needtoptickers": "1",
    "additionalindexes": "4cd6f71fbf22424d,937d600b0d0d4e23,3bfbe40caee7443e,626fdfcd96444f28"
}

goal = ["cn:title", "_pubDate", "cn:liveURL", "description"]


def main(url):
    with requests.Session() as req:
        allin = []
        for page, item in enumerate(range(0, 1100, 100)):
            print(f"Extracting Page# {page +1}")
            params["endindex"] = item
            r = req.get(url, params=params).json()
            for loop in r['results']:
                allin.append([loop[x] for x in goal])
        new = pd.DataFrame(
            allin, columns=["Title", "Date", "Url", "Description"])
        new.to_csv("data.csv", index=False)


main("https://api.queryly.com/cnbc/json.aspx")

输出:在线查看

在此处输入图像描述


推荐阅读