使用requests、re、BeautifulSoup、线程池爬取携程酒店信息并保存到Excel中

import requests

import json

import re

import csv

import threadpool

import time, random

from bs4 import BeautifulSoup

from fake_useragent import UserAgent

def hotel(city_letter, city_num, city_name):

    with open('has_address.json', 'a+', encoding="utf-8") as f:

        f.write(str(city_num) + '\n')

    f.close()

    ss = 0

    with open('携程/%s.csv' % city_name, 'w+', encoding='utf-8-sig') as hotel_xie:

        k = csv.writer(hotel_xie, dialect='excel')

        k.writerow(['序号', '名称', '价格', '星级', '地址', '酒店介绍'])

        for i in range(1, 100):

            url = "http://hotels.ctrip.com/Domestic/Tool/AjaxHotelList.aspx"

            headers = {

                "Connection": "keep-alive",

                "origin": "http://hotels.ctrip.com",

                "Host": "hotels.ctrip.com",

                "referer": "http://hotels.ctrip.com/hotel/%s" % city_letter,

                "user-agent": UserAgent(verify_ssl=False).random,

                "Content-Type": "application/x-www-form-urlencoded",

            }

            data = {

                "StartTime": "2019-02-25",

                "DepTime": "2019-02-26",

                "RoomGuestCount": "1,1,0",

                "city": city_num,

                "page": i,

            }

            try:

                time.sleep(random.randint(1, 5))

                html = requests.post(url, headers=headers, data=data)

                regex = re.compile(r'\\(?![/u"])')

                fixed = regex.sub(r"\\\\", html.text)

                aa = json.loads(fixed)

            except Exception:

                pass

            for n in range(0, 25):

                try:

                    hotel_name = aa["hotelPositionJSON"][n]["name"]

                    hotel_id = aa["hotelPositionJSON"][n]["id"]

                    hotel_address = aa["hotelPositionJSON"][n]["address"]

                    price = eval(aa["HotelMaiDianData"]["value"]["htllist"])[n]["amount"]

                    star_class = aa["hotelPositionJSON"][n]["star"][-2:]

                    time.sleep(random.randint(1, 3))

                    hotel_intro = requests.get('http://hotels.ctrip.com/hotel/%s.html' % hotel_id)

                    res_req = BeautifulSoup(hotel_intro.text, "html5lib")

                    iss = re.sub('资质备案', '', re.sub('联系方式', '', res_req.find('div', id='htlDes').findAll('p')[0].get_text()))

                    ins = iss.replace('\n', '').replace(' ', '').replace('&nbsp;', '')

                    s = res_req.find('span', id='J_realContact')['data-real'].replace('\n', ',')

                    tel = s[s.rfind("电话"): s.rfind("<a") - 2]

                    duction = res_req.find('span', id='ctl00_MainContentPlaceHolder_hotelDetailInfo_lbDesc').get_text().replace('\n', ',')

                    introduction = str(ins) + str(tel) + str(duction)

                    ss += 1

                    k.writerow([ss, hotel_name,  price + "元起", star_class, hotel_address, introduction])

                except Exception:

                    continue

                time.sleep(random.randint(1, 4))

    hotel_xie.close()

if __name__ == '__main__':

    has_num = []

    will_req_list = []

    for line in open("address.json", encoding='utf-8'):

        single_list = line.replace("\n", "").split(',')

        for has in open("has_address.json", encoding='utf-8'):

            has_num.append(int(has.replace("\n", "")))

        if int(single_list[1]) in has_num:

            continue

        single_tuple = (single_list, None)

        will_req_list.append(single_tuple)

    pool = threadpool.ThreadPool(8)

    request_list = threadpool.makeRequests(hotel, will_req_list)

    [pool.putRequest(req) for req in request_list]

    pool.wait()

    # 爬取地址

    # h = {

    #         "Connection": "keep-alive",

    #         "origin": "http://hotels.ctrip.com",

    #         "Host": "hotels.ctrip.com",

    #         "referer": "http://hotels.ctrip.com/hotel/beijing1",

    #         "user-agent": "Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/71.0.3578.98 Safari/537.36",

    #         "Content-Type": "application/x-www-form-urlencoded",

    #     }

    # res = requests.get('http://hotels.ctrip.com/Domestic/Tool/AjaxGetCitySuggestion.aspx', headers=h)

    # a_list = re.findall('data:(.*?),group:', res.text)

    # with open('address.json', 'w+',  encoding="utf-8") as f:

    #     for address in a_list:

    #         i = 0

    #         al = address.split(',')

    #         for a in al:

    #             city_letter = ''.join(re.findall(r'[A-Za-z]', a))

    #             f.write(city_letter + ',')

    #             city_num = re.sub("\D", "", a)

    #             f.write(str(city_num))

    #             city_name = re.sub('[A-Za-z0-9\"\|]', "", a)

    #             f.write(',' + str(city_name))

    #             f.write('\n')

    #         i += 1

    # f.close()

使用requests、re、BeautifulSoup、线程池爬取携程酒店信息并保存到Excel中的更多相关文章

使用requests、BeautifulSoup、线程池爬取艺龙酒店信息并保存到Excel中
import requests import time, random, csv from fake_useragent import UserAgent from bs4 import Beauti ...
Python爬取猫眼电影100榜并保存到excel表格
首先我们前期要导入的第三方类库有; 通过猫眼电影100榜的源码可以看到很有规律如: 亦或者是: 根据规律我们可以得到非贪婪的正则表达式 """<div class ...
爬取拉勾网所有python职位并保存到excel表格对象方式
# 1.把之间案例,使用bs4,正则,xpath,进行数据提取. # 2.爬取拉钩网上的所有python职位. from urllib import request,parse import json ...
爬取淘宝商品数据并保存在excel中
1.re实现 import requests from requests.exceptions import RequestException import re,json import xlwt,x ...
基于requests模块的cookie,session和线程池爬取
目录基于requests模块的cookie,session和线程池爬取基于requests模块的cookie操作基于requests模块的代理操作基于multiprocessing.dummy ...
Python+Requests+异步线程池爬取视频到本地
1.本次项目为获取梨视频中的视频,再使用异步线程池下载视频到本地 2.获取视频时,其地址中的Url是会动态变化,不播放时src值为图片的地址,播放时src值为mp4格式 3.查看视频链接是否存在aja ...
python爬取数据保存到Excel中
# -*- conding:utf-8 -*- # 1.两页的内容 # 2.抓取每页title和URL # 3.根据title创建文件,发送URL请求,提取数据 import requests fro ...
使用pandas中的raad_html函数爬取TOP500超级计算机表格数据并保存到csv文件和mysql数据库中
参考链接:https://www.makcyun.top/web_scraping_withpython2.html #!/usr/bin/env python # -*- coding: utf-8 ...
「拉勾网」薪资调查的小爬虫，并将抓取结果保存到excel中
学习Python也有一段时间了,各种理论知识大体上也算略知一二了,今天就进入实战演练:通过Python来编写一个拉勾网薪资调查的小爬虫. 第一步:分析网站的请求过程我们在查看拉勾网上的招聘信息的时候 ...

随机推荐

Activiti服务任务（serviceTask）
Activiti服务任务(serviceTask) 作者:Jesai 都有一段沉默的时间,等待厚积薄发应用场景: 当客户有这么一个需求:下一个任务我需要自动执行一些操作,并且这个节点不需要任何的人工 ...
全网最详细！Centos7.X 搭建Grafana+Jmeter+Influxdb 性能实时监控平台
背景日常工作中,经常会用到Jmeter去压测,毕竟LR还要钱(@￥&*...),而最常用的接口压力测试,我们都是通过聚合报告去查看压测结果的,然鹅聚合报告的真的是丑到家了,作为程序猿这当然不 ...
python 验证客户端的合法性
目的:对连接服务器的客户端进行判断 # Server import socket import hmac import os secret_key = bytes('tom', encoding='u ...
1.Java和Python的选择
我认为高级语言分为Java/c系列和其他. Java:1995年,让程序员设计一些大型分布式复杂应用. Python:1991年,面向系统管理.科研教育.等非程序员群体用的多. C系列语言:奠定了现在 ...
去除空白字符串trim
let str = ' foo ' //去除开头空格 console.log(str.trimLeft()) console.log(str.trimStart()) //去除尾部空格 console ...
Collections中的常用方法
collections中的常用方法 public class CollectionsTest { public static void main(String[] args) { List list ...
「从0到1学习微服务SpringCloud 」05服务消费者Fegin
系列文章(更新ing): 「从0到1学习微服务SpringCloud 」01 一起来学呀! 「从0到1学习微服务SpringCloud 」02 Eureka服务注册与发现「从0到1学习微服务S ...
工具之sed
转自:http://www.cnblogs.com/dong008259/archive/2011/12/07/2279897.html sed是一个很好的文件处理工具,本身是一个管道命令,主要是以行 ...
python(从放弃到从头开始)
本节内容 Python介绍发展史 Python 2 or 3? Hello World程序变量用户输入 .pyc是个什么鬼? 数据类型初识数据运算表达式if ...else语句表达式for ...
Oracle数据库、实例、用户、表空间、表之间的关系
完整的Oracle数据库通常由两部分组成:Oracle数据库和数据库实例. 1) 数据库是一系列物理文件的集合(数据文件,控制文件,联机日志,参数文件等): 2) Oracle数据库实例则是一组Ora ...

使用requests、re、BeautifulSoup、线程池爬取携程酒店信息并保存到Excel中

使用requests、re、BeautifulSoup、线程池爬取携程酒店信息并保存到Excel中的更多相关文章

随机推荐

热门专题