python 爬虫实例（四）

环境：

OS：Window10

python：3.7

爬取链家地产上面的数据，两个画面上的数据的爬取

效果，下面的两个网页中的数据取出来

代码

import datetime

import threading

import requests

from bs4 import BeautifulSoup

class LianjiaHouseInfo:

    '''

        初期化变量的值

    '''

    def __init__(self):

        # 定义自己要爬取的URL

        self.url = "https://dl.lianjia.com/ershoufang/pg{0}"

        self.path = r"C:\pythonProject\Lianjia_House"

        self.headers = {

            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/71.0.3578.98 Safari/537.36"}

    '''

        访问URL

    '''

    def request(self, param):

        # 如果不加的话可能会出现403的错误，所以尽量的都加上header，模仿网页来访问

        req = requests.get(param, headers=self.headers)

        # req.raise_for_status()

        # req.encoding = req.apparent_encoding

        return req.text

    '''

        page設定

    '''

    def all_pages(self, pageCn):

        dataListA = []

        for i in range(1, pageCn+1):

            if pageCn == 1:

                dataListA = dataListA + self.getData(self.url[0:self.url.find("pg")])

            else:

                url = self.url.format(i)

                dataListA = dataListA + self.getData(url)

        # self.dataOrganize(dataListA)

    '''

       数据取得

    '''

    def getData(self, url):

        dataList = []

        thread_lock.acquire()

        req = self.request(url)

        # driver = webdriver.Chrome()

        # driver.get(self.url)

        # iframe_html = driver.page_source

        # driver.close()

        # print(iframe_html)

        soup = BeautifulSoup(req, 'lxml')

        countHouse = soup.find(class_="total fl").find("span")

        print("共找到 ", countHouse.string, " 套大连二手房")

        sell_all = soup.find(class_="sellListContent").find_all("li")

        for sell in sell_all:

            title = sell.find(class_="title")

            if title is not None:

                print("------------------------概要--------------------------------------------")

                title = title.find("a")

                print("title:", title.string)

                housInfo = sell.find(class_="houseInfo").get_text()

                print("houseInfo:", housInfo)

                positionInfo = sell.find(class_="positionInfo").get_text()

                print("positionInfo:", positionInfo)

                followInfo = sell.find(class_="followInfo").get_text()

                print("followInfo:", followInfo)

                print("------------------------詳細信息--------------------------------------------")

                url_detail = title["href"]

                req_detail = self.request(url_detail)

                soup_detail = BeautifulSoup(req_detail, "lxml")

                total = soup_detail.find(class_="total")

                unit = soup_detail.find(class_="unit").get_text()

                dataList.append(total.string+unit)

                print("总价:", total.string, unit)

                unitPriceValue = soup_detail.find(class_="unitPriceValue").get_text()

                dataList.append(unitPriceValue)

                print("单价:", unitPriceValue)

                room_mainInfo = soup_detail.find(class_="room").find(class_="mainInfo").get_text()

                dataList.append(room_mainInfo)

                print("户型:", room_mainInfo)

                type_mainInfo = soup_detail.find(class_="type").find(class_="mainInfo").get_text()

                dataList.append(type_mainInfo)

                print("朝向:", type_mainInfo)

                area_mainInfo = soup_detail.find(class_="area").find(class_="mainInfo").get_text()

                dataList.append(area_mainInfo)

                print("面积:", area_mainInfo)

            else:

                print("広告です")

        thread_lock.release()

        return dataList

    #

    # def dataOrganize(self, dataList):

    #

    #     data2 = pd.DataFrame(dataList)

    #     data2.to_csv(r'C:\Users\peiqiang\Desktop\lagoujob.csv', header=False, index=False, mode='a+')

    #     data3 = pd.read_csv(r'C:\Users\peiqiang\Desktop\lagoujob.csv', encoding='gbk')

thread_lock = threading.BoundedSemaphore(value=100)

house_Info = LianjiaHouseInfo()

startTime = datetime.datetime.now()

house_Info.all_pages(1)

endTime = datetime.datetime.now()

print("実行時間：", (endTime - startTime).seconds)

　　运行之后的效果

python 爬虫实例（四）的更多相关文章

Python爬虫实例：爬取B站《工作细胞》短评——异步加载信息的爬取
很多网页的信息都是通过异步加载的,本文就举例讨论下此类网页的抓取. <工作细胞>最近比较火,bilibili 上目前的短评已经有17000多条. 先看分析下页面右边 li 标签中的就是短 ...
Python爬虫实例：爬取猫眼电影——破解字体反爬
字体反爬字体反爬也就是自定义字体反爬,通过调用自定义的字体文件来渲染网页中的文字,而网页中的文字不再是文字,而是相应的字体编码,通过复制或者简单的采集是无法采集到编码后的文字内容的. 现在貌似不少网 ...
Python爬虫实例：爬取豆瓣Top250
入门第一个爬虫一般都是爬这个,实在是太简单.用了 requests 和 bs4 库. 1.检查网页元素,提取所需要的信息并保存.这个用 bs4 就可以,前面的文章中已经有详细的用法阐述. 2.找到下一 ...
Python爬虫实战四之抓取淘宝MM照片
原文:Python爬虫实战四之抓取淘宝MM照片其实还有好多,大家可以看 Python爬虫学习系列教程福利啊福利,本次为大家带来的项目是抓取淘宝MM照片并保存起来,大家有没有很激动呢? 本篇目标 1. ...
Python爬虫进阶四之PySpider的用法
审时度势 PySpider 是一个我个人认为非常方便并且功能强大的爬虫框架,支持多线程爬取.JS动态解析,提供了可操作界面.出错重试.定时爬取等等的功能,使用非常人性化. 本篇内容通过跟我做一个好玩的 ...
Python爬虫入门四之Urllib库的高级用法
1.设置Headers 有些网站不会同意程序直接用上面的方式进行访问,如果识别有问题,那么站点根本不会响应,所以为了完全模拟浏览器的工作,我们需要设置一些Headers 的属性. 首先,打开我们的浏览 ...
转 Python爬虫入门四之Urllib库的高级用法
静觅 » Python爬虫入门四之Urllib库的高级用法 1.设置Headers 有些网站不会同意程序直接用上面的方式进行访问,如果识别有问题,那么站点根本不会响应,所以为了完全模拟浏览器的工作,我 ...
Python 爬虫实例
下面是我写的一个简单爬虫实例 1.定义函数读取html网页的源代码 2.从源代码通过正则表达式挑选出自己需要获取的内容 3.序列中的htm依次写到d盘 #!/usr/bin/python import ...
shell及Python爬虫实例展示
1.shell爬虫实例: [root@db01 ~]# vim pa.sh #!/bin/bash www_link=http://www.cnblogs.com/clsn/default.html? ...
python爬虫实例——爬取歌单
学习自<<从零开始学python网络爬虫>> 爬取酷狗歌单,保存入csv文件直接上源代码:(含注释) import requests #用于请求网页获取网页数据 from b ...

随机推荐

GDB的安装
1.下载GDB7.10.1安装包 #wget http://ftp.gnu.org/gnu/gdb/gdb-7.10.1.tar.gz或者可以远程看下有哪些版本 http://ftp.gnu.org/ ...
如何在Processing中用鼠标获取RGB颜色数值
要做一个抠图应用,所以随手做了个鼠标取色,代码如下: void mousePressed(){ int imgC = get(mouseX,mouseY); int R = (imgC >> ...
c++和cuda混合编程实现传统神经网络
直接放代码了... 实现的是x1+x2=y的预测,但梯度下降很慢...233333,gpu运行时间很快!! // // main.cpp // bp // // Created by jzc on 2 ...
小程序在选择某一个东西的时候,可以用if,else 来做
<view class='fake-select-item-text brand-selected' wx:if='{{selectedBrandName}}'> {{selectedBr ...
kubernetes（K8S）创建自签TLS证书
TLS证书用于进行通信使用,组件需要证书关系如下: 组件需要使用的证书 etcd ca.pem server.pem server-key.pem flannel ca.pem server.pem ...
mac环境使用python处理protobuf
安装 brew install protobuf 然后再安装protobuf需要的依赖 brew install autoconf automake libtool 验证是否安装成功 protoc – ...
Chapter Two
Web容器配置 ~Tomcat配置 server.port配置了Web容器的端口号 error.path配置了当项目出错时跳转去的页面 session.timeout配置了session失效的时间 c ...
Python之pygame学习绘制文字制作滚动文字
pygame绘制文字 ✕ 今天来学习绘制文本内容,毕竟游戏中还是需要文字对玩家提示一些有用的信息. 字体常用的不是很多,在pygame中大多用于提示文字,或者记录分数等事件. 字体绘制基本分为以下几个 ...
strace命令二
让我们看一台高负载服务器的 top 结果: top 技巧:运行 top 时,按「1」打开 CPU 列表,按「shift+p」以 CPU 排序. 在本例中大家很容易发现 CPU 主要是被若干个 PHP ...
常见的医学基因筛查检测 | genetic testing | 相癌症早筛 | 液体活检
NIPT, Non-invasive Prenatal Testing - 无创产前基因检测 (学术名词) NIFTY,胎儿染色体异常无创产前基因检测 (注册商标)华大的明显产品新生儿耳聋基因检测 ...

python 爬虫实例（四）

python 爬虫实例（四）的更多相关文章

随机推荐

热门专题