Scrapy实战：爬取http://quotes.toscrape.com网站数据

需要学习的地方：

1.Scrapy框架流程梳理，各文件的用途等

2.在Scrapy框架中使用MongoDB数据库存储数据

3.提取下一页链接，回调自身函数再次获取数据

重点：从当前页获取下一页的链接，传给函数自身继续发起请求

next = response.css('.pager .next a::attr(href)').extract_first() # 获取下一页的相对链接
url = response.urljoin(next) # 生成完整的下一页链接
yield scrapy.Request(url=url, callback=self.parse) # 把下一页的链接回调给自身再次请求

站点:http://quotes.toscrape.com

该站点网页结构比较简单，需要的数据都在div标签中

操作步骤：

1.创建项目

# scrapy startproject quotetutorial

此时目录结构如下：

2.生成爬虫文件

# cd quotetutorial
# scrapy genspider quotes quotes.toscrape.com # 若是有多个爬虫多次操作该命令即可

3.编辑items.py文件，获取需要输出的数据

import scrapy

class QuoteItem(scrapy.Item):

    # define the fields for your item here like:

    # name = scrapy.Field()

    text = scrapy.Field()

    author = scrapy.Field()

    tags = scrapy.Field()

4.编辑quotes.py文件，爬取网站数据

# -*- coding: utf-8 -*-

import scrapy

from quotetutorial.items import QuoteItem

class QuotesSpider(scrapy.Spider):

    name = 'quotes'

    allowed_domains = ['quotes.toscrape.com']

    start_urls = ['http://quotes.toscrape.com/']

    def parse(self, response):

        # print(response.status) # 200

        quotes = response.css('.quote')

        for quote in quotes:

            item = QuoteItem()

            text = quote.css('.text::text').extract_first()

            author = quote.css('.author::text').extract_first()

            tags = quote.css('.tags .tag::text').extract()

            item['text'] = text

            item['author'] = author

            item['tags'] = tags

            yield item

        next = response.css('.pager .next a::attr(href)').extract_first()  # 获取下一页的相对链接

        url = response.urljoin(next)  # 生成完整的下一页链接

        yield scrapy.Request(url=url, callback=self.parse)  # 把下一页的链接回调给自身再次请求

5.编写pipelines.py文件，进一步处理item数据，保存到mongodb数据库

# -*- coding: utf-8 -*-

# Define your item pipelines here

#

# Don't forget to add your pipeline to the ITEM_PIPELINES setting

# See: https://doc.scrapy.org/en/latest/topics/item-pipeline.html

# 使用的话需要在settings文件中设置

import pymongo as pymongo

from scrapy.exceptions import DropItem

class TextPipeline(object):

    """对输出的item进行进一步的处理"""

    def __init__(self):

        self.limit = 50

    def process_item(self, item, spider):

        if item['text']:

            if len(item['text']) > self.limit:

                item['text'] = item['text'][0:self.limit].rstrip() + '......'

            return item

        else:

            return DropItem('Missing Text!')

class MongoPipeline(object):

    """把输出的item保存到MongoDB数据库"""

    def __init__(self, mongo_url, mongo_db):

        self.mongo_uri = mongo_url

        self.mongo_db = mongo_db

    @classmethod

    def from_crawler(cls, crawler):

        """从settings文件获取配置信息"""

        return cls(

            mongo_url=crawler.settings.get('MONGO_URI'),

            mongo_db=crawler.settings.get('MONGO_DB')

        )

    def open_spider(self, spider):

        """初始化mongodb"""

        self.client = pymongo.MongoClient(self.mongo_uri)

        self.db = self.client[self.mongo_db]  # 为啥用[],而不是()

    def process_item(self, item, spider):

        name = item.__class__.__name__  # 获取item的名称用作表名,也就是QuoteItem

        self.db[name].insert(dict(item))  # 为啥要用dict(item)

        return item

    def close_spider(self, spider):

        self.client.close()

6.编辑配置文件，增加mongodb数据库参数，以及使用的pipeline管道参数

ITEM_PIPELINES = {

   # 'quotetutorial.pipelines.TextPipeline': 300,

   'quotetutorial.pipelines.MongoPipeline': 400,

}

MONGO_URI = 'localhost'

MONGO_DB = 'quotestutorial'

7.执行程序

# scrapy crawl quotes

8.保存到文件

# scrapy crawl quotes -o quotes.json # 保存成json文件
# scrapy crawl quotes -o quotes.csv # 保存成csv文件
# scrapy crawl quotes -o quotes.xml # 保存成xml文件
# scrapy crawl quotes -o quotes.jl # 保存成jl文件
# scrapy crawl quotes -o quotes.pickle # 保存成pickle文件
# scrapy crawl quotes -o quotes.marshal # 保存成marshal文件
# scrapy crawl quotes -o ftp://user:password@ftp.example.com/path/quotes.csv # 生成csv文件保存到远程FTP上

效果：

源码下载地址：https://files.cnblogs.com/files/sanduzxcvbnm/quotetutorial.7z

Scrapy实战：爬取http://quotes.toscrape.com网站数据的更多相关文章

简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息系统环境:Fedora22(昨天已安装scrapy环境) 爬取的开始URL:ht ...
教程+资源,python scrapy实战爬取知乎最性感妹子的爆照合集(12G)!
一.出发点: 之前在知乎看到一位大牛(二胖)写的一篇文章:python爬取知乎最受欢迎的妹子(大概题目是这个,具体记不清了),但是这位二胖哥没有给出源码,而我也没用过python,正好顺便学一学,所以 ...
scrapy实战--爬取最新美剧
现在写一个利用scrapy爬虫框架爬取最新美剧的项目. 准备工作: 目标地址:http://www.meijutt.com/new100.html 爬取项目:美剧名称.状态.电视台.更新时间 1.创建 ...
<scrapy爬虫>爬取quotes.toscrape.com
1.创建scrapy项目 dos窗口输入: scrapy startproject quote cd quote 2.编写item.py文件(相当于编写模板,需要爬取的数据在这里定义) import ...
Scrapy Learning笔记（四）- Scrapy双向爬取
摘要:介绍了使用Scrapy进行双向爬取(对付分类信息网站)的方法. 所谓的双向爬取是指以下这种情况,我要对某个生活分类信息的网站进行数据爬取,譬如要爬取租房信息栏目,我在该栏目的索引页看到如下页面, ...
第三百三十节，web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解
第三百三十节,web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解封装模块 #!/usr/bin/env python # -*- coding: utf- ...
使用scrapy框架爬取自己的博文（2）
之前写了一篇用scrapy框架爬取自己博文的博客,后来发现对于中文的处理一直有问题- - 显示的时候 [u'python\u4e0b\u722c\u67d0\u4e2a\u7f51\u9875\u76 ...
如何提高scrapy的爬取效率
提高scrapy的爬取效率增加并发: 默认scrapy开启的并发线程为32个,可以适当进行增加.在settings配置文件中修改CONCURRENT_REQUESTS = 100值为100,并发设置 ...
九 web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解
封装模块 #!/usr/bin/env python # -*- coding: utf-8 -*- import urllib from urllib import request import j ...

随机推荐

Redux 源码解读--createStore,js
一.依赖:$$observable.ActionTypes.isPlainObject 二.接下来看到直接 export default 一个 createStore 函数,下面根据代码以及注释来分析 ...
数据库html 数据的分句
Python 中文分句 - CSDN博客 https://blog.csdn.net/laoyaotask/article/details/9260263 # 设置分句的标志符号:可以根据实际需要进行 ...
idea2016的使用心得 --- 太棒了
今天打开myeclipse感觉里面全是project,也懒着换地方了,因为这些代码还要时常看,索性安装了idea试试水,感觉还不错,用起来并不比myeclipse差,跟webstorm差不多,他俩就是 ...
【Apio2009】Bzoj1179 Atm
目录 List Description Input Output Sample Input Sample Output HINT Solution Code Dfs 记忆化搜索 Position: h ...
神经网络结构设计指导原则——输入层：神经元个数=feature维度输出层：神经元个数=分类类别数，默认只用一个隐层如果用多个隐层，则每个隐层的神经元数目都一样
神经网络结构设计指导原则原文 http://blog.csdn.net/ybdesire/article/details/52821185 下面这个神经网络结构设计指导原则是Andrew N ...
在IIS上搭建WebSocket服务器（二）
服务器端代码编写 1.新建一个ASP.net Web MVC5项目 2.新建一个“一般处理程序” 3.Handler1.ashx代码如下: using System; using System.Col ...
Codeforces--597A--Divisibility（数学）
DivisibilityCrawling in process... Crawling failed Time Limit:1000MS Memory Limit:262144KB ...
JAVA Swing 组件演示***
下面是Swing组件的演示: package a_swing; import java.awt.BorderLayout; import java.awt.Color; import java.awt ...
AcWing算法基础1.5
前缀和与差分两个内容都比较少,就放一起写了设数组 a 的前 n 项为a1 , a2 , a3 ... an 前缀和数组就是每一项是a数组的前i项和,比如前缀和数组res,res[ 1 ] = a[ ...
vue-resource 拦截器的使用
园友参考 https://www.cnblogs.com/lhl66/p/8022823.html vue-resource 拦截器使用在vue项目使用vue-resource的过程中,临时增加了一 ...

Scrapy实战：爬取http://quotes.toscrape.com网站数据

Scrapy实战：爬取http://quotes.toscrape.com网站数据的更多相关文章

随机推荐

热门专题