20.multi_协程方法抓取总阅读量

# 用asyncio和aiohttp抓取博客的总阅读量 (提示:先用接又找到每篇文章的链接)

# https://www.jianshu.com/u/130f76596b02

import re

import asyncio

import aiohttp

import requests

import ssl

from lxml import etree

from asyncio.queues import Queue

from aiosocksy import Socks5Auth

from aiosocksy.connector import ProxyConnector, ProxyClientRequest

class Common():

    task_queue = Queue()

    result_queue = Queue()

    result_queue_1 = []

async def session_get(session, url, socks):

    auth = Socks5Auth(login='...', password='...')

    headers = {'User-Agent': 'Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)'}

    timeout = aiohttp.ClientTimeout(total=20)

    response = await session.get(

        url,

        proxy=socks,

        proxy_auth=auth,

        timeout=timeout,

        headers=headers,

        ssl=ssl.SSLContext()

    )

    return await response.text(), response.status

async def download(url):

    connector = ProxyConnector()

    socks = None

    async with aiohttp.ClientSession(

            connector=connector,

            request_class=ProxyClientRequest

    ) as session:

        ret, status = await session_get(session, url, socks)

        if 'window.location.href' in ret and len(ret) < 1000:

            url = ret.split("window.location.href='")[1].split("'")[0]

            ret, status = await session_get(session, url, socks)

        return ret, status

async def parse_html(content):

    read_num_pattern = re.compile(r'"views_count":\d+')

    read_num = int(read_num_pattern.findall(content)[0].split(':')[-1])

    return read_num

def get_all_article_links():

    links_list = []

    for i in range(1, 21):

        url = 'https://www.jianshu.com/u/130f76596b02?order_by=shared_at&page={}'.format(

            i)

        header = {

            'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',

            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 '

            '(KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36'}

        response = requests.get(url,

                                headers=header,

                                timeout=5

                                )

        tree = etree.HTML(response.text)

        article_links = tree.xpath(

            '//div[@class="content"]/a[@class="title"]/@href')

        for item in article_links:

            article_link = 'https://www.jianshu.com' + item

            links_list.append(article_link)

            print(article_link)

    return links_list

async def down_and_parse_task(queue):

    while True:

        try:

            url = queue.get_nowait()

        except BaseException:

            return

        error = None

        for retry_cnt in range(3):

            try:

                html, status = await download(url)

                if status != 200:

                    html, status = await download(url)

                read_num = await parse_html(html)

                print(read_num)

                # await Common.result_queue.put(read_num)

                Common.result_queue_1.append(read_num)

                break

            except Exception as e:

                error = e

                await asyncio.sleep(0.2)

                continue

        else:

            raise error

async def count_sum():

    while True:

        try:

            print(Common.result_queue_1)

            print('总阅读量 = ', sum(Common.result_queue_1))

            await asyncio.sleep(3)

        except BaseException:

            pass

async def main():

    all_links = get_all_article_links()

    for item in set(all_links):

        await Common.task_queue.put(item)

    for _ in range(10):

        loop.create_task(down_and_parse_task(Common.task_queue))

        loop.create_task(count_sum())

if __name__ == '__main__':

    loop = asyncio.get_event_loop()

    loop.create_task(main())

    loop.run_forever()

20.multi_协程方法抓取总阅读量的更多相关文章

代理池抓取基础版-（python协程）--抓取网站（西刺-后期会持续更新）
# coding = utf- __autor__ = 'litao' import urllib.request import urllib.request import urllib.error ...
成功抓取csdn阅读量过万博文
http://images.cnblogs.com/cnblogs_com/elesos/1120632/o_111.png var commentscount = 1; 嵌套的评论算一条,这个可能有 ...
比物理线程都好用的C++20的协程，你会用吗？
摘要:事件驱动(event driven)是一种常见的代码模型,其通常会有一个主循环(mainloop)不断的从队列中接收事件,然后分发给相应的函数/模块处理.常见使用事件驱动模型的软件包括图形用户界 ...
开启gzip压缩/cdn是否会影响抓取和收录量
http://www.wocaoseo.com/thread-291-1-1.html 服务器开启gzip压缩是否会影响蜘蛛抓取和收录量?站点开了CDN,对百度SEO影响有多大?我发现我们站自从开了C ...
(20)gevent协程
协程: 也叫纤程,协程是线程的一种实现,指的是一条线程能够在多任务之间来回切换的一种实现,对于CPU.操作系统来说,协程并不存在任务之间的切换会花费时间.目前电脑配置一般线程开到200会阻塞卡顿 ...
scrapy实战4 GET方法抓取ajax动态页面(以糗事百科APP为例子)：
一般来说爬虫类框架抓取Ajax动态页面都是通过一些第三方的webkit库去手动执行html页面中的js代码, 最后将生产的html代码交给spider分析.本篇文章则是通过利用fiddler抓包获取j ...
scrapy实战5 POST方法抓取ajax动态页面(以慕课网APP为例子)：
在手机端打开慕课网,fiddler查看如图注意圈起来的位置经过分析只有画线的page在变化上代码: items.py import scrapy class ImoocItem(scrapy.It ...
ADB logcat 过滤方法(抓取日志)
1. Log信息级别 Log.v- VERBOSE : 黑色 Log.d- DEBUG : 蓝色 Log.i- INFO : 绿色 Log.w- WARN : 橙色 Log.e- ERROR ...
python3用BeautifulSoup用字典的方法抓取a标签内的数据
# -*- coding:utf-8 -*- #python 2.7 #XiaoDeng #http://tieba.baidu.com/p/2460150866 #标签操作 from bs4 imp ...

随机推荐

【转载】查看Linux进程CPU过高具体的线程堆栈(不中断程序)
具体的命令经常忘记,毕竟用的不是很多.为了避免去找备份一下 1.TOP命令,找到占用CPU最高的进程 $ top top - 20:11:45 up 850 days, 1:18, 3 users ...
iOS开发UIResponder简介API
#import <Foundation/Foundation.h> #import <UIKit/UIKitDefines.h> #import <UIKit/UIEve ...
【JS】 +function(){} 作用
原文地址:https://www.jianshu.com/p/a2666014a280 瞎扯在JS中,经常会遇到下面这种代码, 到底在 function 前面加一个一元操作符, 有什么作用呢? ...
C89,C99: C数组&结构体&联合体快速初始化
1. 背景 C89标准规定初始化语句的元素以固定顺序出现,该顺序即待初始化数组或结构体元素的定义顺序. C99标准新增指定初始化(Designated Initializer),即可按照任意顺序对数组 ...
剑指offer——26反转链表
题目描述输入一个链表,反转链表后,输出新链表的表头. 题解: 每次只反转一个节点,先记住cur->next, 然后pre->cur,即可; class Solution { pu ...
Python print命令/ 解压序列
Python 命令参数 print 命令 : #默认的print是有个空格,和换行的 # print(sep= ' ') # print(end = '/n') a = 'sunjinchao' ...
金三银四铜五铁六，Offer收到手软！
作者:鲁班大师来源:cnblogs.com/zhuoqingsen/p/interview.html 文中的鲁班简称LB 据说,金三银四,截止今天为止面试黄金时间已经过去十之八九,而LB恰逢是这批面 ...
idea引入项目下所有文件（ps：包括静态文件夹）
打开项目的目录结构点击finish 最后删除目录下多余的src就可以了
MVC 传递数据从前台到后台，包括单个对象，多个对象，集合
MVC 传递数据从前台到后台,包括单个对象,多个对象,集合 1.基本数据类型我们常见有传递 int, string, bool, double, decimal 等类型. 需要注意的是前台传递的参 ...
Could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR tensorflow-1.13.1和1.14windows版本目前不支持CUDA10.0
报错出现 Could not create cudnn handle: CUDNN_STATUS_INTERNAL_ERROR tensorflow-1.13.1和1.14windows版本目前不支持 ...

20.multi_协程方法抓取总阅读量

20.multi_协程方法抓取总阅读量的更多相关文章

随机推荐

热门专题