Urllib--爬虫

1.简单爬虫

from urllib import request

def f(url):

    print('GET: %s' % url)

    resp = request.urlopen(url) #赋给一个实例,请求

    data = resp.read() #把结果读出来

    f=open('url.html','wb')

    f.write(data)

    f.close()

    print('%d bytes received from %s.' % (len(data), url))

f('http://www.cnblogs.com/alex3714/articles/5248247.html')

运行结果：

C:\abccdxddd\Oldboy\python-3.5.2-embed-amd64\python.exe C:/abccdxddd/Oldboy/Py_Exercise/Day10/爬虫.py

GET: http://www.cnblogs.com/alex3714/articles/5248247.html

91829 bytes received from http://www.cnblogs.com/alex3714/articles/5248247.html.

Process finished with exit code 0

2.爬多个网页

from urllib import request

import gevent

def f(url):

    print('GET: %s' % url)

    resp = request.urlopen(url) #赋给一个实例,请求

    data = resp.read() #把结果读出来

    print('%d bytes received from %s.' % (len(data), url))

#启动3个协程并且传参数

gevent.joinall([

        gevent.spawn(f, 'https://www.python.org/'),

        gevent.spawn(f, 'https://www.yahoo.com/'),

        gevent.spawn(f, 'https://github.com/'),

])

运行结果：

GET: https://www.python.org/

48751 bytes received from https://www.python.org/.

GET: https://www.yahoo.com/

479631 bytes received from https://www.yahoo.com/.

GET: https://github.com/

55394 bytes received from https://github.com/.

Process finished with exit code 0

3.测试运行时间：

from urllib import request

import gevent

import time

def f(url):

    print('GET: %s' % url)

    resp = request.urlopen(url) #赋给一个实例,请求

    data = resp.read() #把结果读出来

    print('%d bytes received from %s.' % (len(data), url))

start_time=time.time()

#启动3个协程并且传参数

gevent.joinall([

        gevent.spawn(f, 'https://www.python.org/'),

        gevent.spawn(f, 'https://www.yahoo.com/'),

        gevent.spawn(f, 'https://github.com/'),

])

print('cost is %s:'%(time.time()-start_time))

运行结果：通过时间看到也是串行运行的。gevent默认检测不到 urllib 进行的是否是io操作。

C:\abccdxddd\Oldboy\python-3.5.2-embed-amd64\python.exe C:/abccdxddd/Oldboy/Py_Exercise/Day10/爬虫.py

GET: https://www.python.org/

48751 bytes received from https://www.python.org/.

GET: https://www.yahoo.com/

488624 bytes received from https://www.yahoo.com/.

GET: https://github.com/

55394 bytes received from https://github.com/.

cost is 4.5304529666900635:

Process finished with exit code 0

4.同步与异步的时间比较：

from urllib import request

import gevent

import time

#from gevent import monkey

#monkey.patch_all() #把当前程序的所有io操作给我单独地做上标记

def f(url):

    print('GET: %s' % url)

    resp = request.urlopen(url) #赋给一个实例,请求

    data = resp.read() #把结果读出来

    print('%d bytes received from %s.' % (len(data), url))

urls=['https://www.python.org/','https://www.yahoo.com/','https://github.com/']

start_time=time.time()

for url in urls:

    f(url)

print('同步cost is %s:'%(time.time()-start_time))

async_time_start=time.time() #异步的起始时间

gevent.joinall([

        gevent.spawn(f, 'https://www.python.org/'),

        gevent.spawn(f, 'https://www.yahoo.com/'),

        gevent.spawn(f, 'https://github.com/'),

])

print('异步cost is %s:'%(time.time()-async_time_start))

运行时间：几乎差不多，看不出异步的优势。

C:\abccdxddd\Oldboy\python-3.5.2-embed-amd64\python.exe C:/abccdxddd/Oldboy/Py_Exercise/Day10/爬虫.py

GET: https://www.python.org/

48751 bytes received from https://www.python.org/.

GET: https://www.yahoo.com/

480499 bytes received from https://www.yahoo.com/.

GET: https://github.com/

55394 bytes received from https://github.com/.

同步cost is 7.112711191177368:

GET: https://www.python.org/

48751 bytes received from https://www.python.org/.

GET: https://www.yahoo.com/

485666 bytes received from https://www.yahoo.com/.

GET: https://github.com/

55390 bytes received from https://github.com/.

异步cost is 4.510450839996338:

Process finished with exit code 0

5.因为gevent默认检测不到 urllib 进行的是否是io操作。要想让两者关联起来，需要再导入一个新函数（补丁）

from gevent import monkey，

monkey.patch_all()

from urllib import request

import gevent

import time

from gevent import monkey

monkey.patch_all() #把当前程序的所有io操作给我单独地做上标记

def f(url):

    print('GET: %s' % url)

    resp = request.urlopen(url) #赋给一个实例,请求

    data = resp.read() #把结果读出来

    print('%d bytes received from %s.' % (len(data), url))

urls=['https://www.python.org/','https://www.yahoo.com/','https://github.com/']

start_time=time.time()

for url in urls:

    f(url)

print('同步cost is %s:'%(time.time()-start_time))

async_time_start=time.time() #异步的起始时间

gevent.joinall([

        gevent.spawn(f, 'https://www.python.org/'),

        gevent.spawn(f, 'https://www.yahoo.com/'),

        gevent.spawn(f, 'https://github.com/'),

])

print('异步cost is %s:'%(time.time()-async_time_start))

运行结果：

C:\abccdxddd\Oldboy\python-3.5.2-embed-amd64\python.exe C:/abccdxddd/Oldboy/Py_Exercise/Day10/爬虫.py

GET: https://www.python.org/

48751 bytes received from https://www.python.org/.

GET: https://www.yahoo.com/

487577 bytes received from https://www.yahoo.com/.

GET: https://github.com/

55392 bytes received from https://github.com/.

同步cost is 5.784578323364258:

GET: https://www.python.org/

GET: https://www.yahoo.com/

GET: https://github.com/

480662 bytes received from https://www.yahoo.com/.

48751 bytes received from https://www.python.org/.

55394 bytes received from https://github.com/.

异步cost is 1.8721871376037598:

Process finished with exit code 0

Urllib--爬虫的更多相关文章

urllib爬虫（流程+案例）
网络爬虫是一种按照一定规则自动抓取万维网信息的程序.在如今网络发展,信息爆炸的时代,信息的处理变得尤为重要.而这之前就需要获取到数据.有关爬虫的概念可以到网上查看详细的说明,今天在这里介绍一下使用ur ...
urllib爬虫模块
网络爬虫也称为网络蜘蛛.网络机器人,抓取网络的数据.其实就是用Python程序模仿人点击浏览器并访问网站,而且模仿的越逼真越好.一般爬取数据的目的主要是用来做数据分析,或者公司项目做数据测试,公司业务 ...
【Python】python3中urllib爬虫开发
以下是三种方法 ①First Method 最简单的方法 ②添加data,http header 使用Request对象 ③CookieJar import urllib.request from h ...
python爬虫 urllib模块url编码处理
案例:爬取使用搜狗根据指定词条搜索到的页面数据(例如爬取词条为‘周杰伦'的页面数据) import urllib.request # 1.指定url url = 'https://www.sogou. ...
[Python]新手写爬虫全过程（已完成）
今天早上起来,第一件事情就是理一理今天该做的事情,瞬间get到任务,写一个只用python字符串内建函数的爬虫,定义为v1.0,开发中的版本号定义为v0.x.数据存放?这个是一个练手的玩具,就写在tx ...
[Python]爬虫v0.1
#coding:utf-8 import urllib ###### #爬虫v0.1 利用urlib2 和字符串内建函数 ###### # 获取网页内容 def getHtml(url): page ...
[Python]新手写爬虫全过程（转）
今天早上起来,第一件事情就是理一理今天该做的事情,瞬间get到任务,写一个只用python字符串内建函数的爬虫,定义为v1.0,开发中的版本号定义为v0.x.数据存放?这个是一个练手的玩具,就写在tx ...
Python 网络爬虫
爬虫介绍爬取图片爬取文本爬虫相关模块:re 爬虫相关模块:urllib 爬虫相关模块:urllib2 爬虫相关模块:cookielib 爬虫相关模块:requests 爬取需要登录的页面
vue+node+mongoDB火车票H5（七）-- nodejs 爬12306查票接口
菜鸟一枚,业余一直想做个火车票查票的H5,前端页面什么的已经写好了,node+mongoDB 也写了一个车站的接口,但接下来的爬12306获取车次信息数据一直卡住,网上的爬12306的大部分是pyt ...
爬虫---request+++urllib
网络爬虫(又被称为网页蜘蛛,网络机器人,在FOAF社区中间,更经常的称为网页追逐者),是一种按照一定的规则,自动地抓取万维网信息的程序或者脚本.另外一些不常使用的名字还有蚂蚁.自动索引.模拟程序或者蠕 ...

随机推荐

Python：numpy中的tile函数
在学习机器学习实教程时,实现KNN算法的代码中用到了numpy的tile函数,因此对该函数进行了一番学习: tile函数位于python模块 numpy.lib.shape_base中,他的功能是重复 ...
「日常训练」Divisibility by Eight（Codeforces Round 306 Div.2 C）
题意与分析极简单的数论+思维题. 代码 #include <bits/stdc++.h> #define MP make_pair #define PB emplace_back #de ...
Git 与 GitHub
Git 这个年代,不会点Git真不行啦,少年别问问什么,在公司你就知道了~ Git是一个协同开发的工具,主要作用是进行版本控制,而且还能自动检测代码是否发生变化. 一. 安装下载地址:https:/ ...
第五模块：WEB开发基础第2章·JavaScript基础
01-JavaScript的历史发展过程 02-js的引入方式和输出 03-命名规范和变量的声明定义 04-五种基本数据类型 05-运算符 06-字符串处理 07-数据类型转换 08-流程控制语句if ...
BFC与合并浅析
BFC BFC 全称 Block Formatting Context.每个渲染区域用formatting context表示,它决定了其子元素将如何定位,以及和其他元素的关系和相互作用在正常流中的盒 ...
[转载]启动tomcat时，一直卡在Deploying web application directory这块的解决方案
转载:https://www.cnblogs.com/mycifeng/p/6972446.html 本来今天正常往服务器上扔一个tomcat 部署一个项目的, 最后再启动tomcat 的时候发现项 ...
最短路径算法（II）
什么??你问我为什么不在一篇文章写完所有方法?? Hmm…其实我是想的,但是博皮的加载速度再带上文章超长图片超多的话… 可能这辈子都打不开了吧… 上接https://www.cnblogs.com/U ...
opencv-学习笔记(4)-模糊
opencv-学习笔记(4)-模糊本章要点: 4种模糊方式 2d卷积 Cv2.filter2D(‘图像对象’,‘目标图像这里直接设为-1即可’,kernal,anchor(-1,-1)) 一般后一个 ...
Case 降序升序排列
select nc.Class_Name,hn.home_news_id,hn.hemo_id,hn.hemo_Date, hn.hemo_title,hemo_order from Hemo_New ...
基于Kubernetes（k8s）的RabbitMQ 集群
目前,有很多种基于Kubernetes搭建RabbitMQ集群的解决方案.今天笔者今天将要讨论我们在Fuel CCP项目当中所采用的方式.这种方式加以转变也适用于搭建RabbitMQ集群的一般方法.所 ...

Urllib--爬虫

Urllib--爬虫的更多相关文章

随机推荐

热门专题