利用Requests库写爬虫

基本Get请求：

#-*- coding:utf-8 -*-

import requests

url = 'http://www.baidu.com'

r = requests.get(url)

print r.text

带参数Get请求：

#-*- coding:utf-8 -*-

import requests

url = 'http://www.baidu.com'

payload = {'key1': 'value1', 'key2': 'value2'}

r = requests.get(url, params=payload)

print r.text

POST请求模拟登陆及一些返回对象的方法：

#-*- coding:utf-8 -*-

import requests

url1 = 'http://www.exanple.com/login'#登陆地址

url2 = "http://www.example.com/main"#需要登陆才能访问的地址

data={"user":"user","password":"pass"}

headers = { "Accept":"text/html,application/xhtml+xml,application/xml;",

            "Accept-Encoding":"gzip",

            "Accept-Language":"zh-CN,zh;q=0.8",

            "Referer":"http://www.example.com/",

            "User-Agent":"Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.90 Safari/537.36"

            }

res1 = requests.post(url1, data=data, headers=headers)

res2 = requests.get(url2, cookies=res1.cookies, headers=headers)

print res2.content#获得二进制响应内容

print res2.raw#获得原始响应内容,需要stream=True

print res2.raw.read(50)

print type(res2.text)#返回解码成unicode的内容

print res2.url

print res2.history#追踪重定向

print res2.cookies

print res2.cookies['example_cookie_name']

print res2.headers

print res2.headers['Content-Type']

print res2.headers.get('content-type')

print res2.json#讲返回内容编码为json

print res2.encoding#返回内容编码

print res2.status_code#返回http状态码

print res2.raise_for_status()#返回错误状态码

使用Session()对象的写法（Prepared Requests）:

#-*- coding:utf-8 -*-

import requests

s = requests.Session()

url1 = 'http://www.exanple.com/login'#登陆地址

url2 = "http://www.example.com/main"#需要登陆才能访问的地址

data={"user":"user","password":"pass"}

headers = { "Accept":"text/html,application/xhtml+xml,application/xml;",

            "Accept-Encoding":"gzip",

            "Accept-Language":"zh-CN,zh;q=0.8",

            "Referer":"http://www.example.com/",

            "User-Agent":"Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.90 Safari/537.36"

            }

prepped1 = requests.Request('POST', url1,

    data=data,

    headers=headers

).prepare()

s.send(prepped1)

'''

也可以这样写

res = requests.Request('POST', url1,

data=data,

headers=headers

)

prepared = s.prepare_request(res)

# do something with prepped.body

# do something with prepped.headers

s.send(prepared)

'''

prepare2 = requests.Request('POST', url2,

    headers=headers

).prepare()

res2 = s.send(prepare2)

print res2.content

另一种写法 :

#-*- coding:utf-8 -*-

import requests

s = requests.Session()

url1 = 'http://www.exanple.com/login'#登陆地址

url2 = "http://www.example.com/main"#需要登陆才能访问的页面地址

data={"user":"user","password":"pass"}

headers = { "Accept":"text/html,application/xhtml+xml,application/xml;",

            "Accept-Encoding":"gzip",

            "Accept-Language":"zh-CN,zh;q=0.8",

            "Referer":"http://www.example.com/",

            "User-Agent":"Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.90 Safari/537.36"

            }

res1 = s.post(url1, data=data)

res2 = s.post(url2)

print(resp2.content)

SessionApi

其他的一些请求方式

>>> r = requests.put("http://httpbin.org/put")

>>> r = requests.delete("http://httpbin.org/delete")

>>> r = requests.head("http://httpbin.org/get")

>>> r = requests.options("http://httpbin.org/get")

遇到的问题:

在cmd下执行，遇到个小错误:

UnicodeEncodeError:'gbk' codec can't encode character u'\xbb' in

position 23460: illegal multibyte sequence

分析:
1、Unicode是编码还是解码

UnicodeEncodeError

很明显是在编码的时候出现了错误

2、用了什么编码

'gbk' codec can't encode character

使用GBK编码出错

解决办法：

确定当前字符串，比如

#-*- coding:utf-8 -*-

import requests

url = 'http://www.baidu.com'

r = requests.get(url)

print r.encoding

>utf-8

已经确定html的字符串是utf-8的，则可以直接去通过utf-8去编码。

print r.text.encode('utf-8')

作者：Jelvis
链接：http://www.jianshu.com/p/e1f8b690b951
來源：简书
著作权归作者所有。商业转载请联系作者获得授权，非商业转载请注明出处。

利用Requests库写爬虫的更多相关文章

requests库写接口测试框架初学习
学习网址: https://docs.microsoft.com/en-us/openspecs/windows_protocols/ms-dscpm/ff75b907-415d-4220-89 ...
爬虫入门实例：利用requests库爬取笔趣小说网
w3cschool上的来练练手,爬取笔趣看小说http://www.biqukan.com/, 爬取<凡人修仙传仙界篇>的所有章节 1.利用requests访问目标网址,使用了get方法 ...
Requests库网络爬虫实战
实例一:页面的爬取 >>> import requests>>> r= requests.get("https://item.jd.com/1000037 ...
python利用requests库模拟post请求时json的使用
我们都见识过requests库在静态网页的爬取上展现的威力,我们日常见得最多的为get和post请求,他们最大的区别在于安全性上: 1.GET是通过URL方式请求,可以直接看到,明文传输. 2.POS ...
requests库（爬虫）
北京理工大学嵩天老师的课程:http://www.icourse163.org/course/BIT-1001870001 官方文档:http://docs.python-requests.org/e ...
python脚本实例002－利用requests库实现应用登录
#! /usr/bin/python # coding:utf-8 #导入requests库 import requests #获取会话 s = requests.session() #创建登录数据 ...
使用python requests库写接口自动化测试--记录学习过程中遇到的坑（1）
一直听说python requests库对于接口自动化测试特别合适,但由于自身代码基础薄弱,一直没有实践: 这次赶上公司项目需要,同事小伙伴们一起学习写接口自动化脚本,听起来特别给力,赶紧实践一把: ...
Python爬虫：HTTP协议、Requests库（爬虫学习第一天）
HTTP协议: HTTP(Hypertext Transfer Protocol):即超文本传输协议.URL是通过HTTP协议存取资源的Internet路径,一个URL对应一个数据资源. HTTP协议 ...
利用requests库访问网站
1.关于requests库函数 Response对象包含服务器返回的所有信息,也包含请求的Request信息. 访问百度二十次 import requests def getHTMLText(url ...

随机推荐

[opencv] 图像几何变换：旋转，缩放，斜切
几何变换几何变换可以看成图像中物体(或像素)空间位置改变,或者说是像素的移动. 几何运算需要空间变换和灰度级差值两个步骤的算法,像素通过变换映射到新的坐标位置,新的位置可能是在几个像素之间,即不一定 ...
自定义ribbon规则
关于ribbon的知识:. 在微服务架构中,业务都会被拆分成一个独立的服务,服务与服务的通讯是基于http restful的.Spring cloud有两种服务调用方式,一种是ribbon+restT ...
【Asp.net入门2-01】C#基本功能
C#是一种功能强大的语言,但并不是所有程序员都熟悉我们将在本书中讨论的所有功能.因此, 本章将介绍优秀的Web窗体程序员需要了解的C#语言功能. 本章仅简要介绍每一项功能.有关C#语言本身的知识不是本 ...
设置CMD默认路径
用CMD每一次都得切换路径,很麻烦. 所以,需要设置一下CMD默认路径: 1.打开注册表编辑器(WIN+R打开运行.输入regedit) 2.定位到: “HKEY_CURRENT_USER\Softw ...
RCNN,fast R-CNN,faster R-CNN
转自:https://www.cnblogs.com/skyfsm/p/6806246.html object detection我的理解,就是在给定的图片中精确找到物体所在位置,并标注出物体的类别. ...
Shell记录-Shell脚本基础（六）
watch是一个非常实用的命令,基本所有的Linux发行版都带有这个小工具,如同名字一样,watch可以帮你监测一个命令的运行结果,省得你一遍遍的手动运行. 1．命令格式 watch[参数][命令] ...
【总结】CSS透明度大汇总
近年来,CSS不透明算得上是一种相当流行的技术,但在跨浏览器支持上,对于开发者来说,可以说是一件令人头疼的事情.目前还没有一个通用方法,以确保透明度设置可以在目前使用的所有浏览器上有效. 这篇汇总主要 ...
活学活用，CSS清除浮动的4种方法
清除浮动这个问题,做前端的应该再熟悉不过了,咱是个新人,所以还是记个笔记,做个积累,努力学习向大神靠近. CSS清除浮动的方法网上一搜,大概有N多种,用过几种,说下个人感受. 1.结尾处加空div标签 ...
Tomcat与Spring中的事件机制详解
最近在看tomcat源码,源码中出现了大量事件消息,可以说整个tomcat的启动流程都可以通过事件派发机制串起来,研究透了tomcat的各种事件消息,基本上对tomcat的启动流程也就有了一个整体的认 ...
SMTP——MIME
MIME 基础知识 MIME 表示多用途 Internet 邮件扩允协议.MIME 扩允了基本的面向文本的 Internet 邮件系统,以便可以在消息中包含二进制附件. MIME 信息由正常的 Int ...

利用Requests库写爬虫

遇到的问题:

解决办法：

利用Requests库写爬虫的更多相关文章

随机推荐

热门专题