爬虫五 Beautifulsoup模块

一介绍

Beautiful Soup 是一个可以从HTML或XML文件中提取数据的Python库.它能够通过你喜欢的转换器实现惯用的文档导航,查找,修改文档的方式.Beautiful Soup会帮你节省数小时甚至数天的工作时间.你可能在寻找 Beautiful Soup3 的文档,Beautiful Soup 3 目前已经停止开发,官网推荐在现在的项目中使用Beautiful Soup 4, 移植到BS4

#安装 Beautiful Soup

pip install beautifulsoup4

#安装解析器

Beautiful Soup支持Python标准库中的HTML解析器,还支持一些第三方的解析器,其中一个是 lxml .根据操作系统不同,可以选择下列方法来安装lxml:

$ apt-get install Python-lxml

$ easy_install lxml

$ pip3 install lxml

另一个可供选择的解析器是纯Python实现的 html5lib , html5lib的解析方式与浏览器相同,可以选择下列方法来安装html5lib:

$ apt-get install Python-html5lib

$ easy_install html5lib

$ pip3 install html5lib

下表列出了主要的解析器,以及它们的优缺点,官网推荐使用lxml作为解析器,因为效率更高. 在Python2.7.3之前的版本和Python3中3.2.2之前的版本,必须安装lxml或html5lib, 因为那些Python版本的标准库中内置的HTML解析方法不够稳定.

解析器	使用方法	优势	劣势
Python标准库	`BeautifulSoup(markup, "html.parser")`	Python的内置标准库执行速度适中文档容错能力强	Python 2.7.3 or 3.2.2)前的版本中文档容错能力差
lxml HTML 解析器	`BeautifulSoup(markup, "lxml")`	速度快文档容错能力强	需要安装C语言库
lxml XML 解析器	`BeautifulSoup(markup, ["lxml", "xml"])` `BeautifulSoup(markup, "xml")`	速度快唯一支持XML的解析器	需要安装C语言库
html5lib	`BeautifulSoup(markup, "html5lib")`	最好的容错性以浏览器的方式解析文档生成HTML5格式的文档	速度慢不依赖外部扩展

二基本使用

html_doc = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were

<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,

<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;

and they lived at the bottom of a well.</p>

<p class="story">...</p>

"""

#基本使用：容错处理,文档的容错能力指的是在html代码不完整的情况下,使用该模块可以识别该错误。使用BeautifulSoup解析上述代码,能够得到一个 BeautifulSoup 的对象,并能按照标准的缩进格式的结构输出

from bs4 import BeautifulSoup

soup=BeautifulSoup(html_doc,'lxml') #具有容错功能

res=soup.prettify() #处理好缩进，结构化显示

print(res)

三标签选择器

#1、标签选择器：即直接通过标签名字选择，选择速度快，如果存在多个相同的标签则只返回第一个

#2、获取标签的名称

#3、获取标签的属性

#4、获取标签的内容

#5、嵌套选择

#6、子节点、子孙节点

#7、父节点、祖先节点

#8、兄弟节点

#1、标签选择器：即直接通过标签名字选择，选择速度快，如果存在多个相同的标签则只返回第一个

html_doc = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were

<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,

<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;

and they lived at the bottom of a well.</p>

<p class="story">...</p>

"""

from bs4 import BeautifulSoup

soup=BeautifulSoup(html_doc,'lxml')

print(soup.head)

print(type(soup.head)) #<class 'bs4.element.Tag'>

print(soup.p) #存在多个相同的标签则只返回第一个

print(soup.a) #存在多个相同的标签则只返回第一个

#2、获取标签的名称

print(soup.p.name)

#3、获取标签的属性

print(soup.p.attrs)

#4、获取表的内容

print(soup.p.string)

'''

对下面的这种结构，soup.p.string 返回为None,因为里面有a

<p id='list-1'>

    哈哈哈哈

    <a class='sss'>

        <span>

            <h1>aaaa</h1>

        </span>

    </a>

    <b>bbbbb</b>

</p>

'''

#5、嵌套选择

print(soup.head.title.string)

print(soup.body.a.string)

#6、子节点、子孙节点

html_doc = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title">

    <b>The Dormouse's story</b>

    Once upon a time there were three little sisters; and their names were

    <a href="http://example.com/elsie" class="sister" id="link1">

        <span>Elsie</span>

    </a>,

    <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

    <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;

    and they lived at the bottom of a well.

</p>

<p class="story">...</p>

"""

from bs4 import BeautifulSoup

soup=BeautifulSoup(html_doc,'lxml')

print(soup.p.contents) #p下所有子节点

print(soup.p.children) #得到一个迭代器,包含p下所有子节点

for i,child in enumerate(soup.p.children):

    print(i,child)

print(soup.p.descendants) #获取子孙节点,p下所有的标签都会选择出来

for i,child in enumerate(soup.p.descendants):

    print(i,child)

#7、父节点、祖先节点

print(soup.a.parent) #获取a标签的父节点

print(soup.a.parents) #找到a标签所有的祖先节点，父亲的父亲，父亲的父亲的父亲...

#8、兄弟节点

print(soup.a.next_siblings) #得到生成器对象

print(soup.a.previous_siblings) #得到生成器对象

四标准选择器

#find与findall:用法完全一样,可根据标签名,属性,内容查找文档，但是find只找第一个元素

html_doc = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title">

    <b>The Dormouse's story</b>

    Once upon a time there were three little sisters; and their names were

    <a href="http://example.com/elsie" class="sister" id="link1">

        <span>Elsie</span>

    </a>

    <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

    <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;

    and they lived at the bottom of a well.

</p>

<p class="story">...</p>

"""

from bs4 import BeautifulSoup

soup=BeautifulSoup(html_doc,'lxml')

#1、按照标签名查找

# print(soup.find_all('a'))

# print(soup.find_all('a',id='link3'))

# print(soup.find_all('a',id='link3',attrs={'class':"sister"}))

#

# print(soup.find_all('a')[0].find('span')) #嵌套查找

#2、按照属性查找

# print(soup.p.find_all(attrs={'id':'link1'})) #等同于print(soup.find_all(id='link1'))

# print(soup.p.find_all(attrs={'class':'sister'}))

#

# print(soup.find_all(class_='sister'))

#3、按照文本内容查找

print(soup.p.find_all(text="The Dormouse's story")) # 按照完整内容匹配（是==而不是in）,得到的结果也是内容

#更多：https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html#find

五 CSS选择器

##该模块提供了select方法来支持css

html_doc = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title">

    <b>The Dormouse's story</b>

    Once upon a time there were three little sisters; and their names were

    <a href="http://example.com/elsie" class="sister" id="link1">

        <span>Elsie</span>

    </a>

    <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

    <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;

    <div class='panel-1'>

        <ul class='list' id='list-1'>

            <li class='element'>Foo</li>

            <li class='element'>Bar</li>

            <li class='element'>Jay</li>

        </ul>

        <ul class='list list-small' id='list-2'>

            <li class='element'><h1 class='yyyy'>Foo</h1></li>

            <li class='element xxx'>Bar</li>

            <li class='element'>Jay</li>

        </ul>

    </div>

    and they lived at the bottom of a well.

</p>

<p class="story">...</p>

"""

from bs4 import BeautifulSoup

soup=BeautifulSoup(html_doc,'lxml')

#1、CSS选择器

print(soup.p.select('.sister'))

print(soup.select('.sister span'))

print(soup.select('#link1'))

print(soup.select('#link1 span'))

print(soup.select('#list-2 .element.xxx'))

print(soup.select('#list-2')[0].select('.element')) #可以一直select,但其实没必要,一条select就可以了

# 2、获取属性

print(soup.select('#list-2 h1')[0].attrs)

# 3、获取内容

print(soup.select('#list-2 h1')[0].get_text())

六总结

# 总结:

#1、推荐使用lxml解析库

#2、讲了三种选择器:标签选择器,find与find_all，css选择器

    1、标签选择器筛选功能弱,但是速度快

    2、建议使用find,find_all查询匹配单个结果或者多个结果

    3、如果对css选择器非常熟悉建议使用select

#3、记住常用的获取属性attrs和文本值get_text()的方法

爬虫五 Beautifulsoup模块的更多相关文章

爬虫五 Beautifulsoup模块详细
一.基本使用 from bs4 import BeautifulSoup htmlCharset = "GB2312" soup=BeautifulSoup(html_doc,'l ...
Python爬虫之Beautifulsoup模块的使用
一 Beautifulsoup模块介绍 Beautiful Soup 是一个可以从HTML或XML文件中提取数据的Python库.它能够通过你喜欢的转换器实现惯用的文档导航,查找,修改文档的方式.Be ...
Python 爬虫三 beautifulsoup模块
beautifulsoup模块 BeautifulSoup模块 BeautifulSoup是一个模块,该模块用于接收一个HTML或XML字符串,然后将其进行格式化,之后遍可以使用他提供的方法进行快速查 ...
Python网络爬虫之BeautifulSoup模块
一.介绍: Beautiful Soup 是一个可以从HTML或XML文件中提取数据的Python库.它能够通过你喜欢的转换器实现惯用的文档导航,查找,修改文档的方式.Beautiful Soup会帮 ...
爬虫利器BeautifulSoup模块使用
一.简介 BeautifulSoup 是一个可以从HTML或XML文件中提取数据的Python库.它能够通过你喜欢的转换器实现惯用的文档导航,查找,修改文档的方式,同时应用场景也是非常丰富,你可以使用 ...
python爬虫主要就是五个模块：爬虫启动入口模块，URL管理器存放已经爬虫的URL和待爬虫URL列表，html下载器，html解析器，html输出器同时可以掌握到urllib2的使用、bs4（BeautifulSoup）页面解析器、re正则表达式、urlparse、python基础知识回顾（set集合操作）等相关内容。
本次python爬虫百步百科,里面详细分析了爬虫的步骤,对每一步代码都有详细的注释说明,可通过本案例掌握python爬虫的特点: 1.爬虫调度入口(crawler_main.py) # coding: ...
【爬虫入门手记03】爬虫解析利器beautifulSoup模块的基本应用
[爬虫入门手记03]爬虫解析利器beautifulSoup模块的基本应用 1.引言网络爬虫最终的目的就是过滤选取网络信息,因此最重要的就是解析器了,其性能的优劣直接决定这网络爬虫的速度和效率.Bea ...
【网络爬虫入门03】爬虫解析利器beautifulSoup模块的基本应用
[网络爬虫入门03]爬虫解析利器beautifulSoup模块的基本应用 1.引言网络爬虫最终的目的就是过滤选取网络信息,因此最重要的就是解析器了,其性能的优劣直接决定这网络爬虫的速度和效率.B ...
第三百二十五节，web爬虫，scrapy模块标签选择器下载图片，以及正则匹配标签
第三百二十五节,web爬虫,scrapy模块标签选择器下载图片,以及正则匹配标签标签选择器对象 HtmlXPathSelector()创建标签选择器对象,参数接收response回调的html对象需 ...

随机推荐

SQL——使用游标进行遍历
前两天一个同事大叔问了这样一个问题,他要对表做个类似foreach的效果,问我怎么搞.我想了想,就拿游标回答他,当时事实上也没用过数据库中的游标,可是曾经用过ADO里面的,感觉应该几乎相同. 今天闲下 ...
1. Two Sum【easy】
1. Two Sum[easy] Given an array of integers, return indices of the two numbers such that they add up ...
CentOS 7 安装以及配置桌面环境
一.安装 GNOME 桌面 1.安装命令: yum groupinstall "GNOME Desktop" "X Window System" " ...
Charles 的几个调试技巧
Charles 是一个网络调试的代理工具,类似 Windows 下的 Fildder,这里介绍下几个常用的调试技巧,使用的版本是 Charles 4. 1.移动端抓包在移动开发中,经常会遇到在手机上 ...
row format delimited fields terminated by ','
row format delimited fields terminated by ',' 以','结尾的行格式分隔字段
PHP编译选项
PHP安装 ./configure --prefix=/usr/local/php --with-config-file-path=/usr/local/php/etc --with-mysql=/u ...
PHP 关掉浏览器还会执行代码
ignore_user_abort();//关掉浏览器,PHP脚本也可以继续执行. set_time_limit(0);// 通过set_time_limit(0)可以让程序无限制的执行下去 $int ...
Spring MVC错误处理
以下示例显示如何在使用Spring Web MVC框架的表单中使用错误处理和验证器.首先使用Eclipse IDE来创建一个WEB工程,实现一个输入用户信息提交验证提示的功能.并按照以下步骤使用Spr ...
Eclipse 视图(View)
Eclipse 视图关于视图 Eclipse视图允许用户以图表形式更直观的查看项目的元数据. 例如,项目导航视图中显示的文件夹和文件图形表示在另外一个编辑窗口中相关的项目和属性视图. Eclipse ...
Codeforces Round #398 (Div. 2) BCD
B:The Queue 题目大意:你要去办签证,那里上班时间是[s,t), 你知道那一天有n个人会来办签证,他们分别是在时间点ai来的.每个人办业务要花相同的时间x,问你什么时候来排队等待的时间最少 ...

爬虫五 Beautifulsoup模块

一 介绍

二 基本使用

三 标签选择器

四 标准选择器