python 爬虫系列09-selenium+拉钩

使用selenium爬取拉勾网职位

 from selenium import webdriver

 from lxml import etree

 import re

 import time

 from selenium.webdriver.support.ui import WebDriverWait

 from selenium.webdriver.support import expected_conditions as EC

 from selenium.webdriver.common.by import By

 class LagouSpider(object):

     driver_path = r"D:\driver\chromedriver.exe"

     def __init__(self):

         self.driver = webdriver.Chrome(executable_path=LagouSpider.driver_path)

         self.url = 'https://www.lagou.com/jobs/list_%E4%BA%91%E8%AE%A1%E7%AE%97?labelWords=&fromSearch=true&suginput='

         self.positions = []

     def run(self):

         self.driver.get(self.url)

         while True:

             source = self.driver.page_source

             WebDriverWait(driver=self.driver,timeout=10).until(

                 EC.presence_of_element_located((By.XPATH, "//div[@class='pager_container']/span[last()]"))

             )

             self.parse_list_page(source)

             try:

                 next_btn = self.driver.find_element_by_xpath("//div[@class='pager_container']/span[last()]")

                 if "pager_next_disabled" in next_btn.get_attribute("class"):

                     break

                 else:

                     next_btn.click()

             except:

                 print(source)

             time.sleep(1)

     def parse_list_page(self,source):

         html = etree.HTML(source)

         links = html.xpath("//a[@class='position_link']/@href")

         for link in links:

             self.request_detail_page(link)

             time.sleep(1)

     def request_detail_page(self,url):

         # self.driver.get(url)

         print()

         print(url)

         print()

         self.driver.execute_script("window.open('%s')" % url)

         self.driver.switch_to.window(self.driver.window_handles[1])

         WebDriverWait(self.driver,timeout=10).until(

             EC.presence_of_element_located((By.XPATH,"//div[@class='job-name']/span[@class='name']"))

         )

         source = self.driver.page_source

         self.parse_detail_page(source)

         self.driver.close()

         self.driver.switch_to.window(self.driver.window_handles[0])

     def parse_detail_page(self,source):

         html = etree.HTML(source)

         position_name = html.xpath("//span[@class='name']/text()")[0]

         job_request_spans = html.xpath("//dd[@class='job_request']//span")

         salary = job_request_spans[0].xpath('.//text()')[0].strip()

         city = job_request_spans[1].xpath(".//text()")[0].strip()

         city = re.sub(r"[\s/]", "", city)

         work_years = job_request_spans[2].xpath(".//text()")[0].strip()

         work_years = re.sub(r"[\s/]", "", work_years)

         education = job_request_spans[3].xpath(".//text()")[0].strip()

         education = re.sub(r"[\s/]", "", education)

         desc = "".join(html.xpath("//dd[@class='job_bt']//text()")).strip()

         company_name = html.xpath("//h2[@class='f1']/text()")

         position = {

             'name': position_name,

             'company_name': company_name,

             'salary': salary,

             'city': city,

             'work_years': work_years,

             'education': education,

             'desc': desc

         }

         self.positions.append(position)

         print(position)

 if __name__ == '__main__':

     spider = LagouSpider()

     spider.run()

python 爬虫系列09-selenium+拉钩的更多相关文章

python爬虫动态html selenium.webdriver
python爬虫:利用selenium.webdriver获取渲染之后的页面代码! 1 首先要下载浏览器驱动: 常用的是chromedriver 和phantomjs chromedirver下载地址 ...
Python爬虫之设置selenium webdriver等待
Python爬虫之设置selenium webdriver等待 ajax技术出现使异步加载方式呈现数据的网站越来越多,当浏览器在加载页面时,页面上的元素可能并不是同时被加载完成,这给定位元素的定位增加 ...
Python爬虫系列-Selenium详解
自动化测试工具,支持多种浏览器.爬虫中主要用来解决JavaScript渲染的问题. 用法讲解模拟百度搜索网站过程: from selenium import webdriver from selen ...
PYTHON 爬虫笔记七:Selenium库基础用法
知识点一:Selenium库详解及其基本使用什么是Selenium selenium 是一套完整的web应用程序测试系统,包含了测试的录制(selenium IDE),编写及运行(Selenium ...
python爬虫之初始Selenium
1.初始 Selenium[1] 是一个用于Web应用程序测试的工具.Selenium测试直接运行在浏览器中,就像真正的用户在操作一样.支持的浏览器包括IE(7, 8, 9, 10, 11),Moz ...
python 爬虫系列教程方法总结及推荐
爬虫,是我学习的比较多的,也是比较了解的.打算写一个系列教程,网上搜罗一下,感觉别人写的已经很好了,我没必要重复造轮子了. 爬虫不过就是访问一个页面然后用一些匹配方式把自己需要的东西摘出来. 而访问页 ...
$python爬虫系列（2）—— requests和BeautifulSoup库的基本用法
本文主要介绍python爬虫的两大利器:requests和BeautifulSoup库的基本用法. 1. 安装requests和BeautifulSoup库可以通过3种方式安装: easy_inst ...
Python爬虫系列 - 初探：爬取旅游评论
Python爬虫目前是基于requests包,下面是该包的文档,查一些资料还是比较方便. http://docs.python-requests.org/en/master/ POST发送内容格式爬 ...
python爬虫系列（2）—— requests和BeautifulSoup
本文主要介绍python爬虫的两大利器:requests和BeautifulSoup库的基本用法. 1. 安装requests和BeautifulSoup库可以通过3种方式安装: easy_inst ...
Python爬虫系列（七）：提高解析效率
如果仅仅因为想要查找文档中的<a>标签而将整片文档进行解析,实在是浪费内存和时间.最快的方法是从一开始就把<a>标签以外的东西都忽略掉. SoupStrainer 类可以定义文 ...

随机推荐

C# JSON使用的常用技巧(二)
JSON在php里一句json_encode就可以得到在C#里我们同样也很容易的可以得到用到的类库:Newtonsoft.Json.dll 实体类: class Cat { public stri ...
升级Ubuntu 12.04下的gcc到4.7
我们知道C++11标准开始支持类内初始化(in-class initializer),Qt creator编译出现error,不支持这个特性,原因在于,Ubuntu12.04默认的是使用gcc4.6, ...
解决iReport打不开的一种方法
解决iReport打不开的一种方法 iReport版本:iReport-5.6.0-windows-installer.exe 系统:Win7 64位 JDK:1.7 在公司电脑安装没问题,能打开,但 ...
.net IAsyncResult 异步操作
//定义一个委托 public delegate int DoSomething(int count); //BeginInvoke 的回调函数 private static void Execute ...
那些年我们追过的SQL
SQL是大学必修课程之一二维表结构,看着就是一种美感. 针对近期感情,聊一聊,在平时容易犯的一个错误,看看你是不是中枪了. 我们还是选用传统的student表(请不要考虑表的结构是否合理)ID ...
ComicEnhancerPro 系列教程十七：二值化图像去毛刺
作者:马健邮箱:stronghorse_mj@hotmail.com 主页:http://www.comicer.com/stronghorse/ 发布:2017.07.23 教程十七:二值化图像去毛 ...
solidity_mapping_implementation
solidity 中 mapping 是如何存储的为了探测 solidity mapping 如何实现,我构造了一个简单的合约. 先说结论,实际上 mapping的访问成本并不比直接访问storag ...
数据库抽象层 pdo
一 . PDO的连接 $host = "localhost"; $dbname = "hejuntest"; $username = "root&qu ...
Ubuntu 14.10，准备C/C++的编译环境
Ubuntu缺省情况下,并没有提供C/C++的编译环境,因此还需要手动安装. 如果单独安装gcc以及g++比较麻烦,幸运的是,为了能够编译Ubuntu的内核,Ubuntu提供了一个build-esse ...
Mybatis 的动态 SQL 语句
<if>标签我们根据实体类的不同取值,使用不同的 SQL 语句来进行查询. 比如在 id 如果不为空时可以根据 id 查询, 如果 username 不同空时还要加入用户名作为条件.这种 ...

python 爬虫系列09-selenium+拉钩

python 爬虫系列09-selenium+拉钩的更多相关文章

随机推荐

热门专题