爬虫抓取5大门户网站和电商数据day1:基础环境搭建

最新想用爬虫实现抓取五大门户网站（搜狐、新浪、网易、腾讯、凤凰网）和电商数据（天猫，京东，聚美等），今天第一天先搭建下环境和测试。

采用maven+xpath+ HttpClient+正则表达式。

maven pom.xml配置文件信息

<dependency>

      <groupId>junit</groupId>

      <artifactId>junit</artifactId>

      <version>4.12</version>

      <scope>test</scope>

    </dependency>

    <dependency>

       <groupId>org.apache.spark</groupId>

       <artifactId>spark-core_2.10</artifactId>

       <version>1.6.0</version>

    </dependency>

    <dependency>

        <groupId>org.apache.spark</groupId>

        <artifactId>spark-sql_2.10</artifactId>

        <version>1.6.0</version>

    </dependency>

    <dependency>

      <groupId>org.apache.spark</groupId>

      <artifactId>spark-hive_2.10</artifactId>

      <version>1.6.0</version>

    </dependency>

    <dependency>

          <groupId>org.apache.spark</groupId>

          <artifactId>spark-streaming_2.10</artifactId>

          <version>1.6.0</version>

    </dependency>

    <dependency>

          <groupId>org.apache.hadoop</groupId>

          <artifactId>hadoop-client</artifactId>

          <version>2.6.0</version>

    </dependency>

    <dependency>

          <groupId>org.apache.spark</groupId>

          <artifactId>spark-streaming-kafka_2.10</artifactId>

          <version>1.6.0</version>

    </dependency>

    <dependency>

          <groupId>org.apache.spark</groupId>

          <artifactId>spark-graphx_2.10</artifactId>

          <version>1.6.0</version>

    </dependency>

    <!-- httpclient4.4 -->

        <dependency>

            <groupId>org.apache.httpcomponents</groupId>

            <artifactId>httpclient</artifactId>

            <version>4.4</version>

        </dependency>

        <!-- htmlcleaner -->

        <dependency>

            <groupId>net.sourceforge.htmlcleaner</groupId>

            <artifactId>htmlcleaner</artifactId>

            <version>2.10</version>

        </dependency>

        <!-- json -->

        <dependency>

            <groupId>org.json</groupId>

            <artifactId>json</artifactId>

            <version>20140107</version>

        </dependency>

        <!-- hbase -->

        <dependency>

            <groupId>org.apache.hbase</groupId>

            <artifactId>hbase-client</artifactId>

            <version>0.96.1.1-hadoop2</version>

        </dependency>

            <dependency>

            <groupId>org.apache.hbase</groupId>

            <artifactId>hbase-server</artifactId>

            <version>0.96.1.1-hadoop2</version>

        </dependency>

        <!-- redis 2.7.0-->

        <dependency>

            <groupId>redis.clients</groupId>

            <artifactId>jedis</artifactId>

            <version>2.7.0</version>

        </dependency>

        <!-- slf4j -->

        <dependency>

            <groupId>org.slf4j</groupId>

            <artifactId>slf4j-api</artifactId>

            <version>1.7.10</version>

        </dependency>

        <dependency>

            <groupId>org.slf4j</groupId>

            <artifactId>slf4j-log4j12</artifactId>

            <version>1.7.10</version>

        </dependency>

        <!-- quartz1.8.4 -->

        <dependency>

            <groupId>org.quartz-scheduler</groupId>

            <artifactId>quartz</artifactId>

            <version>1.8.4</version>

        </dependency>

        <!-- curator -->

        <dependency>

            <groupId>org.apache.curator</groupId>

            <artifactId>curator-framework</artifactId>

            <version>2.7.1</version>

        </dependency>

新建一个测试类：SpiderTest

  /**

     * url 入口，下载页面

     * @param url

     */

    public  static String downLoadCrawlurl(String url){

        String context = null;

        Logger logger = LoggerFactory.getLogger(SpiderTest.class);

        HttpClientBuilder create = HttpClientBuilder.create();

        HttpGet httpGet = new HttpGet(url);

        CloseableHttpClient build = create.build();

        try {

            CloseableHttpResponse response = build.execute( httpGet);

            HttpEntity entity = response.getEntity();

            context = EntityUtils.toString( entity );

            System.out.println("context:" + context);

        }

        catch ( ClientProtocolException e ) {

            e.printStackTrace();

        }

        catch ( IOException e ) {

            logger.info("download...." );

        }

        return context;

    }

public static void main( String[] args ) {

　　
   String url = "http://money.163.com/";
   downLoadCrawlurl(url);
 

}

爬虫抓取5大门户网站和电商数据day1:基础环境搭建的更多相关文章

PID控制器的应用：控制网络爬虫抓取速度
一.初识PID控制器冬天乡下人喜欢烤火取暖,常见的情形就是四人围着麻将桌,桌底放一盆碳火.有人觉得火不够大,那加点木炭吧,还不够,再加点.片刻之后,又觉得火太大,脚都快被烤熟了,那就取出一些木碳…… ...
爬虫抓取页面数据原理（php爬虫框架有很多）
爬虫抓取页面数据原理(php爬虫框架有很多 ) 一.总结 1.php爬虫框架有很多,包括很多傻瓜式的软件 2.照以前写过java爬虫的例子来看,真的非常简单,就是一个获取网页数据的类或者方法(这里的话 ...
python3.4学习笔记(十四) 网络爬虫实例代码，抓取新浪爱彩双色球开奖数据实例
python3.4学习笔记(十四) 网络爬虫实例代码,抓取新浪爱彩双色球开奖数据实例新浪爱彩双色球开奖数据URL:http://zst.aicai.com/ssq/openInfo/ 最终输出结果格 ...
爬虫技术 -- 进阶学习（七）简单爬虫抓取示例（附c#代码）
这是我的第一个爬虫代码...算是一份测试版的代码.大牛大神别喷... 通过给定一个初始的地址startPiont然后对网页进行捕捉,然后通过正则表达式对网址进行匹配. List<string&g ...
python 爬虫抓取心得
quanwei9958 转自 python 爬虫抓取心得分享 urllib.quote('要编码的字符串') 如果你要在url请求里面放入中文,对相应的中文进行编码的话,可以用: urllib.quo ...
爬虫技术（四）-- 简单爬虫抓取示例（附c#代码）
这是我的第一个爬虫代码...算是一份测试版的代码.大牛大神别喷... 通过给定一个初始的地址startPiont然后对网页进行捕捉,然后通过正则表达式对网址进行匹配. List<string&g ...
Java 实现 HttpClients+jsoup，Jsoup，htmlunit，Headless Chrome 爬虫抓取数据
最近整理一下手头上搞过的一些爬虫,有HttpClients+jsoup,Jsoup,htmlunit,HeadlessChrome 一,HttpClients+jsoup,这是第一代比较low,很快就 ...
Python爬虫抓取东方财富网股票数据并实现MySQL数据库存储
Python爬虫可以说是好玩又好用了.现想利用Python爬取网页股票数据保存到本地csv数据文件中,同时想把股票数据保存到MySQL数据库中.需求有了,剩下的就是实现了. 在开始之前,保证已经安装好 ...
python爬虫抓取哈尔滨天气信息（静态爬虫）
python 爬虫爬取哈尔滨天气信息 - http://www.weather.com.cn/weather/101050101.shtml 环境: windows7 python3.4(pip i ...

随机推荐

CCflow与基础框架组织机构整合
SELECT No,Name,Pass,FK_Dept,SID FROM Port_Emp SELECT No,Name,ParentNo FROM Port_Dept SELECT No,Name, ...
关于audio不能拖放
图一,图二均为wav格式文件图一为播放本地的音频,可以拖放图二为放在后台的音频,不可以拖放把这两个图片发给后台,让后台分析下两个的headers不同之处
JS 多个条件判断
// 多个条件判断 // 对象序列(Object) 推荐使用这一种 var obj = {'CJ':'成交', 'WCJ':'未成交'}; if (key in obj) { // TODO } // ...
html - body标签中相关标签
body标签中相关标签今日内容: 字体标签: h1~h6..... ...
bootstrap学习（四）表格
基础样式: 自适应沾满浏览器 <table class="table"> <tr> <th>序号</th> <th>姓名 ...
android5.1 隐藏状态栏
修改frameworks/base/core/res/res/values/dimens.xml文件中  <!-- ...
tomcat的首次登录配置
登录tomcat时需要输入账号密码,而账号密码需要在配置文件中配置好才能使用. 此处我们先点击取消,tomcat会弹出一个提示界面: 这个界面的大致意思是: 401未经授权您无权查看此页面. 如果您 ...
Bash: Removing leading zeroes from a variable
old=" # sed removes leading zeroes from stdin new=$(echo $old | sed 's/^0*//')
STL_Algorithm
#include <algorithm> #include <cstdio> using namespace std; /*虽然最后一个排列没有下一个排列,用next_perm ...
cboard进行访问，汉化

爬虫抓取5大门户网站和电商数据day1:基础环境搭建

爬虫抓取5大门户网站和电商数据day1:基础环境搭建的更多相关文章

随机推荐

热门专题