MapReduce之WordCount

用户统计文件中的单词出现的个数

注意各个文件的导包，job的封装步骤

WordCountMapper.java

package top.wintp.mapreduce.wordcount;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.LongWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Mapper;

import java.io.IOException;

/**

 * @description: description:

 * <p>

 * @author: upuptop

 * <p>

 * @qq: 337081267

 * <p>

 * @CSDN: http://blog.csdn.net/pyfysf

 * <p>

 * @cnblogs: http://www.cnblogs.com/upuptop

 * <p>

 * @blog: http://wintp.top

 * <p>

 * @email: pyfysf@163.com

 * <p>

 * @time: 2019/05/2019/5/21

 * <p>

 */

public class WordCountMapper extends Mapper<LongWritable, Text, Text, IntWritable> {

    private Text K = new Text();

    private IntWritable V = new IntWritable(1);

    @Override

    protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException {

        String line = value.toString();

        String[] words = line.split(" ");

        for (String word : words) {

            K.set(word);

            context.write(K, V);

        }

    }

}

WordCountReduce

package top.wintp.mapreduce.wordcount;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Reducer;

import java.io.IOException;

/**

 * @description: description:

 * <p>

 * @author: upuptop

 * <p>

 * @qq: 337081267

 * <p>

 * @CSDN: http://blog.csdn.net/pyfysf

 * <p>

 * @cnblogs: http://www.cnblogs.com/upuptop

 * <p>

 * @blog: http://wintp.top

 * <p>

 * @email: pyfysf@163.com

 * <p>

 * @time: 2019/05/2019/5/21

 * <p>

 */

public class WordCountReduce extends Reducer<Text, IntWritable, Text, IntWritable> {

    private IntWritable V = new IntWritable();

    @Override

    protected void reduce(Text key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException {

        int sum = 0;

        for (IntWritable value : values) {

            sum += value.get();

        }

        V.set(sum);

        context.write(key, V);

    }

}

WordCountRunner

package top.wintp.mapreduce.wordcount;

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.fs.Path;

import org.apache.hadoop.io.IntWritable;

import org.apache.hadoop.io.Text;

import org.apache.hadoop.mapreduce.Job;

import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;

import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

import org.apache.hadoop.util.Tool;

import org.apache.hadoop.util.ToolRunner;

/**

 * @description: description:

 * <p>

 * @author: upuptop

 * <p>

 * @qq: 337081267

 * <p>

 * @CSDN: http://blog.csdn.net/pyfysf

 * <p>

 * @cnblogs: http://www.cnblogs.com/upuptop

 * <p>

 * @blog: http://wintp.top

 * <p>

 * @email: pyfysf@163.com

 * <p>

 * @time: 2019/05/2019/5/21

 * <p>

 */

public class WordCountRunner implements Tool {

    private Configuration conf;

    public int run(String[] strings) throws Exception {

        //封装job

        Job job = Job.getInstance(this.conf);

        job.setJarByClass(WordCountRunner.class);

        job.setMapperClass(WordCountMapper.class);

        job.setReducerClass(WordCountReduce.class);

        job.setMapOutputKeyClass(Text.class);

        job.setMapOutputValueClass(IntWritable.class);

        job.setOutputKeyClass(Text.class);

        job.setOutputValueClass(IntWritable.class);

        FileInputFormat.setInputPaths(job, new Path("E:/input/wordcount/"));

        FileOutputFormat.setOutputPath(job, new Path("E:/output/wordcount/" + System.currentTimeMillis()));

        //提交任务

        int result = job.waitForCompletion(true) ? 0 : 1;

        return result;

    }

    public void setConf(Configuration configuration) {

        this.conf = configuration;

    }

    public Configuration getConf() {

        return this.conf;

    }

    public static void main(String[] args) throws Exception {

        int status = ToolRunner.run(new WordCountRunner(), args);

        System.exit(status);

    }

}

log4j.properties

log4j.rootLogger=INFO, stdout

log4j.appender.stdout=org.apache.log4j.ConsoleAppender

log4j.appender.stdout.layout=org.apache.log4j.PatternLayout

log4j.appender.stdout.layout.ConversionPattern=%d %p [%c] - %m%n

log4j.appender.logfile=org.apache.log4j.FileAppender

log4j.appender.logfile.File=target/spring.log

log4j.appender.logfile.layout=org.apache.log4j.PatternLayout

log4j.appender.logfile.layout.ConversionPattern=%d %p [%c] - %m%n

MapReduce之WordCount的更多相关文章

Java编程MapReduce实现WordCount
Java编程MapReduce实现WordCount 1.编写Mapper package net.toocruel.yarn.mapreduce.wordcount; import org.apac ...
eclipse运行mapreduce的wordcount
1,eclipse安装hadoop插件插件下载地址:链接: https://pan.baidu.com/s/1U4_6kLFNiKeLsGfO7ahXew 提取码: as9e 下载hadoop-ec ...
MapReduce实现WordCount
package algorithm; import java.io.IOException; import java.util.StringTokenizer; import org.apache.h ...
Hadoop实战5:MapReduce编程-WordCount统计单词个数-eclipse-java-windows环境
Hadoop研发在java环境的拓展一背景由于一直使用hadoop streaming形式编写mapreduce程序,所以目前的hadoop程序局限于python语言.下面为了拓展java语言研 ...
Hadoop实战3:MapReduce编程-WordCount统计单词个数-eclipse-java-ubuntu环境
之前习惯用hadoop streaming环境编写python程序,下面总结编辑java的eclipse环境配置总结,及一个WordCount例子运行. 一下载eclipse安装包及hadoop插件 ...
Hadoop 6、第一个mapreduce程序 WordCount
1.程序代码 Map: import java.io.IOException; import org.apache.hadoop.io.IntWritable; import org.apache.h ...
Hadoop Mapreduce中wordcount 过程解析
将文件split 文件1: 分割结果: hello world ...
三.hadoop mapreduce之WordCount例子
目录: 目录见文章1 这个案列完成对单词的计数,重写map,与reduce方法,完成对mapreduce的理解. Mapreduce初析 Mapreduce是一个计算框架,既然是做计算的框架,那么表现 ...
大数据技术 - 通俗理解MapReduce之WordCount（三）
上一章我们编写了简单的 MapReduce 程序,掌握这些就能编写大多数数据处理的代码.但是 MapReduce 框架提供给用户的能力并不止如此,本章我们仍然以上一章 word count 为例,继续 ...
大数据技术 - 通俗理解MapReduce之WordCount（二）
上一章我们搭建了分布式的 Hadoop 集群.本章我们介绍 Hadoop 框架中的一个核心模块 - MapReduce.MapReduce 是并行计算模块,顾名思义,它包含两个主要的阶段,map 阶段 ...

随机推荐

在Delphi中关于UDP协议的实现
原文地址:在Delphi中关于UDP协议的实现作者:菜心首先我把UDP无连接协议的套接字调用时序图表示出来在我把在Delphi中使用UDP协议实现数据通讯收发的实现方法总结如下: 例子描述:下 ...
VS2013编译Qt5.6.0静态库，并提供了百度云下载（乌合之众）good
获取qt5.6.0源码包直接去www.qt.io下载就好了,这里就不详细说了. 这里是我已经编译好的** 链接:http://pan.baidu.com/s/1pLb6wVT 密码: ak7y ** ...
QT中的SOCKET编程
转自:http://mylovejsj.blog.163.com/blog/static/38673975200892010842865/ QT中的SOCKET编程 2008-10-07 23:13 ...
Mac上使用brew安装nvm来支持多版本的Nodejs
brew方式如果机器没有安装过node,那么首先brew install nvm安装nvm. 其次需要在shell的配置文件(~/.bashrc, ~/.profile, or ~/.zshrc)中 ...
Zookeeper详解-安装（四）
ZooKeeper服务器是用Java创建的,它在JVM上运行.你需要使用JDK 6或更高版本. 步骤1:验证Java安装相信你已经在系统上安装了Java环境.现在只需使用以下命令验证它. $ jav ...
Laravel --- 【转】安装调试利器 Laravel Debugbar
[转]http://www.tuicool.com/articles/qYfmmur 1.简介 Laravel Debugbar 在 Laravel 5 中集成了 PHP Debug Bar ,用于显 ...
spring 5.x 系列第17篇 —— 整合websocket (xml配置方式)
源码Gitub地址:https://github.com/heibaiying/spring-samples-for-all 一.说明 1.1 项目结构说明项目模拟一个简单的群聊功能,为区分不同的聊 ...
Redis Ubuntu 安装
1.使用 root 用户登录 Ubuntu 2. wget http://download.redis.io/releases/redis-5.0.3.tar.gz 下载最新的稳定版本到 redis ...
docker相关使用
安装docker 在CentOS 7上安装docker-ce,首先检查系统中是否已经安装过docker及相关依赖: $ sudo yum remove docker docker-client doc ...
There is no getter for property named 'username' in 'class Model1.User'-----报错解决
There is no getter for property named 'username' in 'class Model1.User' -----Model Model1.User'中没有名为 ...

MapReduce之WordCount

MapReduce之WordCount的更多相关文章

随机推荐

热门专题