HDFS源码分析数据块复制之PendingReplicationBlocks

PendingReplicationBlocks实现了所有正在复制的数据块的记账工作。它实现以下三个主要功能：

1、记录此时正在复制的块；

2、一种对复制请求进行跟踪的粗粒度计时器；

3、一个定期识别未执行复制请求的线程。

我们先看下它内部有哪些成员变量，如下：

// 块和正在进行的块复制信息的映射集合
private final Map<Block, PendingBlockInfo> pendingReplications;
// 复制请求超时的块列表
private final ArrayList<Block> timedOutItems;
// 后台工作线程
Daemon timerThread = null;
// 文件系统是否正在运行的标志位
private volatile boolean fsRunning = true;
//
// It might take anywhere between 5 to 10 minutes before
// a request is timed out.
// 在一个请求超时之前可能需要5到10分钟
// 请求超时阈值，默认为5分钟
private long timeout = 5 * 60 * 1000;
// 超时检查固定值：5分钟
private final static long DEFAULT_RECHECK_INTERVAL = 5 * 60 * 1000;

首先是pendingReplications，它是块和正在进行的块复制信息的映射集合，所有正在复制的数据块及其对应复制信息都会被加入到这个集合。数据块复制信息PendingBlockInfo是对数据块开始复制时间timeStamp、待复制的目标数据节点列表List<DatanodeDescriptor>实例targets的一个封装，代码如下：

/**
* An object that contains information about a block that
* is being replicated. It records the timestamp when the
* system started replicating the most recent copy of this
* block. It also records the list of Datanodes where the
* replication requests are in progress.
*
* 正在被复制的块信息。它记录系统开始复制块最新副本的时间，也记录复制请求正在执行的数据节点列表。
*/
static class PendingBlockInfo {
// 时间戳
private long timeStamp;
// 待复制的目标数据节点列表
private final List<DatanodeDescriptor> targets;
// 构造方法
PendingBlockInfo(DatanodeDescriptor[] targets) {
// 时间戳赋值为当前时间
this.timeStamp = now();
this.targets = targets == null ? new ArrayList<DatanodeDescriptor>()
: new ArrayList<DatanodeDescriptor>(Arrays.asList(targets));
}
long getTimeStamp() {
return timeStamp;
}
// 设置时间戳为当前时间
void setTimeStamp() {
timeStamp = now();
}
// 增加复制数量，即增加目标数据节点
void incrementReplicas(DatanodeDescriptor... newTargets) {
if (newTargets != null) {
for (DatanodeDescriptor dn : newTargets) {
targets.add(dn);
}
}
}
// 减少复制数量，即减少目标数据节点
void decrementReplicas(DatanodeDescriptor dn) {
targets.remove(dn);
}
// 获取复制数量，即或许待复制的数据节点数目
int getNumReplicas() {
return targets.size();
}
}

它的构造方法中，即将时间戳timeStamp赋值为当前时间，并且提供了设置时间戳为当前时间的setTimeStamp()方法。同时提供了增加复制数量、减少复制数量、获取复制数量相关的三个方法，均是对待复制的目标数据节点列表的增加、减少与计数操作，上面注释很清楚，不再详述！

另外两个比较重要的变量就是复制请求超时的块列表timedOutItems和后台工作线程timerThread。由后台工作线程周期性的检查pendingReplications列表中的待复制数据块，看看其是否超时，如果超时的话，将其加入timedOutItems列表。后台工作线程timerThread的初始化如下：

// 启动块复制监控线程
void start() {
timerThread = new Daemon(new PendingReplicationMonitor());
timerThread.start();
}

它实际上是借助PendingReplicationMonitor来完成的。PendingReplicationMonitor实现了Runnable接口，是一个周期性工作的线程，用于浏览从未完成它们复制请求的数据块，这个从未完成实际上就是在规定时间内还未完成的数据块复制信息。PendingReplicationMonitor的实现如下：

/*
* A periodic thread that scans for blocks that never finished
* their replication request.
* 一个周期性线程，用于浏览从未完成它们复制请求的数据块
*/
class PendingReplicationMonitor implements Runnable {
@Override
public void run() {
// 如果标志位fsRunning为true，即文件系统正常运行，则while循环一直进行
while (fsRunning) {
// 检查周期：取timeout，最高为5分钟
long period = Math.min(DEFAULT_RECHECK_INTERVAL, timeout);
try {
// 检查方法
pendingReplicationCheck();
// 线程休眠period
Thread.sleep(period);
} catch (InterruptedException ie) {
if(LOG.isDebugEnabled()) {
LOG.debug("PendingReplicationMonitor thread is interrupted.", ie);
}
}
}
}
/**
* Iterate through all items and detect timed-out items
* 通过所有项目迭代检测超时项目
*/
void pendingReplicationCheck() {
// 使用synchronized关键字对pendingReplications进行同步
synchronized (pendingReplications) {
// 获取集合pendingReplications的迭代器
Iterator<Map.Entry<Block, PendingBlockInfo>> iter =
pendingReplications.entrySet().iterator();
// 记录当前时间now
long now = now();
if(LOG.isDebugEnabled()) {
LOG.debug("PendingReplicationMonitor checking Q");
}
// 遍历pendingReplications集合中的每个元素
while (iter.hasNext()) {
// 取出每个<Block, PendingBlockInfo>条目
Map.Entry<Block, PendingBlockInfo> entry = iter.next();
// 取出Block对应的PendingBlockInfo实例pendingBlock
PendingBlockInfo pendingBlock = entry.getValue();
// 判断pendingBlock自其生成时的timeStamp以来到现在，是否已超过timeout时间
if (now > pendingBlock.getTimeStamp() + timeout) {
// 超过的话，
// 取出timeout实例block
Block block = entry.getKey();
// 使用synchronized关键字对timedOutItems进行同步
synchronized (timedOutItems) {
// 将block添加入复制请求超时的块列表timedOutItems
timedOutItems.add(block);
}
LOG.warn("PendingReplicationMonitor timed out " + block);
// 从迭代器中移除该条目
iter.remove();
}
}
}
}
}

在它的run()方法内，如果标志位fsRunning为true，即文件系统正常运行，则while循环一直进行，然后在while循环内：

1、先取检查周期period：取timeout，最高为5分钟；

2、调用pendingReplicationCheck()方法进行检查；

3、线程休眠period时间，再次进入while循环。

pendingReplicationCheck的实现逻辑也很简单，如下：

使用synchronized关键字对pendingReplications进行同步：

1、获取集合pendingReplications的迭代器iter；

2、记录当前时间now；

3、遍历pendingReplications集合中的每个元素：

3.1、取出每个<Block, PendingBlockInfo>条目；

3.2、取出Block对应的PendingBlockInfo实例pendingBlock；

3.3、判断pendingBlock自其生成时的timeStamp以来到现在，是否已超过timeout时间，超过的话：

3.3.1、取出timeout实例block；

3.3.2、使用synchronized关键字对timedOutItems进行同步，使用synchronized关键字对timedOutItems进行同步；

3.3.3、从迭代器中移除该条目。

PendingReplicationBlocks还提供了获取复制超时块数组的getTimedOutBlocks()方法，代码如下：

/**
* Returns a list of blocks that have timed out their
* replication requests. Returns null if no blocks have
* timed out.
* 返回一个其复制请求已超时的数据块列表，如果没有则返回null
*/
Block[] getTimedOutBlocks() {
/ 使用synchronized关键字对timedOutItems进行同步
synchronized (timedOutItems) {
// 如果timedOutItems中没有数据，则直接返回null
if (timedOutItems.size() <= 0) {
return null;
}
// 将Block列表timedOutItems转换成Block数组
Block[] blockList = timedOutItems.toArray(
new Block[timedOutItems.size()]);
// 清空Block列表timedOutItems
timedOutItems.clear();
// 返回Block数组
return blockList;
}
}

PendingReplicationBlocks另外还提供了增加一个块到正在进行的块复制信息列表中的increment()方法和减少正在复制请求的数量的decrement()方法，代码如下：

/**
* Add a block to the list of pending Replications
* 增加一个块到正在进行的块复制信息列表中
*
* @param block The corresponding block
* @param targets The DataNodes where replicas of the block should be placed
*/
void increment(Block block, DatanodeDescriptor[] targets) {
// 使用synchronized关键字对pendingReplications进行同步
synchronized (pendingReplications) {
// 根据Block实例block先从集合pendingReplications中查找
PendingBlockInfo found = pendingReplications.get(block);
if (found == null) {
// 如果没有找到，直接put进去，利用DatanodeDescriptor[]的实例targets构造PendingBlockInfo对象
pendingReplications.put(block, new PendingBlockInfo(targets));
} else {
// 如果之前存在，增加复制数量，即增加目标数据节点
found.incrementReplicas(targets);
// 设置时间戳为当前时间
found.setTimeStamp();
}
}
}
/**
* One replication request for this block has finished.
* Decrement the number of pending replication requests
* for this block.
* 针对给定数据块的一个复制请求已完成。针对该数据块，减少正在复制请求的数量。
*
* @param The DataNode that finishes the replication
*/
void decrement(Block block, DatanodeDescriptor dn) {
// 使用synchronized关键字对pendingReplications进行同步
synchronized (pendingReplications) {
// 根据Block实例block先从集合pendingReplications中查找
PendingBlockInfo found = pendingReplications.get(block);
if (found != null) {
if(LOG.isDebugEnabled()) {
LOG.debug("Removing pending replication for " + block);
}
// 减少复制数量，即减少目标数据节点
found.decrementReplicas(dn);
// 如果数据块对应的复制数量总数小于等于0，复制工作完成，
// 直接从pendingReplications集合中移除该数据块及其对应信息
if (found.getNumReplicas() <= 0) {
pendingReplications.remove(block);
}
}
}
}

以及统计块数量和块复制数量的方法，如下：

/**
* The total number of blocks that are undergoing replication
* 正在被复制的块的总数
*/
int size() {
return pendingReplications.size();
}
/**
* How many copies of this block is pending replication?
* 块复制的总量
*/
int getNumReplicas(Block block) {
synchronized (pendingReplications) {
PendingBlockInfo found = pendingReplications.get(block);
if (found != null) {
return found.getNumReplicas();
}
}
return 0;
}

上述方法代码逻辑都很简单，而且注释也很详细，此处不再过多赘述！

HDFS源码分析数据块复制之PendingReplicationBlocks的更多相关文章

HDFS源码分析数据块复制监控线程ReplicationMonitor（二）
HDFS源码分析数据块复制监控线程ReplicationMonitor(二)
HDFS源码分析数据块复制监控线程ReplicationMonitor（一）
ReplicationMonitor是HDFS中关于数据块复制的监控线程,它的主要作用就是计算DataNode工作,并将复制请求超时的块重新加入到待调度队列.其定义及作为线程核心的run()方法如下: ...
HDFS源码分析数据块复制选取复制源节点
数据块的复制当然需要一个源数据节点,从其上拷贝数据块至目标数据节点.那么数据块复制是如何选取复制源节点的呢?本文我们将针对这一问题进行研究. 在BlockManager中,chooseSourceDa ...
HDFS源码分析数据块校验之DataBlockScanner
DataBlockScanner是运行在数据节点DataNode上的一个后台线程.它为所有的块池管理块扫描.针对每个块池,一个BlockPoolSliceScanner对象将会被创建,其运行在一个单独 ...
HDFS源码分析数据块汇报之损坏数据块检测checkReplicaCorrupt()
无论是第一次,还是之后的每次数据块汇报,名字名字节点都会对汇报上来的数据块进行检测,看看其是否为损坏的数据块.那么,损坏数据块是如何被检测的呢?本文,我们将研究下损坏数据块检测的checkReplic ...
HDFS源码分析数据块之CorruptReplicasMap
CorruptReplicasMap用于存储文件系统中所有损坏数据块的信息.仅当它的所有副本损坏时一个数据块才被认定为损坏.当汇报数据块的副本时,我们隐藏所有损坏副本.一旦一个数据块被发现完好副本达到 ...
HDFS源码分析之数据块及副本状态BlockUCState、ReplicaState
关于数据块.副本的介绍,请参考文章<HDFS源码分析之数据块Block.副本Replica>. 一.数据块状态BlockUCState 数据块状态用枚举类BlockUCState来表示,代 ...
HDFS源码分析心跳汇报之数据块汇报
在<HDFS源码分析心跳汇报之数据块增量汇报>一文中,我们详细介绍了数据块增量汇报的内容,了解到它是时间间隔更长的正常数据块汇报周期内一个smaller的数据块汇报,它负责将DataNod ...
HDFS源码分析心跳汇报之数据块增量汇报
在<HDFS源码分析心跳汇报之BPServiceActor工作线程运行流程>一文中,我们详细了解了数据节点DataNode周期性发送心跳给名字节点NameNode的BPServiceAct ...

随机推荐

WPF 自动选择dll,以SQLite为例
在学习sqlite的过程中,发现它的dll是区分32位和64位的,起初觉得很恼火,但是仔细看了下, 发现让程序自行选择dll其实也不是一件很麻烦的事情,如下: 1>创建一个sqlite数据 2& ...
css sticky footer 布局手机端
什么是css sticky footer 布局? 通常在手机端写页面会遇到如下情况页面长度很短不足以撑起一屏,此时希望页脚在页面的底部而当页面超过一屏时候,页脚会在文章的底部 ,网上有许多办法, ...
MVC中的过滤器/拦截器怎么写
创建一个AuthenticateFilterAttribute(即过滤器/拦截器) 引用System.Web.Mvc; public class AuthenticateFilterAttribute ...
常见的 Git 命令：
开始一个工作区(参见:git help tutorial) clone 克隆一个仓库到一个新目录 init 创建一个空的 Git 仓库或重新初始化一个已存在的仓库在当前变更上工作(参见:git he ...
Qualcomm download 所需要的 contents.xml
Platform MSM8917 PM8937 PMI8940 在 Qualcomm code base 中, amss下有許多 MSM89xx 之類的 folder, 這些是為了不同 chip 所產 ...
mdev详解【转】
转自:http://blog.chinaunix.net/uid-29401328-id-5019678.html 一.概述 mdev是busybox提供的一个工具,用在嵌入式系统中,相当于简化版的u ...
C++学习（一）：现代C++尝试
C++是一门与时俱进的语言. 早期的C++关注的主要问题是通用性,却没有太多关注易用性的问题,使得C++成为了一门多范式语言,但是使用门槛较高. 从2011开始,C++的标准进行了较大的更新,开始更多 ...
【C/C++】快速排序的两种实现思路
方法一:不断填坑,一次确定一个值.http://blog.csdn.net/morewindows/article/details/6684558 #include<stdio.h> vo ...
转：Java并发编程：volatile关键字解析
Java并发编程:volatile关键字解析 Java并发编程:volatile关键字解析 volatile这个关键字可能很多朋友都听说过,或许也都用过.在Java 5之前,它是一个备受争议的关键字, ...
利用js实现table增加一行
简单的方法: 用jquery插件,比如设置该table的id为mytable <table id="mytable"> <tr> <td> 第一 ...

HDFS源码分析数据块复制之PendingReplicationBlocks

HDFS源码分析数据块复制之PendingReplicationBlocks的更多相关文章

随机推荐

热门专题