Spark Streaming性能优化系列-怎样获得和持续使用足够的集群计算资源？

一：数据峰值的巨大影响

1. 数据确实不稳定，比如晚上的时候訪问流量特别大

2. 在处理的时候比如GC的时候耽误时间会产生delay延迟

二：Backpressure：数据的反压机制

基本思想：依据上一次计算的Job的一些信息评估来决定下一个Job数据接收的速度。

怎样限制Spark接收数据的速度？

Spark Streaming在接收数据的时候必须把当前的数据接收完毕才干接收下一条数据。

源代码解析

RateController：

1. RateController是监听器。继承自StreamingListener.

/**

 * A StreamingListener that receives batch completion updates, and maintains

 * an estimate of the speed at which this stream should ingest messages,

 * given an estimate computation from a `RateEstimator`

 */

private[streaming] abstract class RateController(val streamUID: Int, rateEstimator: RateEstimator)

    extends StreamingListener with Serializable {

问题来了。RateContoller什么时候被调用的呢？

BackPressure是依据上一次计算的Job信息来评估下一个Job数据接收的速度。

因此肯定是在JobScheduler中被调用的。

1. 在JobScheduler的start方法中rateController方法是从inputStream中获取的。

// attach rate controllers of input streams to receive batch completion updates

for {

  inputDStream <- ssc.graph.getInputStreams

  rateController <- inputDStream.rateController

} ssc.addStreamingListener(rateController)

2.  然后将此消息增加到listenerBus中。

/** Add a [[org.apache.spark.streaming.scheduler.StreamingListener]] object for

  * receiving system events related to streaming.

  */

def addStreamingListener(streamingListener: StreamingListener) {

  scheduler.listenerBus.addListener(streamingListener)

}

}

3. 在StreamingListenerBus源代码例如以下：

/** Asynchronously passes StreamingListenerEvents to registered StreamingListeners. */

private[spark] class StreamingListenerBus

  extends AsynchronousListenerBus[StreamingListener, StreamingListenerEvent]("StreamingListenerBus")

  with Logging {

  private val logDroppedEvent = new AtomicBoolean(false)

  override def onPostEvent(listener: StreamingListener, event: StreamingListenerEvent): Unit = {

    event match {

      case receiverStarted: StreamingListenerReceiverStarted =>

        listener.onReceiverStarted(receiverStarted)

      case receiverError: StreamingListenerReceiverError =>

        listener.onReceiverError(receiverError)

      case receiverStopped: StreamingListenerReceiverStopped =>

        listener.onReceiverStopped(receiverStopped)

      case batchSubmitted: StreamingListenerBatchSubmitted =>

        listener.onBatchSubmitted(batchSubmitted)

      case batchStarted: StreamingListenerBatchStarted =>

        listener.onBatchStarted(batchStarted)

      case batchCompleted: StreamingListenerBatchCompleted =>

        listener.onBatchCompleted(batchCompleted)

4.  在RateController就实现了onBatchCompleted

5. RateController中onBatchCompleted详细实现例如以下：

override def onBatchCompleted(batchCompleted: StreamingListenerBatchCompleted) {

  val elements = batchCompleted.batchInfo.streamIdToInputInfo

  for {

    processingEnd <- batchCompleted.batchInfo.processingEndTime

    workDelay <- batchCompleted.batchInfo.processingDelay

    waitDelay <- batchCompleted.batchInfo.schedulingDelay

    elems <- elements.get(streamUID).map(_.numRecords)

  } computeAndPublish(processingEnd, elems, workDelay, waitDelay)

}

6.  RateController中computeAndPulish源代码例如以下：

/**

 * Compute the new rate limit and publish it asynchronously.

 */

private def computeAndPublish(time: Long, elems: Long, workDelay: Long, waitDelay: Long): Unit =

  Future[Unit] {

//评估新的更加合适Rate速度。

val newRate = rateEstimator.compute(time, elems, workDelay, waitDelay)

    newRate.foreach { s =>

      rateLimit.set(s.toLong)

      publish(getLatestRate())

    }

  }

7.  当中publish实现是在ReceiverRateController中。

8. 将pulish消息给ReceiverTracker.

/**

 * A RateController that sends the new rate to receivers, via the receiver tracker.

 */

private[streaming] class ReceiverRateController(id: Int, estimator: RateEstimator)

    extends RateController(id, estimator) {

  override def publish(rate: Long): Unit =

//由于会有非常多RateController所以会有详细Id

    ssc.scheduler.receiverTracker.sendRateUpdate(id, rate)

}

9.  在ReceiverTracker中sendRateUpdate源代码例如以下：

此时的endpoint是ReceiverTrackerEndpoint.

/** Update a receiver's maximum ingestion rate */

def sendRateUpdate(streamUID: Int, newRate: Long): Unit = synchronized {

  if (isTrackerStarted) {

    endpoint.send(UpdateReceiverRateLimit(streamUID, newRate))

  }

}

10. 在ReceiverTrackerEndpoint的receive方法中就接收到了发来的消息。

case UpdateReceiverRateLimit(streamUID, newRate) =>

//依据receiverTrackingInfos获取info信息，然后依据endpoint获取通信句柄。

//此时endpoint是ReceiverSupervisor的endpoint通信实体。

  for (info <- receiverTrackingInfos.get(streamUID); eP <- info.endpoint) {

    eP.send(UpdateRateLimit(newRate))

  }

11. 因此在ReceiverSupervisorImpl中接收到ReceiverTracker发来的消息。

/** RpcEndpointRef for receiving messages from the ReceiverTracker in the driver */

private val endpoint = env.rpcEnv.setupEndpoint(

  "Receiver-" + streamId + "-" + System.currentTimeMillis(), new ThreadSafeRpcEndpoint {

    override val rpcEnv: RpcEnv = env.rpcEnv

    override def receive: PartialFunction[Any, Unit] = {

      case StopReceiver =>

        logInfo("Received stop signal")

        ReceiverSupervisorImpl.this.stop("Stopped by driver", None)

      case CleanupOldBlocks(threshTime) =>

        logDebug("Received delete old batch signal")

        cleanupOldBlocks(threshTime)

      case UpdateRateLimit(eps) =>

        logInfo(s"Received a new rate limit: $eps.")

        registeredBlockGenerators.foreach { bg =>

          bg.updateRate(eps)

        }

    }

  })

12. RateLimiter中updateRate源代码例如以下：

/**

 * Set the rate limit to `newRate`. The new rate will not exceed the maximum rate configured by

//这里有最大限制，由于你的集群处理规模是有限的。

//Spark Streaming可能执行在YARN之上。由于多个计算框架都在执行的话。资源就//更有限了。

 * {{{spark.streaming.receiver.maxRate}}}, even if `newRate` is higher than that.

 *

 * @param newRate A new rate in events per second. It has no effect if it's 0 or negative.

 */

private[receiver] def updateRate(newRate: Long): Unit =

  if (newRate > 0) {

    if (maxRateLimit > 0) {

      rateLimiter.setRate(newRate.min(maxRateLimit))

    } else {

      rateLimiter.setRate(newRate)

    }

  }

整体流程图例如以下：

总结:

每次上一个Batch Duration的Job执行完毕之后。都会返回JobCompleted等信息，基于这些信息产生一个新的Rate，然后将新的Rate通过远程通信交给了Executor中，而Executor也会依据Rate又一次设置Rate大小。

Spark Streaming性能优化系列-怎样获得和持续使用足够的集群计算资源？的更多相关文章

Spark Streaming性能优化: 如何在生产环境下应对流数据峰值巨变
1.为什么引入Backpressure 默认情况下,Spark Streaming通过Receiver以生产者生产数据的速率接收数据,计算过程中会出现batch processing time > ...
Spark Streaming性能调优
数据接收并行度调优(一) 通过网络接收数据时(比如Kafka.Flume),会将数据反序列化,并存储在Spark的内存中.如果数据接收称为系统的瓶颈,那么可以考虑并行化数据接收.每一个输入DStrea ...
SparkSQL的一些用法建议和Spark的性能优化
1.写在前面 Spark是专为大规模数据处理而设计的快速通用的计算引擎,在计算能力上优于MapReduce,被誉为第二代大数据计算框架引擎.Spark采用的是内存计算方式.Spark的四大核心是Spa ...
[MySQL性能优化系列]提高缓存命中率
1. 背景通常情况下,能用一条sql语句完成的查询,我们尽量不用多次查询完成.因为,查询次数越多,通信开销越大.但是,分多次查询,有可能提高缓存命中率.到底使用一个复合查询还是多个独立查询,需要根据 ...
[MySQL性能优化系列]巧用索引
1. 普通青年的索引使用方式假设我们有一个用户表 tb_user,内容如下: name age sex jack 22 男 rose 21 女 tom 20 男 ... ... ... 执行SQL语 ...
[MySQL性能优化系列]LIMIT语句优化
1. 背景假设有如下SQL语句: SELECT * FROM table1 LIMIT offset, rows 这是一条典型的LIMIT语句,常见的使用场景是,某些查询返回的内容特别多,而客户端处 ...
PLSQL_性能优化系列14_Oracle High Water Level高水位分析
2014-10-04 Created By BaoXinjian 一.摘要 PLSQL_性能优化系列14_Oracle High Water Level高水位分析高水位线好比水库中储水的水位线,用于 ...
[Android 性能优化系列]降低你的界面布局层次结构的一部分
大家假设喜欢我的博客,请关注一下我的微博,请点击这里(http://weibo.com/kifile),谢谢转载请标明出处(http://blog.csdn.net/kifile),再次感谢原文地 ...
Spark Streaming性能调优详解
Spark Streaming性能调优详解 Spark 2015-04-28 7:43:05 7896℃ 0评论分享到微博下载为PDF 2014 Spark亚太峰会会议资料下载.< ...

随机推荐

python基础——16（re模块，内存管理）
一.内存管理 1.垃圾回收机制不能被程序访问到的数据,就称之为垃圾. 1.1.引用计数引用计数是用来记录值的内存地址被记录的次数的. 每一次对值地址的引用都使该值的引用计数+1:每一次对值地址的释 ...
【10】css hack原理及常用hack
[10]css hack原理及常用hack 原理:利用不同浏览器对CSS的支持和解析结果不一样编写针对特定浏览器样式.常见的hack有1)属性hack.2)选择器hack.3)IE条件注释 IE条件注 ...
【01】报错：webpack 不是内部或不可执行命令
[02] webpack 不是内部或不可执行命令一般来安装完之后是可以直接执行的你可以执行 webpack -v 或者是 webpack --help 这样的就是正确的,我的问题的解决办法是将 ...
脑阔疼的双层SQLserver游标
本来简单的双层游标没啥的,内层游标需要读取的是视图的内容,一直报“当前命令发生了严重错误.应放弃任何可能产生的结果.”的错误.无可奈何尝试先将视图的数据放到表变量中,之后再用游标遍历表变量. 简直很怀 ...
北京师范大学第十五届ACM决赛-重现赛
Another Server 时间限制:1秒空间限制:262144K 题目描述何老师某天在机房里搞事情的时候,发现机房里有n台服务器,从1到n标号,同时有2n-2条网线,从1到2n-2标号,其中第 ...
C# 方法冒号this的用法
public Class1(string host, int port, string password = null):this() { this.Host=host; this.Port=port ...
HDU-5319 Painter，深搜标记！
Painter 题意:有一个棋盘n行,列数不超过50,用red和blue给这个棋盘涂色,每个格子每种颜色最多涂一次,如果两种颜色都涂了则该格子颜色为Green;red以斜杠'\'方式涂色,bule以' ...
BZOJ 3750: [POI2015]Pieczęć 【模拟】
Description 一张n*m的方格纸,有些格子需要印成黑色,剩下的格子需要保留白色. 你有一个a*b的印章,有些格子是凸起(会沾上墨水)的.你需要判断能否用这个印章印出纸上的图案.印的过程中需要 ...
bzoj2850巧克力王国
巧克力王国 Time Limit: 60 Sec Memory Limit: 512 MBSubmit: 861 Solved: 325[Submit][Status][Discuss] Desc ...
oracle 当中，（+）是什么意思
SELECT A.id, B.IDDFROM A, BWHERE A.id(+)=B.IDD等价于SELECT A.id, B.IDDFROM A RIGHT OUTER JOIN B ON ( A. ...

Spark Streaming性能优化系列-怎样获得和持续使用足够的集群计算资源？

Spark Streaming性能优化系列-怎样获得和持续使用足够的集群计算资源？的更多相关文章

随机推荐

热门专题