通过 spark.files 传入spark任务依赖的文件源码分析

版本：spak2.3

相关源码：org.apache.spark.SparkContext

在创建spark任务时候，往往会指定一些依赖文件，通常我们可以在spark-submit脚本使用--files /path/to/file指定来实现。

但是公司产品的架构是通过livy来调spark任务，livy的实现其实是对spark-submit的一个包装，所以如何指定依赖文件归根到底还是在spark这边。既然不能通过命令行--files指定，那在编程中怎么指定？任务在各个节点上运行时又是如何获取到这些文件的呢？

根据spark-submit的参数传递源码分析得知，spark-submit --files其实是由参数"spark.files"接收，所以在代码中可以通过sparkConf设置该参数。

比如：

SparkConf conf = new SparkConf();

conf.set("spark.files","/path/to/file");

//如果文件是放在hdfs上，可以通过conf.set("spark.files","hdfs:/path/to/file")指定，注意这里只需要加上个hdfs的schema即可，不需要ip port

spark官网关于该参数的解释：

spark.files　　Comma-separated list of files to be placed in the working directory of each executor. Globs are allowed.

具体怎么读取用户指定的文件相关源码在SparkContext.scala中，如下（--jars指定依赖jar包同理）：

def jars: Seq[String] = _jars

def files: Seq[String] = _files

...

_jars = Utils.getUserJars(_conf)

_files = _conf.getOption("spark.files").map(_.split(",")).map(_.filter(_.nonEmpty))

  .toSeq.flatten

...

// Add each JAR given through the constructor

if (jars != null) {

  jars.foreach(addJar)

}

if (files != null) {

  files.foreach(addFile)

}

addFile实现如下：

/**

* Add a file to be downloaded with this Spark job on every node.

*

* If a file is added during execution, it will not be available until the next TaskSet starts.

*

* @param path can be either a local file, a file in HDFS (or other Hadoop-supported

* filesystems), or an HTTP, HTTPS or FTP URI. To access the file in Spark jobs,

* use `SparkFiles.get(fileName)` to find its download location.

* @param recursive if true, a directory can be given in `path`. Currently directories are

* only supported for Hadoop-supported filesystems.

*    1. 文件会下载到每一个节点

*    2. 如果在运行中增加文件，那么只有到下一批taskset开始执行时有效

*    3. 文件的位置可以是本地文件，HDFS文件或者其他hadoop支持的文件系统上，HTTP,HTTPS或者FTP URI也可以。在spark jobs中可以通过

*        SparkFiles.get(fileName)访问此文件

*    4. 如果要递归获取文件，那么可以给定一个目录，但是这种方式只对Hadoop-supported filesystems有效。

*/

def addFile(path: String, recursive: Boolean): Unit = {

val uri = new Path(path).toUri

val schemeCorrectedPath = uri.getScheme match {

    //如果路径中不指定schema，也就是null.

    //在命令行指定--files 时候，--files /home/kong/log4j.properties等同于--files local:/home/kong/log4j.properties

  case null | "local" => new File(path).getCanonicalFile.toURI.toString

  case _ => path

}

val hadoopPath = new Path(schemeCorrectedPath)

val scheme = new URI(schemeCorrectedPath).getScheme

if (!Array("http", "https", "ftp").contains(scheme)) {

  val fs = hadoopPath.getFileSystem(hadoopConfiguration)

  val isDir = fs.getFileStatus(hadoopPath).isDirectory

  if (!isLocal && scheme == "file" && isDir) {

    throw new SparkException(s"addFile does not support local directories when not running " +

      "local mode.")

  }

  if (!recursive && isDir) {

    throw new SparkException(s"Added file $hadoopPath is a directory and recursive is not " +

      "turned on.")

  }

} else {

  // SPARK-17650: Make sure this is a valid URL before adding it to the list of dependencies

  Utils.validateURL(uri)

}

val key = if (!isLocal && scheme == "file") {

  env.rpcEnv.fileServer.addFile(new File(uri.getPath))

} else {

  schemeCorrectedPath

}

val timestamp = System.currentTimeMillis

if (addedFiles.putIfAbsent(key, timestamp).isEmpty) {

  logInfo(s"Added file $path at $key with timestamp $timestamp")

  // Fetch the file locally so that closures which are run on the driver can still use the

  // SparkFiles API to access files.

  Utils.fetchFile(uri.toString, new File(SparkFiles.getRootDirectory()), conf,

    env.securityManager, hadoopConfiguration, timestamp, useCache = false)

  postEnvironmentUpdate()

}

}

在addJar和addFile方法的最后都调用了postEnvironmentUpdate方法，而且在SparkContext初始化过程的
最后也会调用postEnvironmentUpdate，代码如下：

  /** Post the environment update event once the task scheduler is ready */

  private def postEnvironmentUpdate() {

    if (taskScheduler != null) {

      val schedulingMode = getSchedulingMode.toString

      val addedJarPaths = addedJars.keys.toSeq

      val addedFilePaths = addedFiles.keys.toSeq

        // 通过调用SparkEnv的方法environmentDetails将环境的JVM参数、Spark 属性、系统属性、classPath等信息设置为环境明细信息。

      val environmentDetails = SparkEnv.environmentDetails(conf, schedulingMode, addedJarPaths,

        addedFilePaths)

        // 生成SparkListenerEnvironmentUpdate事件，并投递到事件总线

      val environmentUpdate = SparkListenerEnvironmentUpdate(environmentDetails)

      listenerBus.post(environmentUpdate)

    }

  }

environmentDetails方法：

  /**

   * Return a map representation of jvm information, Spark properties, system properties, and

   * class paths. Map keys define the category, and map values represent the corresponding

   * attributes as a sequence of KV pairs. This is used mainly for SparkListenerEnvironmentUpdate.

   */

  private[spark]

  def environmentDetails(

      conf: SparkConf,

      schedulingMode: String,

      addedJars: Seq[String],

      addedFiles: Seq[String]): Map[String, Seq[(String, String)]] = {

    import Properties._

    val jvmInformation = Seq(

      ("Java Version", s"$javaVersion ($javaVendor)"),

      ("Java Home", javaHome),

      ("Scala Version", versionString)

    ).sorted

    // Spark properties

    // This includes the scheduling mode whether or not it is configured (used by SparkUI)

    val schedulerMode =

      if (!conf.contains("spark.scheduler.mode")) {

        Seq(("spark.scheduler.mode", schedulingMode))

      } else {

        Seq.empty[(String, String)]

      }

    val sparkProperties = (conf.getAll ++ schedulerMode).sorted

    // System properties that are not java classpaths

    val systemProperties = Utils.getSystemProperties.toSeq

    val otherProperties = systemProperties.filter { case (k, _) =>

      k != "java.class.path" && !k.startsWith("spark.")

    }.sorted

    // Class paths including all added jars and files

    val classPathEntries = javaClassPath

      .split(File.pathSeparator)

      .filterNot(_.isEmpty)

      .map((_, "System Classpath"))

    val addedJarsAndFiles = (addedJars ++ addedFiles).map((_, "Added By User"))

    val classPaths = (addedJarsAndFiles ++ classPathEntries).sorted

    Map[String, Seq[(String, String)]](

      "JVM Information" -> jvmInformation,

      "Spark Properties" -> sparkProperties,

      "System Properties" -> otherProperties,

      "Classpath Entries" -> classPaths)

  }

通过 spark.files 传入spark任务依赖的文件源码分析的更多相关文章

Spark技术内幕：Stage划分及提交源码分析
http://blog.csdn.net/anzhsoft/article/details/39859463 当触发一个RDD的action后,以count为例,调用关系如下: org.apache. ...
Spark大师之路：广播变量（Broadcast）源码分析
概述最近工作上忙死了……广播变量这一块其实早就看过了,一直没有贴出来. 本文基于Spark 1.0源码分析,主要探讨广播变量的初始化.创建.读取以及清除. 类关系 BroadcastManager类 ...
65、Spark Streaming：数据接收原理剖析与源码分析
一.数据接收原理二.源码分析入口包org.apache.spark.streaming.receiver下ReceiverSupervisorImpl类的onStart()方法 ### overr ...
Spark(二)【sc.textfile的分区策略源码分析】
sparkcontext.textFile()返回的是HadoopRDD! 关于HadoopRDD的官方介绍,使用的是旧版的hadoop api ctrl+F12搜索 HadoopRDD的getPar ...
小白都能看懂的 Spring 源码揭秘之依赖注入(DI)源码分析
目录前言依赖注入的入口方法依赖注入流程分析 AbstractBeanFactory#getBean AbstractBeanFactory#doGetBean AbstractAutowireC ...
Spark技术内幕: Task向Executor提交的源码解析
在上文<Spark技术内幕:Stage划分及提交源码分析>中,我们分析了Stage的生成和提交.但是Stage的提交,只是DAGScheduler完成了对DAG的划分,生成了一个计算拓扑, ...
Spark Scheduler模块源码分析之DAGScheduler
本文主要结合Spark-1.6.0的源码,对Spark中任务调度模块的执行过程进行分析.Spark Application在遇到Action操作时才会真正的提交任务并进行计算.这时Spark会根据Ac ...
Spark源码分析之三：Stage划分
继上篇<Spark源码分析之Job的调度模型与运行反馈>之后,我们继续来看第二阶段--Stage划分. Stage划分的大体流程如下图所示: 前面提到,对于JobSubmitted事件,我 ...
spark 源码分析之十七 -- Spark磁盘存储剖析
上篇文章 spark 源码分析之十六 -- Spark内存存储剖析主要剖析了Spark 的内存存储.本篇文章主要剖析磁盘存储. 总述磁盘存储相对比较简单,相关的类关系图如下: 我们先从依赖类 Di ...

随机推荐

如何用AU3调用自己用VC++写的dll函数
这问题困扰我一个上午了,终于找到原因了,不敢藏私,和大家分享一下. 大家都知道,AU3下调用dll文件里的函数是很方便的,只要一个dllcall语句就可以了. 比如下面这个: $result = Dl ...
JAVA培训—线程同步--卖票问题
线程同步方法: (1).同步代码块,格式: synchronized (同步对象){ //同步代码 } (2).同步方法,格式: 在方法前加synchronized修饰问题: 多个人同时买票. 1. ...
HTTP关键词收集
[HTTP协议][客户端][服务器端][HTTPS][Web服务器][域名][DNS][IP地址][虚拟服务器][虚拟主机][中转服务器][HTTP/1.1规范][域名解析][Web托管服务][代理] ...
C# Show()与ShowDialog()的区别-----转载
A.WinForm中窗体显示显示窗体可以有以下2种方法: Form.ShowDialog方法 (窗体显示为模式窗体) Form.Show方法 (窗体显示为无模式窗体) 两者具体区别如下: 1 ...
JavaScript引用类型与对象
1.引用类型引用类型的值(对象)是引用类型的一个实例.引用类型有时候也被称为对象定义,因为它们描述的是一类对象所具有的属性和方法. 对象是某个特定引用类型的实例.新对象是使用new操作符后跟一个构造 ...
HTML有2种路径的写法：绝对路径和相对路径
HTML有2种路径的写法:绝对路径和相对路径 2016年11月30日 17:51:20 Bolon0708 阅读数 21775 版权声明:本文为博主原创文章,未经博主允许不得转载. https:/ ...
SIM800L AT command
/*********************************************************** AT+ICF==<format> ,<parity> ...
matlab练习程序（概率路线图PRM）
PRM概率路线图全称 Probabilistic Roadmap,是一种路径规划算法,利用随机撒点的方式将空间抽样并将问题转为图搜索,利用A*或Dijkstra算法找到起始结束节点的最短路径. 可以想 ...
php-计算2个时间之差
//$startdate是开始时间,$enddate是结束时间 <?php $startdate="2011-3-15 11:50:00"; $enddate="2 ...
如何使用Nexus搭建Maven私服
如何使用Nexus搭建Maven私服听语音 | 浏览:47 | 更新:2016-09-29 10:22 1 2 3 4 5 6 7 分步阅读一键约师傅百度师傅最快的到家服务,最优质的电脑清灰! ...

通过 spark.files 传入spark任务依赖的文件源码分析

版本：spak2.3

通过 spark.files 传入spark任务依赖的文件源码分析的更多相关文章

随机推荐

热门专题