本文目的

在介绍estimator分布式的时候，官方文档由于版本更新导致与接口不一致。具体是：在estimator分布式当中，使用dataset作为数据输入，在1.12版本中，数据训练只是dataset的数据，就是所有设备加起来，跑一遍数据。

而在2.0版本中，训练数据是dataset的数据乘以分

布式的设备数。也就是说，在每个设备当中都会完整地跑一遍dataset的所有数据。

1.12版本读取

1. 在主线程当中创建图

下面这段代码中，在client中调用了input function，得到迭代器。这是属于estimator distribute train调用的代码

with ops.Graph().as_default() as g:

      # We want to create the iterations variable outside the distribution scope

      # as that is just stored on the host and mainly used to drive the loop

      # and doesn't need to be a Mirrored/Device variable.

      if is_tpu_strategy:

        steps_per_run_variable = training.get_or_create_steps_per_run_variable()

      with self._train_distribution.scope():

        random_seed.set_random_seed(self._config.tf_random_seed)

        iterator, input_hooks = self._get_iterator_from_input_fn(

            input_fn, model_fn_lib.ModeKeys.TRAIN, self._train_distribution)

_get_iterator_from_input_fn * 这个函数会生成迭代器供后续训练读取数据。

  def _get_iterator_from_input_fn(self, input_fn, mode, distribution=None):

    if distribution is not None:

      result = distribution.distribute_dataset(

          lambda: self._call_input_fn(input_fn, mode))

    else:

      result = self._call_input_fn(input_fn, mode)

    iterator = result.make_initializable_iterator()

    input_hooks = [estimator_util._DatasetInitializerHook(iterator)]  # pylint: disable=protected-access

    return iterator, input_hooks

这里会调用distribute_dataset生成dataset。

再点进去看以后可看到会创建这样一个PerDeviceDataset

class PerDeviceDataset(object):

  """Like `tf.data.Dataset` split devices, producing `PerDevice` data."""

  def __init__(self, dataset, devices, prefetch_on_device=None):

    self._devices = devices

    # Default to using prefetching in graph mode, unless specified.

    # TODO(priyag): Enable prefetching in eager mode.

    self._prefetch_on_device = prefetch_on_device

    if self._prefetch_on_device is None:

      self._prefetch_on_device = not context.executing_eagerly()

    assert not (self._prefetch_on_device and context.executing_eagerly()), (

        "Prefetching is only supported in graph mode currently")

    if self._prefetch_on_device:

      self._dataset = dataset.apply(

          prefetching_ops_v2.prefetch_to_devices(self._devices))

    else:

      # TODO(priyag): If dropping remainder is not appropriate, find another

      # approach to distributing the dataset when not possible to divide evenly.

      # Possibly not an issue when we start using PartitionedDataset.

      self._dataset = dataset.batch(len(devices), drop_remainder=True)

最后一行代码可以看到，在原dataset上又封装了一层batch。将数据根据设备数切分。

后面创建迭代器也是封装为PerDeviceDataIterator，形成一个字典映射，不同设备不同数据，根据batch 的index切分。

分布式训练

在1.12版本中的训练比较简单。对于MirroredStrategy来说，会给每个一个device创建一个线程，

有一个缺点就是，每一次run都会创建线程，在todo里看到，后续会优化掉应该。

下面是在client中从迭代器获取数据，传递给每个device去运算的代码，

self._train_distribution.call_for_each_tower

features, labels = estimator_util.parse_iterator_result(

              iterator.get_next())

          grouped_estimator_spec = self._train_distribution.call_for_each_tower(

              self._call_model_fn,

              features,

              labels,  # although this will be None it seems

              model_fn_lib.ModeKeys.TRAIN,

              self.config)

          loss = self._train_distribution.unwrap(

              self._train_distribution.reduce(

                  distribute_lib.get_loss_reduction(),

                  grouped_estimator_spec.loss,

                  destinations='/device:CPU:0'))[0]

          distributed_train_op = grouped_estimator_spec.train_op

call_for_each_tower是每个设备训练的接口

def _call_for_each_tower(distribution, fn, *args, **kwargs):

  """Run `fn` in separate threads, once per tower/worker device.

  run_concurrently = kwargs.pop("run_concurrently", True)

  if not context.executing_eagerly():

    # Lots of TF library code isn't thread-safe in graph mode, and

    # there is little to be gained by turning on multithreading when

    # constructing a graph.

    run_concurrently = False

    # Needed for per-thread device, etc. contexts in graph mode.

    ops.get_default_graph().switch_to_thread_local()

  elif run_concurrently is None:

    run_concurrently = True

  coord = coordinator.Coordinator(clean_stop_exception_types=(_RequestedStop,))

  shared_variable_store = {}

  # TODO(isaprykin): Create these threads once instead of during every run()

  # call.

  threads = []

  for index, d in enumerate(distribution.worker_devices):

    variable_creator_fn = shared_variable_creator.make_fn(

        shared_variable_store, index)

    t = MirroredStrategy._MirroredTowerThread(  # pylint: disable=protected-access

        distribution, coord, d, variable_creator_fn, fn,

        *values.select_device(d, args), **values.select_device(d, kwargs))

    threads.append(t)

  for t in threads:

    t.start()

其中，select_device就是取对应设备key对应的值。完成整个分布式训练。

TensorFlow Distribution(分布式中的数据读取和训练)的更多相关文章

DataTable to Excel（使用NPOI、EPPlus将数据表中的数据读取到excel格式内存中）
/// <summary> /// DataTable to Excel(将数据表中的数据读取到excel格式内存中) /// </summary> /// <param ...
TensorFlow走过的坑之---数据读取和tf中batch的使用方法
首先介绍数据读取问题,现在TensorFlow官方推荐的数据读取方法是使用tf.data.Dataset,具体的细节不在这里赘述,看官方文档更清楚,这里主要记录一下官方文档没有提到的坑,以示" ...
oracle中的数据读取与查找
数据读取首先数据块读入到Buffer Cache中,并将其放在LRU(Last Recently Used)链表的MRU(Most Recently Used)端,当需要再次访问该块时可以直接从bu ...
c#中使用数据读取器读取查询结果
今天有时间了. 在看<c#数据库入门经典> ,总结数据读取器查询结果. 针对单个结果集使用读取器,有3中方法: String connString =..; String sql =@&q ...
如何在ADO中使用数据读取器（DataReader）读取数据
DbDataReader类型(实现IDataReader接口)是从数据源获取信息最简单也最快速的方法. 数据读取器是只读向前的效据流．井且一次返回一条记录.因此．只有当你向数据源提交 Select 查 ...
《TensorFlow实战》中AlexNet卷积神经网络的训练中
TensorFlow实战中AlexNet卷积神经网络的训练 01 出错 TypeError: as_default() missing 1 required positional argument: ...
Android中Json数据读取与创建
一: Json的特性和在数据交互中的地位就不用说了,直接看案例. 首先在android studio中创建assets文件目录,用于存放Json数据文件,android studio 1.3 默认项 ...
Android中Json数据读取与创建的方法
转自:http://www.jb51.net/article/70875.htm 首先介绍下JSON的定义,JSON是JavaScript Object Notation的缩写. 一种轻量级的数据交换 ...
TensorFlow实践笔记（一）：数据读取
本文整理了TensorFlow中的数据读取方法,在TensorFlow中主要有三种方法读取数据: Feeding:由Python提供数据. Preloaded data:预加载数据. Reading ...

随机推荐

Linux学习笔记04
文件查找命令find 文件查找命令: which locate find which:查找命令字所在的位置 locate:模糊匹配(只要包含关键字的文件都查找出来) 不是实时的,基于数据库查找, up ...
【iOS】receiver type *** for instance message is a forward declaration
错误原因:没有引入相关的头文件 http://stackoverflow.com/questions/8815200/receiver-type-for-instance-message-is-a-f ...
MySQL中一些关于索引的知识点
什么是索引索引是一种数据结构,其作用就是用来提高数据查询效率.比较常用的比喻就是将其类比为书籍的目录.通过目录可以精确的找到某一章节的内容所在页. 在数据量较小的时候使用索引其实也没有什么意义,即使 ...
详解 Diff 算法以及循环要加 key 值问题
上一篇文章我简述了什么是 Virtual DOM,这一章我会详细讲 Diff 算法以及为什么在 React 和 Vue 中循环都需要 key 值. 什么是 DOM Diff 算法 Web 界面其实就是 ...
maven 下载安装环境配置
电脑系统:win10 64位 idea 2019 Java 1.8 1.链接地址,我一般都找官网 http://maven.apache.org/download.cgi 截图:注意mav ...
js中数组和对象的合并
1 数组合并 1.1 concat 方法 1 2 3 4 var a=[1,2,3],b=[4,5,6]; var c=a.concat(b); console.log(c);// 1,2,3,4,5 ...
java多线程基础（二）--sleep(),wait,()yield()和join()方法
1.sleep()方法在指定时间内让当前正在执行的线程暂停执行,但不会释放“锁标志”.不推荐使用. sleep()使当前线程进入阻塞状态,在指定时间内不会执行. 2.wait()方法在其他线程调用 ...
素数筛法（Eratosthenes筛法）
介绍 Eratosthenes筛法,又名埃氏筛法,对于求1~n区间内的素数,时间复杂度为n log n,对于10^6^ 以内的数比较合适,再超出此范围的就不建议用该方法了. 筛法的思想特别简单: 对于 ...
Sqlmap过waf命令tamper各脚本的适用环境
0x00 相信很多小伙伴和我一样感同身受,站上明明有注入可是被万恶的WAF拦截了或者过滤了,这时候就需要用到SQLMAP强大的tamper了. 0x01 使用方法--tamper xxx.py apo ...
vscode 支持 threejs 的智能提示
VSCode Typings and Intellisense: Dummy Learning VS-Code 1 Jun 20, 2016 Updated on Jun 20 2016 for 1. ...