1. 经常使用类

class tf.contrib.rnn.BasicLSTMCell

BasicLSTMCell 是最简单的一个LSTM类。没有实现clipping，projection layer。peep-hole等一些LSTM的高级变种，仅作为一个主要的basicline结构存在，假设要使用这些高级变种，需用class tf.contrib.rnn.LSTMCell这个类。

使用方式：

lstm = rnn.BasicLSTMCell(lstm_size, forget_bias=1.0, state_is_tuple=True)

Args:

num_units: int, The number of units in the LSTM cell.

forget_bias: float, The bias added to forget gates.

state_is_tuple: If True, accepted and returned states are 2-tuples of the c_state and m_state. If False, they are concatenated along the column axis. The latter behavior will soon be deprecated.

activation: Activation function of the inner states.

说明：

num_units 是指一个Cell中神经元的个数，并非循环层的Cell个数。

这里有人可能会疑问：循环层的Cell数目怎么表示？答案是通过例如以下代码中的 time_step_size确定（X_split 中划分出的arrays数量为循环层的Cell个数）：

    X_split = tf.split(XR, time_step_size, 0)

在随意时刻 t ，LSTM Cell会产生两个内部状态 ct和ht （关于RNN与LSTM的介绍可參考：循环神经网络与LSTM）。当state_is_tuple=True时，上面讲到的状态ct和ht 就是分开记录，放在一个二元tuple中返回，假设这个參数没有设定或设置成False，两个状态就按列连接起来返回。官方说这样的形式立即就要被deprecated了，全部我们在使用LSTM的时候要加上state_is_tuple=True。

class tf.contrib.rnn.DropoutWrapper

RNN中的dropout和cnn不同，在RNN中。时间序列方向不进行dropout，也就是说从t-1时刻的状态传递到t时刻进行计算时，这个中间不进行memory的dropout。例如以下图所看到的，Dropout仅应用于虚线方向的输入，即仅针对于上一层的输出做Dropout。

因此。我们在代码中定义完Cell之后，在Cell外部包裹上dropout，这个类叫DropoutWrapper，这样我们的Cell就有了dropout功能！

lstm = tf.nn.rnn_cell.DropoutWrapper(lstm, output_keep_prob=keep_prob)

Args:

cell: an RNNCell, a projection to output_size is added to it.

input_keep_prob: unit Tensor or float between 0 and 1, input keep probability; if it is float and 1, no input dropout will be added.

output_keep_prob: unit Tensor or float between 0 and 1, output keep probability; if it is float and 1, no output dropout will be added.

seed: (optional) integer, the randomness seed.

class tf.contrib.rnn.MultiRNNCell

假设希望整个网络的层数很多其它（比如上图表示一个两层的RNN，第一层Cell的output还要作为下一层Cell的输入），应该堆叠多个LSTM Cell，tensorflow给我们提供了MultiRNNCell，因此堆叠多层网络仅仅生成这个类就可以：

lstm = tf.nn.rnn_cell.MultiRNNCell([lstm] * num_layers, state_is_tuple=True)

2. 代码

MNIST数据集的格式与数据预处理代码 input_data.py的解说请參考 :Tutorial (2)

# -*- coding: utf-8 -*-

import tensorflow as tf

from tensorflow.contrib import rnn

import numpy as np

import input_data

# configuration

#                        O * W + b -> 10 labels for each image, O[? 28], W[28 10], B[10]

#                       ^ (O: output 28 vec from 28 vec input)

#                       |

#      +-+  +-+       +--+

#      |1|->|2|-> ... |28| time_step_size = 28

#      +-+  +-+       +--+

#       ^    ^    ...  ^

#       |    |         |

# img1:[28] [28]  ... [28]

# img2:[28] [28]  ... [28]

# img3:[28] [28]  ... [28]

# ...

# img128 or img256 (batch_size or test_size 256)

#      each input size = input_vec_size=lstm_size=28

# configuration variables

input_vec_size = lstm_size = 28 # 输入向量的维度

time_step_size = 28 # 循环层长度

batch_size = 128

test_size = 256

def init_weights(shape):

    return tf.Variable(tf.random_normal(shape, stddev=0.01))

def model(X, W, B, lstm_size):

    # X, input shape: (batch_size, time_step_size, input_vec_size)

    # XT shape: (time_step_size, batch_size, input_vec_size)

    XT = tf.transpose(X, [1, 0, 2])  # permute time_step_size and batch_size,[28, 128, 28]

    # XR shape: (time_step_size * batch_size, input_vec_size)

    XR = tf.reshape(XT, [-1, lstm_size]) # each row has input for each lstm cell (lstm_size=input_vec_size)

    # Each array shape: (batch_size, input_vec_size)

    X_split = tf.split(XR, time_step_size, 0) # split them to time_step_size (28 arrays),shape = [(128, 28),(128, 28)...]

    # Make lstm with lstm_size (each input vector size). num_units=lstm_size; forget_bias=1.0

    lstm = rnn.BasicLSTMCell(lstm_size, forget_bias=1.0, state_is_tuple=True)

    # Get lstm cell output, time_step_size (28) arrays with lstm_size output: (batch_size, lstm_size)

    # rnn..static_rnn()的输出相应于每个timestep。假设仅仅关心最后一步的输出，取outputs[-1]就可以

    outputs, _states = rnn.static_rnn(lstm, X_split, dtype=tf.float32)  # 时间序列上每个Cell的输出:[... shape=(128, 28)..]

    # Linear activation

    # Get the last output

    return tf.matmul(outputs[-1], W) + B, lstm.state_size # State size to initialize the stat

mnist = input_data.read_data_sets("MNIST_data/", one_hot=True) # 读取数据

# mnist.train.images是一个55000 * 784维的矩阵, mnist.train.labels是一个55000 * 10维的矩阵

trX, trY, teX, teY = mnist.train.images, mnist.train.labels, mnist.test.images, mnist.test.labels

# 将每张图用一个28x28的矩阵表示,(55000,28,28,1)

trX = trX.reshape(-1, 28, 28)

teX = teX.reshape(-1, 28, 28) 

X = tf.placeholder("float", [None, 28, 28])

Y = tf.placeholder("float", [None, 10])

# get lstm_size and output 10 labels

W = init_weights([lstm_size, 10])  # 输出层权重矩阵28×10

B = init_weights([10])  # 输出层bais

py_x, state_size = model(X, W, B, lstm_size)

cost = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=py_x, labels=Y))

train_op = tf.train.RMSPropOptimizer(0.001, 0.9).minimize(cost)

predict_op = tf.argmax(py_x, 1)

session_conf = tf.ConfigProto()

session_conf.gpu_options.allow_growth = True

# Launch the graph in a session

with tf.Session(config=session_conf) as sess:

    # you need to initialize all variables

    tf.global_variables_initializer().run()

    for i in range(100):

        for start, end in zip(range(0, len(trX), batch_size), range(batch_size, len(trX)+1, batch_size)):

            sess.run(train_op, feed_dict={X: trX[start:end], Y: trY[start:end]})

        test_indices = np.arange(len(teX))  # Get A Test Batch

        np.random.shuffle(test_indices)

        test_indices = test_indices[0:test_size]

        print(i, np.mean(np.argmax(teY[test_indices], axis=1) ==

                         sess.run(predict_op, feed_dict={X: teX[test_indices]})))

watermark/2/text/aHR0cDovL2Jsb2cuY3Nkbi5uZXQvdTAxMDA4OTQ0NA==/font/5a6L5L2T/fontsize/400/fill/I0JBQkFCMA==/dissolve/70/gravity/SouthEast" width="500" height="300" alt="图 1">

3. 參考资料

Tensorflow - Tutorial (7) : 利用 RNN/LSTM 进行手写数字识别的更多相关文章

Tensorflow笔记——神经网络图像识别（五）手写数字识别
TensorFlow------单层(全连接层)实现手写数字识别训练及测试实例
TensorFlow之单层(全连接层)实现手写数字识别训练及测试实例: import tensorflow as tf from tensorflow.examples.tutorials.mnist ...
5 TensorFlow入门笔记之RNN实现手写数字识别
------------------------------------ 写在开头:此文参照莫烦python教程(墙裂推荐!!!) ---------------------------------- ...
keras和tensorflow搭建DNN、CNN、RNN手写数字识别
MNIST手写数字集 MNIST是一个由美国由美国邮政系统开发的手写数字识别数据集.手写内容是0~9,一共有60000个图片样本,我们可以到MNIST官网免费下载,总共4个.gz后缀的压缩文件,该文件 ...
TensorFlow使用RNN实现手写数字识别
学习,笔记,有时间会加注释以及函数之间的逻辑关系. # https://www.cnblogs.com/felixwang2/p/9190664.html # https://www.cnblogs. ...
TensorFlow下利用MNIST训练模型识别手写数字
本文将参考TensorFlow中文社区官方文档使用mnist数据集训练一个多层卷积神经网络(LeNet5网络),并利用所训练的模型识别自己手写数字. 训练MNIST数据集,并保存训练模型 # Pyth ...
Android+TensorFlow+CNN+MNIST 手写数字识别实现
Android+TensorFlow+CNN+MNIST 手写数字识别实现 SkySeraph 2018 Email:skyseraph00#163.com 更多精彩请直接访问SkySeraph个人站 ...
利用神经网络算法的C＃手写数字识别
欢迎大家前往云+社区,获取更多腾讯海量技术实践干货哦~ 下载Demo - 2.77 MB (原始地址):handwritten_character_recognition.zip 下载源码 - 70. ...
手写数字识别 ----卷积神经网络模型官方案例注释（基于Tensorflow,Python）
# 手写数字识别 ----卷积神经网络模型 import os import tensorflow as tf #部分注释来源于 # http://www.cnblogs.com/rgvb178/p/ ...

随机推荐

appium+python自动化51-adb文件导入和导出（pull push）
前言用手机连电脑的时候,有时候需要把手机(模拟器)上的文件导出到电脑上,或者把电脑的图片导入手机里做测试用,我们可以用第三方的软件管理工具直接复制粘贴,也可以直接通过adb命令导入和导出. adb ...
解决kylin报错 ClassCastException org.apache.hadoop.hive.ql.exec.ConditionalTask cannot be cast to org.apache.hadoop.hive.ql.exec.mr.MapRedTask
方法:去掉参数SET hive.auto.convert.join=true; 从配置文件$KYLIN_HOME/conf/kylin_hive_conf.xml删掉或 kylin-gui的cube ...
非docker的jenkins的master如何使用docker的jenkins的slave
前提 1.存在jenkins的master,这个master不是docker的,是通过yum install jenkins安装的 2.使用docker创建n个jenkins,方法是docker pu ...
RenderMonkey 练习第四天【OpenGL Texture Bump】
BumpTexture 1. 新建一个OpenGL 空effect; 2. 添加相关变量右击Effect节点选择Add Variable->float->float / float3 添 ...
Dx12 occlusion query
https://github.com/Microsoft/DirectX-Graphics-Samples/blob/master/Samples/Desktop/D3D12PredicationQu ...
http://www.cnblogs.com/jqyp/archive/2010/08/20/1805041.html
http://www.cnblogs.com/jqyp/archive/2010/08/20/1805041.html
[Python爬虫] 之十三：Selenium +phantomjs抓取活动树会议活动数据
抓取活动树网站中会议活动数据(http://www.huodongshu.com/html/index.html) 具体的思路是[Python爬虫] 之十一中抓取活动行网站的类似,都是用多线程来抓取, ...
wampserver 下载链接没反应的解决办法
可能有很多小伙伴和我一样使用wampserver时,下载链接点击就是没有反应,当时我以为是因为网络原因,链接没有加载出来,或者是链接的请求不能得到响应,结果百度了一下才发现被“习惯”坑了一把,wamp ...
C#-遍历datatable的几种方法
遍历datatable的方法2009-- :02方法一: DataTable dt = dataSet.Tables[]; ; i < dt.Rows.Count ; i++) { string ...
hdu5293(2015多校1)--Tree chain problem（树状dp）
Tree chain problem Time Limit: 6000/3000 MS (Java/Others) Memory Limit: 65536/65536 K (Java/Other ...

Tensorflow - Tutorial (7) : 利用 RNN/LSTM 进行手写数字识别