(zhuan) Deep Reinforcement Learning Papers
Deep Reinforcement Learning Papers
A list of recent papers regarding deep reinforcement learning.
The papers are organized based on manually-defined bookmarks.
They are sorted by time to see the recent papers first.
Any suggestions and pull requests are welcome.
Bookmarks
- All Papers
- Value
- Policy
- Discrete Control
- Continuous Control
- Text Domain
- Visual Domain
- Robotics
- Games
- Monte-Carlo Tree Search
- Inverse Reinforcement Learning
- Improving Exploration
- Multi-Task and Transfer Learning
- Multi-Agent
- Hierarchical Learning
All Papers
- Model-Free Episodic Control, C. Blundell et al., arXiv, 2016.
- Safe and Efficient Off-Policy Reinforcement Learning, R. Munos et al., arXiv, 2016.
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Curiosity-driven Exploration in Deep Reinforcement Learning via Bayesian Neural Networks, R. Houthooft et al., arXiv, 2016.
- Control of Memory, Active Perception, and Action in Minecraft, J. Oh et al., ICML, 2016.
- Dynamic Frame skip Deep Q Network, A. S. Lakshminarayanan et al., IJCAI Deep RL Workshop, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Benchmarking Deep Reinforcement Learning for Continuous Control, Y. Duan et al., ICML, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection, S. Levine et al., arXiv, 2016.
- Continuous Deep Q-Learning with Model-based Acceleration, S. Gu et al., ICML, 2016.
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, C. Finn et al., arXiv, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Value Iteration Networks, A. Tamar et al., arXiv, 2016.
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks, J. N. Foerster et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Increasing the Action Gap: New Operators for Reinforcement Learning, M. G. Bellemare et al., AAAI, 2016.
- Memory-based control with recurrent neural networks, N. Heess et al., NIPS Workshop, 2015.
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies, V. François-Lavet et al., NIPS Workshop, 2015.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
- Strategic Dialogue Management via Deep Reinforcement Learning, H. Cuayáhuitl et al., NIPS Workshop, 2015.
- MazeBase: A Sandbox for Learning from Games, S. Sukhbaatar et al., arXiv, 2016.
- Learning Simple Algorithms from Examples, W. Zaremba et al., arXiv, 2015.
- Dueling Network Architectures for Deep Reinforcement Learning, Z. Wang et al., arXiv, 2015.
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning, E. Parisotto, et al., ICLR, 2016.
- Better Computer Go Player with Neural Network and Long-term Prediction, Y. Tian et al., ICLR, 2016.
- Policy Distillation, A. A. Rusu et at., ICLR, 2016.
- Prioritized Experience Replay, T. Schaul et al., ICLR, 2016.
- Deep Reinforcement Learning with an Action Space Defined by Natural Language, J. He et al., arXiv, 2015.
- Deep Reinforcement Learning in Parameterized Action Space, M. Hausknecht et al., ICLR, 2016.
- Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control, F. Zhang et al., arXiv, 2015.
- Generating Text with Deep Reinforcement Learning, H. Guo, arXiv, 2015.
- ADAAPT: A Deep Architecture for Adaptive Policy Transfer from Multiple Sources, J. Rajendran et al., arXiv, 2015.
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning, S. Mohamed and D. J. Rezende, arXiv, 2015.
- Deep Reinforcement Learning with Double Q-learning, H. van Hasselt et al., arXiv, 2015.
- Recurrent Reinforcement Learning: A Hybrid Approach, X. Li et al., arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Language Understanding for Text-based Games Using Deep Reinforcement Learning, K. Narasimhan et al., EMNLP, 2015.
- Giraffe: Using Deep Reinforcement Learning to Play Chess, M. Lai, arXiv, 2015.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Learning Continuous Control Policies by Stochastic Value Gradients, N. Heess et al., NIPS, 2015.
- Learning Deep Neural Network Policies with Continuous Memory States, M. Zhang et al., arXiv, 2015.
- Deep Recurrent Q-Learning for Partially Observable MDPs, M. Hausknecht and P. Stone, arXiv, 2015.
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences, H. Mei et al., arXiv, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
- Maximum Entropy Deep Inverse Reinforcement Learning, M. Wulfmeier et al., arXiv, 2015.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation, J. Schulman et al., ICLR, 2016.
- End-to-End Training of Deep Visuomotor Policies, S. Levine et al., arXiv, 2015.
- DeepMPC: Learning Deep Latent Features for Model Predictive Control, I. Lenz, et al., RSS, 2015.
- Universal Value Function Approximators, T. Schaul et al., ICML, 2015.
- Deterministic Policy Gradient Algorithms, D. Silver et al., ICML, 2015.
- Massively Parallel Methods for Deep Reinforcement Learning, A. Nair et al., ICML Workshop, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
- Human-level control through deep reinforcement learning, V. Mnih et al., Nature, 2015.
- Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning, X. Guo et al., NIPS, 2014.
- Playing Atari with Deep Reinforcement Learning, V. Mnih et al., NIPS Workshop, 2013.
Value
- Model-Free Episodic Control, C. Blundell et al., arXiv, 2016.
- Safe and Efficient Off-Policy Reinforcement Learning, R. Munos et al., arXiv, 2016.
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Control of Memory, Active Perception, and Action in Minecraft, J. Oh et al., ICML, 2016.
- Dynamic Frame skip Deep Q Network, A. S. Lakshminarayanan et al., IJCAI Deep RL Workshop, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Continuous Deep Q-Learning with Model-based Acceleration, S. Gu et al., ICML, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Value Iteration Networks, A. Tamar et al., arXiv, 2016.
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks, J. N. Foerster et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Increasing the Action Gap: New Operators for Reinforcement Learning, M. G. Bellemare et al., AAAI, 2016.
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies, V. François-Lavet et al., NIPS Workshop, 2015.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
- Strategic Dialogue Management via Deep Reinforcement Learning, H. Cuayáhuitl et al., NIPS Workshop, 2015.
- Learning Simple Algorithms from Examples, W. Zaremba et al., arXiv, 2015.
- Dueling Network Architectures for Deep Reinforcement Learning, Z. Wang et al., arXiv, 2015.
- Prioritized Experience Replay, T. Schaul et al., ICLR, 2016.
- Deep Reinforcement Learning with an Action Space Defined by Natural Language, J. He et al., arXiv, 2015.
- Deep Reinforcement Learning in Parameterized Action Space, M. Hausknecht et al., ICLR, 2016.
- Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control, F. Zhang et al., arXiv, 2015.
- Generating Text with Deep Reinforcement Learning, H. Guo, arXiv, 2015.
- Deep Reinforcement Learning with Double Q-learning, H. van Hasselt et al., arXiv, 2015.
- Recurrent Reinforcement Learning: A Hybrid Approach, X. Li et al., arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Language Understanding for Text-based Games Using Deep Reinforcement Learning, K. Narasimhan et al., EMNLP, 2015.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Deep Recurrent Q-Learning for Partially Observable MDPs, M. Hausknecht and P. Stone, arXiv, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
- Massively Parallel Methods for Deep Reinforcement Learning, A. Nair et al., ICML Workshop, 2015.
- Human-level control through deep reinforcement learning, V. Mnih et al., Nature, 2015.
- Playing Atari with Deep Reinforcement Learning, V. Mnih et al., NIPS Workshop, 2013.
Policy
- Curiosity-driven Exploration in Deep Reinforcement Learning via Bayesian Neural Networks, R. Houthooft et al., arXiv, 2016.
- Benchmarking Deep Reinforcement Learning for Continuous Control, Y. Duan et al., ICML, 2016.
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection, S. Levine et al., arXiv, 2016.
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, C. Finn et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Memory-based control with recurrent neural networks, N. Heess et al., NIPS Workshop, 2015.
- MazeBase: A Sandbox for Learning from Games, S. Sukhbaatar et al., arXiv, 2016.
- ADAAPT: A Deep Architecture for Adaptive Policy Transfer from Multiple Sources, J. Rajendran et al., arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Learning Continuous Control Policies by Stochastic Value Gradients, N. Heess et al., NIPS, 2015.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation, J. Schulman et al., ICLR, 2016.
- End-to-End Training of Deep Visuomotor Policies, S. Levine et al., arXiv, 2015.
- Deterministic Policy Gradient Algorithms, D. Silver et al., ICML, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
Discrete Control
- Model-Free Episodic Control, C. Blundell et al., arXiv, 2016.
- Safe and Efficient Off-Policy Reinforcement Learning, R. Munos et al., arXiv, 2016.
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Control of Memory, Active Perception, and Action in Minecraft, J. Oh et al., ICML, 2016.
- Dynamic Frame skip Deep Q Network, A. S. Lakshminarayanan et al., IJCAI Deep RL Workshop, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Value Iteration Networks, A. Tamar et al., arXiv, 2016.
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks, J. N. Foerster et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Increasing the Action Gap: New Operators for Reinforcement Learning, M. G. Bellemare et al., AAAI, 2016.
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies, V. François-Lavet et al., NIPS Workshop, 2015.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
- Strategic Dialogue Management via Deep Reinforcement Learning, H. Cuayáhuitl et al., NIPS Workshop, 2015.
- Learning Simple Algorithms from Examples, W. Zaremba et al., arXiv, 2015.
- Dueling Network Architectures for Deep Reinforcement Learning, Z. Wang et al., arXiv, 2015.
- Better Computer Go Player with Neural Network and Long-term Prediction, Y. Tian et al., ICLR, 2016.
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning, E. Parisotto, et al., ICLR, 2016.
- Policy Distillation, A. A. Rusu et at., ICLR, 2016.
- Prioritized Experience Replay, T. Schaul et al., ICLR, 2016.
- Deep Reinforcement Learning with an Action Space Defined by Natural Language, J. He et al., arXiv, 2015.
- Deep Reinforcement Learning in Parameterized Action Space, M. Hausknecht et al., ICLR, 2016.
- Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control, F. Zhang et al., arXiv, 2015.
- Generating Text with Deep Reinforcement Learning, H. Guo, arXiv, 2015.
- ADAAPT: A Deep Architecture for Adaptive Policy Transfer from Multiple Sources, J. Rajendran et al., arXiv, 2015.
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning, S. Mohamed and D. J. Rezende, arXiv, 2015.
- Deep Reinforcement Learning with Double Q-learning, H. van Hasselt et al., arXiv, 2015.
- Recurrent Reinforcement Learning: A Hybrid Approach, X. Li et al., arXiv, 2015.
- Language Understanding for Text-based Games Using Deep Reinforcement Learning, K. Narasimhan et al., EMNLP, 2015.
- Giraffe: Using Deep Reinforcement Learning to Play Chess, M. Lai, arXiv, 2015.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Deep Recurrent Q-Learning for Partially Observable MDPs, M. Hausknecht and P. Stone, arXiv, 2015.
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences, H. Mei et al., arXiv, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
- Universal Value Function Approximators, T. Schaul et al., ICML, 2015.
- Massively Parallel Methods for Deep Reinforcement Learning, A. Nair et al., ICML Workshop, 2015.
- Human-level control through deep reinforcement learning, V. Mnih et al., Nature, 2015.
- Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning, X. Guo et al., NIPS, 2014.
- Playing Atari with Deep Reinforcement Learning, V. Mnih et al., NIPS Workshop, 2013.
Continuous Control
- Curiosity-driven Exploration in Deep Reinforcement Learning via Bayesian Neural Networks, R. Houthooft et al., arXiv, 2016.
- Benchmarking Deep Reinforcement Learning for Continuous Control, Y. Duan et al., ICML, 2016.
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection, S. Levine et al., arXiv, 2016.
- Continuous Deep Q-Learning with Model-based Acceleration, S. Gu et al., ICML, 2016.
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, C. Finn et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Memory-based control with recurrent neural networks, N. Heess et al., NIPS Workshop, 2015.
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning, S. Mohamed and D. J. Rezende, arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Learning Continuous Control Policies by Stochastic Value Gradients, N. Heess et al., NIPS, 2015.
- Learning Deep Neural Network Policies with Continuous Memory States, M. Zhang et al., arXiv, 2015.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation, J. Schulman et al., ICLR, 2016.
- End-to-End Training of Deep Visuomotor Policies, S. Levine et al., arXiv, 2015.
- DeepMPC: Learning Deep Latent Features for Model Predictive Control, I. Lenz, et al., RSS, 2015.
- Deterministic Policy Gradient Algorithms, D. Silver et al., ICML, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
Text Domain
- Strategic Dialogue Management via Deep Reinforcement Learning, H. Cuayáhuitl et al., NIPS Workshop, 2015.
- MazeBase: A Sandbox for Learning from Games, S. Sukhbaatar et al., arXiv, 2016.
- Deep Reinforcement Learning with an Action Space Defined by Natural Language, J. He et al., arXiv, 2015.
- Generating Text with Deep Reinforcement Learning, H. Guo, arXiv, 2015.
- Language Understanding for Text-based Games Using Deep Reinforcement Learning, K. Narasimhan et al., EMNLP, 2015.
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences, H. Mei et al., arXiv, 2015.
Visual Domain
- Model-Free Episodic Control, C. Blundell et al., arXiv, 2016.
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Control of Memory, Active Perception, and Action in Minecraft, J. Oh et al., ICML, 2016.
- Dynamic Frame skip Deep Q Network, A. S. Lakshminarayanan et al., IJCAI Deep RL Workshop, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection, S. Levine et al., arXiv, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Value Iteration Networks, A. Tamar et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Increasing the Action Gap: New Operators for Reinforcement Learning, M. G. Bellemare et al., AAAI, 2016.
- Memory-based control with recurrent neural networks, N. Heess et al., NIPS Workshop, 2015.
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies, V. François-Lavet et al., NIPS Workshop, 2015.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
- Dueling Network Architectures for Deep Reinforcement Learning, Z. Wang et al., arXiv, 2015.
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning, E. Parisotto, et al., ICLR, 2016.
- Better Computer Go Player with Neural Network and Long-term Prediction, Y. Tian et al., ICLR, 2016.
- Policy Distillation, A. A. Rusu et at., ICLR, 2016.
- Prioritized Experience Replay, T. Schaul et al., ICLR, 2016.
- Deep Reinforcement Learning in Parameterized Action Space, M. Hausknecht et al., ICLR, 2016.
- Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control, F. Zhang et al., arXiv, 2015.
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning, S. Mohamed and D. J. Rezende, arXiv, 2015.
- Deep Reinforcement Learning with Double Q-learning, H. van Hasselt et al., arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Giraffe: Using Deep Reinforcement Learning to Play Chess, M. Lai, arXiv, 2015.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Learning Continuous Control Policies by Stochastic Value Gradients, N. Heess et al., NIPS, 2015.
- Deep Recurrent Q-Learning for Partially Observable MDPs, M. Hausknecht and P. Stone, arXiv, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation, J. Schulman et al., ICLR, 2016.
- End-to-End Training of Deep Visuomotor Policies, S. Levine et al., arXiv, 2015.
- Universal Value Function Approximators, T. Schaul et al., ICML, 2015.
- Massively Parallel Methods for Deep Reinforcement Learning, A. Nair et al., ICML Workshop, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
- Human-level control through deep reinforcement learning, V. Mnih et al., Nature, 2015.
- Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning, X. Guo et al., NIPS, 2014.
- Playing Atari with Deep Reinforcement Learning, V. Mnih et al., NIPS Workshop, 2013.
Robotics
- Curiosity-driven Exploration in Deep Reinforcement Learning via Bayesian Neural Networks, R. Houthooft et al., arXiv, 2016.
- Benchmarking Deep Reinforcement Learning for Continuous Control, Y. Duan et al., ICML, 2016.
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection, S. Levine et al., arXiv, 2016.
- Continuous Deep Q-Learning with Model-based Acceleration, S. Gu et al., ICML, 2016.
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, C. Finn et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Memory-based control with recurrent neural networks, N. Heess et al., NIPS Workshop, 2015.
- Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control, F. Zhang et al., arXiv, 2015.
- Learning Continuous Control Policies by Stochastic Value Gradients, N. Heess et al., NIPS, 2015.
- Learning Deep Neural Network Policies with Continuous Memory States, M. Zhang et al., arXiv, 2015.
- High-Dimensional Continuous Control Using Generalized Advantage Estimation, J. Schulman et al., ICLR, 2016.
- End-to-End Training of Deep Visuomotor Policies, S. Levine et al., arXiv, 2015.
- DeepMPC: Learning Deep Latent Features for Model Predictive Control, I. Lenz, et al., RSS, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
Games
- Model-Free Episodic Control, C. Blundell et al., arXiv, 2016.
- Safe and Efficient Off-Policy Reinforcement Learning, R. Munos et al., arXiv, 2016.
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Control of Memory, Active Perception, and Action in Minecraft, J. Oh et al., ICML, 2016.
- Dynamic Frame skip Deep Q Network, A. S. Lakshminarayanan et al., IJCAI Deep RL Workshop, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks, J. N. Foerster et al., arXiv, 2016.
- Asynchronous Methods for Deep Reinforcement Learning, V. Mnih et al., arXiv, 2016.
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Increasing the Action Gap: New Operators for Reinforcement Learning, M. G. Bellemare et al., AAAI, 2016.
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies, V. François-Lavet et al., NIPS Workshop, 2015.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
- MazeBase: A Sandbox for Learning from Games, S. Sukhbaatar et al., arXiv, 2016.
- Dueling Network Architectures for Deep Reinforcement Learning, Z. Wang et al., arXiv, 2015.
- Better Computer Go Player with Neural Network and Long-term Prediction, Y. Tian et al., ICLR, 2016.
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning, E. Parisotto, et al., ICLR, 2016.
- Policy Distillation, A. A. Rusu et at., ICLR, 2016.
- Prioritized Experience Replay, T. Schaul et al., ICLR, 2016.
- Deep Reinforcement Learning with an Action Space Defined by Natural Language, J. He et al., arXiv, 2015.
- Deep Reinforcement Learning in Parameterized Action Space, M. Hausknecht et al., ICLR, 2016.
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning, S. Mohamed and D. J. Rezende, arXiv, 2015.
- Deep Reinforcement Learning with Double Q-learning, H. van Hasselt et al., arXiv, 2015.
- Continuous control with deep reinforcement learning, T. P. Lillicrap et al., ICLR, 2016.
- Language Understanding for Text-based Games Using Deep Reinforcement Learning, K. Narasimhan et al., EMNLP, 2015.
- Giraffe: Using Deep Reinforcement Learning to Play Chess, M. Lai, arXiv, 2015.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Deep Recurrent Q-Learning for Partially Observable MDPs, M. Hausknecht and P. Stone, arXiv, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
- Universal Value Function Approximators, T. Schaul et al., ICML, 2015.
- Massively Parallel Methods for Deep Reinforcement Learning, A. Nair et al., ICML Workshop, 2015.
- Trust Region Policy Optimization, J. Schulman et al., ICML, 2015.
- Human-level control through deep reinforcement learning, V. Mnih et al., Nature, 2015.
- Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning, X. Guo et al., NIPS, 2014.
- Playing Atari with Deep Reinforcement Learning, V. Mnih et al., NIPS Workshop, 2013.
Monte-Carlo Tree Search
- Mastering the game of Go with deep neural networks and tree search, D. Silver et al., Nature, 2016.
- Better Computer Go Player with Neural Network and Long-term Prediction, Y. Tian et al., ICLR, 2016.
- Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning, X. Guo et al., NIPS, 2014.
Inverse Reinforcement Learning
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, C. Finn et al., arXiv, 2016.
- Maximum Entropy Deep Inverse Reinforcement Learning, M. Wulfmeier et al., arXiv, 2015.
Multi-Task and Transfer Learning
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning, E. Parisotto, et al., ICLR, 2016.
- Policy Distillation, A. A. Rusu et at., ICLR, 2016.
- ADAAPT: A Deep Architecture for Adaptive Policy Transfer from Multiple Sources, J. Rajendran et al., arXiv, 2015.
- Universal Value Function Approximators, T. Schaul et al., ICML, 2015.
Improving Exploration
- Unifying Count-Based Exploration and Intrinsic Motivation, M. G. Bellemare et al., arXiv, 2016.
- Curiosity-driven Exploration in Deep Reinforcement Learning via Bayesian Neural Networks, R. Houthooft et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
- Deep Exploration via Bootstrapped DQN, I. Osband et al., arXiv, 2016.
- Action-Conditional Video Prediction using Deep Networks in Atari Games, J. Oh et al., NIPS, 2015.
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models, B. C. Stadie et al., arXiv, 2015.
Multi-Agent
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks, J. N. Foerster et al., arXiv, 2016.
- Multiagent Cooperation and Competition with Deep Reinforcement Learning, A. Tampuu et al., arXiv, 2015.
Hierarchical Learning
- Deep Successor Reinforcement Learning, T. D. Kulkarni et al., arXiv, 2016.
- Hierarchical Reinforcement Learning using Spatio-Temporal Abstractions and Deep Neural Networks, R. Krishnamurthy et al., arXiv, 2016.
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, T. D. Kulkarni et al., arXiv, 2016.
(zhuan) Deep Reinforcement Learning Papers的更多相关文章
- (转) Deep Reinforcement Learning: Playing a Racing Game
Byte Tank Posts Archive Deep Reinforcement Learning: Playing a Racing Game OCT 6TH, 2016 Agent playi ...
- (转) Deep Reinforcement Learning: Pong from Pixels
Andrej Karpathy blog About Hacker's guide to Neural Networks Deep Reinforcement Learning: Pong from ...
- 论文笔记之:Asynchronous Methods for Deep Reinforcement Learning
Asynchronous Methods for Deep Reinforcement Learning ICML 2016 深度强化学习最近被人发现貌似不太稳定,有人提出很多改善的方法,这些方法有很 ...
- 【资料总结】| Deep Reinforcement Learning 深度强化学习
在机器学习中,我们经常会分类为有监督学习和无监督学习,但是尝尝会忽略一个重要的分支,强化学习.有监督学习和无监督学习非常好去区分,学习的目标,有无标签等都是区分标准.如果说监督学习的目标是预测,那么强 ...
- Deep Reinforcement Learning
Reinforcement-Learning-Introduction-Adaptive-Computation http://incompleteideas.net/book/bookdraft20 ...
- Deep Reinforcement Learning with Iterative Shift for Visual Tracking
Deep Reinforcement Learning with Iterative Shift for Visual Tracking 2019-07-30 14:55:31 Paper: http ...
- 深度强化学习(Deep Reinforcement Learning)入门:RL base & DQN-DDPG-A3C introduction
转自https://zhuanlan.zhihu.com/p/25239682 过去的一段时间在深度强化学习领域投入了不少精力,工作中也在应用DRL解决业务问题.子曰:温故而知新,在进一步深入研究和应 ...
- (转) Playing FPS games with deep reinforcement learning
Playing FPS games with deep reinforcement learning 博文转自:https://blog.acolyer.org/2016/11/23/playing- ...
- Learning Roadmap of Deep Reinforcement Learning
1. 知乎上关于DQN入门的系列文章 1.1 DQN 从入门到放弃 DQN 从入门到放弃1 DQN与增强学习 DQN 从入门到放弃2 增强学习与MDP DQN 从入门到放弃3 价值函数与Bellman ...
随机推荐
- python实现的视频下载工具you-get,支持多个国内外主流视频平台
RT,you-get 是一个视频离线下载工具, https://github.com/soimort/you-get 另一个同类工具 youtube-dl 也是python 实现,虽然名为 youtu ...
- POJ 3683 Priest John's Busiest Day (2-SAT)
题意:有n对新人要在同一天结婚.结婚时间为Ti到Di,这里有时长为Si的一个仪式需要神父出席.神父可以在Ti-(Ti+Si)这段时间出席也可以在(Di-Si)-Si这段时间.问神父能否出席所有仪式,如 ...
- 为什么匿名内部类参数必须为final类型
1) 从程序设计语言的理论上:局部内部类(即:定义在方法中的内部类),由于本身就是在方法内部(可出现在形式参数定义处或者方法体处),因而访问方法中的局部变量(形式参数或局部变量)是天经地义的.是很自 ...
- 【three.js详解之一】入门篇
[three.js详解之一]入门篇 开场白 webGL可以让我们在canvas上实现3D效果.而three.js是一款webGL框架,由于其易用性被广泛应用.如果你要学习webGL,抛弃那些复杂的 ...
- iOS8后core location框架启动定位服务的步骤
1.在使用CoreLocation前需要调用如下函数[iOS 8专用]: iOS 8对定位进行了一些修改,其中包括定位授权的方法,CLLocationManager增加了下面的两个方法: (1)始终允 ...
- 深入理解JavaScript系列:JavaScript的构成
此篇文章不是干货类型,也算不上概念阐述,就是简单的进行一个思路上的整理. 要了解一样东西或者完成一件事情,首要的就是先要搞清楚他是什么.作为一个前端开发人员,JavaScript应该算作是最核心之一的 ...
- session过期,登录页被内嵌iframe的解决方案
在登录页的js加上: if(window !=top){ top.location.href = location.href; } 就可以完美解决这个问题!
- Android自定义权限和使用权限
本文来自http://blog.csdn.net/liuxian13183/ ,引用必须注明出处! 自定义权限,主要用于保护被赋予权限的组件.如无权限与有权限,正如public与private的对类保 ...
- [翻译]为什么IIS应用程序池回收时间默认被设置为1740分钟?
作者:斯科特 福赛斯/Scott Forsyth日期:2013/04/06地址:http://weblogs.asp.net/owscott/why-is-the-iis-default-app-po ...
- android摇一摇实现(仿微信)
这个demo模仿的是微信的摇一摇,是一个完整的demo,下载地址在最下面.下面是demo截图: 步驟: 1.手机摇动监听,首先要实现传感器接口SensorEventLi ...