View article

The predictron: End-to-end learning and planning

Authors

David Silver*, Hado van Hasselt*, Matteo Hessel*, Tom Schaul*, Arthur Guez*, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, Thomas Degris

Publication date

2017

Conference

International Conference on Machine Learning (ICML)

Description

One of the key challenges of artificial intelligence is to learn models that are effective in the context of planning. In this document we introduce the predictron architecture. The predictron consists of a fully abstract model, represented by a Markov reward process, that can be rolled forward multiple “imagined” planning steps. Each forward pass of the predictron accumulates internal rewards and values over multiple planning depths. The predictron is trained end-to-end so as to make these accumulated values accurately approximate the true value function. We applied the predictron to procedurally generated random mazes and a simulator for the game of pool. The predictron yielded significantly more accurate predictions than conventional deep neural network architectures.

Total citations

Cited by 345

2016201720182019202020212022202320242 19 52 51 60 44 46 40 26

Scholar articles

The predictron: End-to-end learning and planning

D Silver, H Hasselt, M Hessel, T Schaul, A Guez… - International Conference on Machine Learning, 2017

Hado Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al. The predictron: End-to-end learning and planning*

D Silver - International conference on machine learning, 2017

Cited by 49 Related articles

The Predictron: End-To-End Learning and Planning. arXiv*

MHTSA Guez, THGDA David, RNRAB Thomas… - arXiv preprint arXiv:1612.08810, 2017

Cited by 2 Related articles