Wednesday, November 23, 2016

Hierarchical Reinforcement Learning: A Literature Summary

This is a quick summary of current work on hierarchical reinforcement learning (RL) aimed at students choosing to do hierarchical RL projects under my supervision.

The most common formalisation of hierarchical RL in terms of semi-MDPs was given by Sutton, Precup and Singh

There is also a summary of this area by 

In 2015, Pierre-Luc Bacon, Jean Harb and Doina Precup published an article entitled 'The Option-Critic Architecture', describing an algorithm for automatically sub-dividing ans solving an RL problem.

Wednesday, October 26, 2016

Spatio-Temporal Data from Reinforcement Learning

Applying RL algorithms, in a spatial POMDP domains produces spatio-temporal data that it is necessary to analyse and organise in order to produce effective control policies.

There has recently been a great amount of progress in analysing cortical representations of space and time in terms of place-cells, gird cells.  This work has the potential to inform the area of RL in terms of efficient encoding and reuse of spatial data.
The overlap between RL and the neuroscience of mapping local space is particularly interesting as RL can produce raw spatio-temporal data from local sensors.  This provides us with an opportunity to analyse, explore and identify the computational and behavioural principles that enable efficient learning of spatial behaviours.

Neuroscience - Mapping local space

A great introduction to this work is available through the lectures from three Nobel price winners in this area:
There is also a TED talk from 2011 on this subject by Neil Burgess from UCL (in O'Keefe's group) entitled How your brain tells you where you are.  Burgess has a range of more general papers on spatial cognition, including:

A brief colloquial presentation of this research entitled 'Discovering grid cells' is available from the Kavli Insitute of Systems Neuroscience's Centre for Neural Computation.

There was also a nice review article from in the Annual Review of Neuroscience entitled 'Place Cells, Grid Cells, and the Brain's Spatial Representation System', Vol. 31:69-89, 2008, by Edvard I. Moser, Emilio Kropff and May-Britt Moser.

There was also a Hippocampus special issue in grid-cells in 2008 edited by Michael E. Hasselmo, Edvard I. Moser and May-Britt Moser.

Recently there was another summary article in Nature Reviews Neuroscience entitled 'Grid cells and cortical representation', Vol. 15:466–481, 2014, by Edvard I. Moser, Yasser Roudi, Menno P. Witter, Clifford Kentros, Tobias Bonhoeffer and May-Britt Moser.

Further relevant work has recently been presented in an article entitled 'Grid Cells and Place Cells: An Integrated View of their Navigational and Memory Function' in Trends in Neurosciences, Vol. 38(12):763–775, 2015, by Honi Sanders, César Rennó-Costa, Marco Idiart and John Lisman.

A more general introcution to

Computational Approaches

There is a review article on computational approaches to these issues entitled 'Place Cells, Grid Cells, Attractors, and Remapping' in Neural Plasticity, Vol. 2011, 2011 by Kathryn J. Jeffery.

Other relevant articles:

  • 'Impact of temporal coding of presynaptic entorhinal cortex grid cells on the formation of hippocampal place fields' in Neural Networks, 21(2-3):303-310, 2008, by  Colin Molter and Yoko Yamaguchi.
  • 'An integrated model of autonomous topological spatial cognition' in Autonomous Robots, 40(8):1379–1402, 2016, by Hakan Karaoğuz and Işıl Bozma.
  • In 2003, in a paper entitled 'Subsymbolic action planning for mobile robots: Do plans need to be precise?', John Pisokas and Ulrich Nehmzow used the topology-preserving properties of self-organising maps to create spatial proto-maps that supported sub-symbolic action planning in a mobile robot.
  • A paper entitled Emergence of multimodal action representations from neural network self-organization by German I. Parisi, Jun Tani, Cornelius Weber and Stefan Wermter includes an intteresting section called 'A self-organizing spatiotemporal hierarchy' wich addresses the automated structuring of spetio-temporal data.

Monday, April 11, 2016

A short bibliography BBAI and hierarchical RL with SOMs

This bibliography is meant for anyone who joins my research group to work on hierarchical reinforcement learning algorithms or related areas,

My publications

  • Georgios Pierris and Torbjørn S. Dahl, Learning Robot Control based on a Computational Model of Infant Cognition. In the IEEE Transactions on Cognitive and Developmental Systems, accepted for publication, 2016.
  • Georgios Pierris and Torbjørn S. Dahl, Humanoid Tactile Gesture Production using a Hierarchical SOM-based Encoding. In the IEEE Transactions on Autonomous Mental Development, 6(2):153-167, 2014.
  • Georgios Pierris and Torbjørn S. Dahl, A Developmental Perspective on Humanoid Skill Learning using a Hierarchical SOM-based Encoding. In the Proceedings of the 2014 International Joint Conference on Neural Networks (IJCNN'14), pp708-715, Beijing, China, July 6-11, 2014.
  • Torbjørn S. Dahl, Hierarchical Traces for Reduced NSM Memory Requirements. In the Proceedings of the BCS SGAI International Conference on Artificial Intelligence, pp165-178, Cambridge, UK, December 14-16, 2010.

Relevant papers

  • Daan Wierstra, Alexander Forster, Jan Peters and Jurgen Schmidhuber, Recurrent Policy Gradients.  In Logic Journal of IGPL, 18:620-634, 2010. [pdf from IDSIA]
  • Andrew G. Barto and Sridhar Mahadevan, Recent advances in hierarchical reinforcement learning, Discrete Event Dynamic Systems, 13(4):341-379, 2003. [pdf from Citeseer]
  • Harold H. Chaput and Benjamin Kuipers and Risto Miikkulainenn Constructivist learning: A neural implementation of the schema mechanism.  In the Proceedings of the Workshop on Self-Organizing Maps (WSOM03), Kitakyushu, Japan, 2003. [pdf from Citeseer]
  • Leslie B. Cohen, Harold H. Chaput and Cara H. Cashon, A constructivist model of infant cognition, Cognitive Development, 17:1323–1343, 2002 [pdf from ResearchGate]
  • Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning, In Artificial Intelligence, 112:181–211, 1999. [pdf from the University of Alberta]
  • Patti Maes, How to do the right thing, Connection Science Journal, 1:291-323, 1989. [pdf from Citeseer]
  • Rodney A. Brooks, A Robust Layered Control System for a Mobile Robot, IEEE Journal of Robotics and Automation, 2(1):14-23, 1986. [pdf of MIT AI Memo 864]

Books

  • Joaquin M. Fuster, Cortex and mind: Unifying cognition, Oxford University Press, 2003. [pdf from ResearcgGate]
  • Richard S. Sutton and Andrew G. Barto, Reinforcement learning: An introduction, MIT Press, 1998. [pdf of unfinished 2nd edition]
  • G. L. Drescher, Made-up minds, MIT Press, 1991 [pdf of MIT dissertation] - An actual constructivist architecture.

Wednesday, January 06, 2016

Sequence Similarity for Hidden State Estimation

Little work has been done on comparing long- and short-term memory (LTM and STM) traces in the context of hidden state estimation in POMDPs.  Belief-state algorithms use a probabilistic step-by-step approach which should be optimal, but doesn't scale well and has an unrealistic requirement for knowledge of the state space underlying observations, 

The instance-based Nearest Sequence Memory (NSM) algorithm performs remarkably well without any knowledge of the underlying state space.  Instead is compares previously observed sequences of observations and actions in LTM with a recently observed sequence in STM to estimate the underlying state.  The NSM algorithm uses a count of matching observation-action records as a metric for sequence proximity.

In problems where certain observation-actions are particularly salient, e.g., passing through a 'door' in Sutton's world, or picking up a passenger in the Taxi problem, a simple match count is not a particularly good sequence proximity metric and, as a result, I have recently been casting around for other work on such metrics.
ADD SLTM work reference here.

I have come across some interesting work on sequence comparison by Z. Meral Ozsoyoglu, Cenk Sahinalp and Piotr Indyk on Improving Proximity Search for Sequences.  This work has been done in the context of genome sequencing could be an interesting starting point for going beyond the simple match count metric used by NSM.

Friday, October 30, 2015

Robagogy - Imitating and exploring

Just saw this new article on how human children mix imitation and exploration entitled Imitation and Innovation: The Dual Engines of Cultural Learning.  It seems that the level to which they rely on imitation is reduced reliably with increased performance.  This should inform our work on robot learning.
This and similar insights from psychology makes me think we should define a new field of robagogy, meaning 'how to teach robots'.  This area should be informed by pedagogy 'how to teach children' and the less popular andragogy or 'how to teach adults'.  

Wednesday, February 25, 2015

Beyond Random Motor Babbling: How to explore when learning behaviors

Humans can learn or improve skills by practicing, e.g., tennis strokes improve as you spend time whacking balls against a wall.  This process is likely to involve exploration of a problem space, in the case of tennis, trying different the actuator values to develop appropriate responses to a range of different ball trajectories and speeds.

In robot skill learning, an important question is 'how do we do this exploration?'  Many papers use random motor babbling [refs], typically within a safe envelope of values and most commonly with a flat probability distribution within each range [refs].  It is highly unlikely that this is how humans and animals explore.

It has been recognised for some time [ref] that system noise can provide sufficient variability to support exploration leading to the identification of effective solutions.  Pinheiro et al. [1] just showed that such variability is also a sufficient requirement for learning new skills in humans.
Think more about using this to inspire new, biologically more plausible, exploration strategies for robot learning.

[1] João de Paula Pinheiro, Pricila Garcia Marques, Go Tani and Umberto Cesar Corrêa (2015) Diversification of motor skills rely upon an optimal amount of variability of perceptive and motor task demands.  In Adaptive Behavior, (only online so far http://adb.sagepub.com/content/early/2015/02/23/1059712315571369?papetoc ).


Thursday, May 23, 2013

Other health care projects

There are a number of healthcare research projects that are highly relevant to tele-care.  Below are the ones I know about so far:


  •  Sensor Platform for HEalthcare in a Residential Environment (SPHERE), EPSRC

Thursday, October 04, 2012

Tele-assisted living

The idea behind tele-assisted living is that many services can be provided remotely if there is a mobile manipulator connected to the Internet in someone's home.
Tele-operating this platform will remove the requirement for high levels of robot autonomy which we are not likely to see for decades.

Such a platform could also provide a large amount of training data to speed up the development of robot autonomy.  My interest is in the area of robot learning and the delegation of skills from a tele-operator to the robot.

Many research activities are currently ongoing in this area and this blog post is meant as a list of these.

FP7 projects

  • ACCOMPANY: Acceptable robotiCs COMPanions for AgeiNg Years (includes Hertfordshire and Birmingham)
  • AALIANCE2: The European Ambient Assisted Living Innovation Alliance (includes Tunstall Healthcare Ltd., UK)
  • CAPSIL: International Support of a Common Awareness and Knowledge Platform for Studying and Enabling Independent Living (FP7 Support Action, includes Imperial College, London)
  • CONFIDENCE: Ubiquitous Care System to Support Independent Living (no UK partner)
  • DOMEO: Domestic robot for elderly assisteance (no UK partner) 
  • FLORENCE: Multi-Purpose Mobile Robot for Ambient Assisted Living (no UK partners)
  • KSERA: Knowledgeable Service Robots for Ageing (no UK partners)
  • MOBISERV: An Integrated Intelligent Home Environment for the Provision of Health, Nutrition and Mobility Services to the Elderly (includes Bristol)
  • ROBOT-ERA : Implementation and Integration of Advanced Robotic Systems and Intelligent Environments in Real Scenarios for the Ageing Population (includes Plymouth)
  • SCRIPT: Supervised Care and Rehabilitation Involving Personal Tele-Robotics (includes Hertfordshire and Sheffield)
  • SRS: Multi-Role Shadow Robotic System for Independent Living (includes Cardiff and Bedfordshire)

Interest groups

Thursday, April 26, 2012

Horizontal Connections

I have still to find a neural model of the competitive selection mechanism that finds the winning node of a SOM. The process has been parallelised many times for parallel SOM implementations based on FPGAs or CUDA (such as Aquila).  There is a lot of work on horizontal/lateral connections in the cortex.  It would be very interesting to study a SOM implementation using such lateral connections to select SOM winners.

Endogenously Active Elements

Endogenous activity has been raised as an important mechanism for cognition, both in terms of neural self-organisation (Choa and Choi 2012) and for control (Bechtel and Abrahamsen 2010 ).  Such a mechanism would be a very interesting extension to our PLANCS framework for behavior-based robotics (Dahl and Giraud-Carrier 2005).
  • Bechtel, W. and Abrahamsen, A. (2010) Understanding the brain as an endogenously active mechanism. In Proceedings of the 32nd Annual Conference of the Cognitive Science Society, Austin, Texas. (see Bechtel's web site http://mechanism.ucsd.edu/~bill/index.html)
  • Choa, M. W. and Choi, M. Y. (2012) Spontaneous organization of the cortical structure through endogenous neural firing and gap junction transmission, Neural Networks 31:46–52.
  • T. S. Dahl and C. Giraud-Carrier, "Incremental Development of Adaptive Behaviors using Trees of Self-Contained Solutions," Adaptive Behavior, 13(3)243-260, 2005.

Monday, March 19, 2012

Interpolation with SOMs

The traditional use of a SOM is to reduce the dimensionality of the data through vector quantization.
It may be possible, however, to use SOMs to interpolate across a data space based on a small number of data points.  Turns out that not only is it possible, but it has been done by a number of people!

  • Goppert and Rosenstiel (1997) The Continuous Interpolating Self-Organizing Map, Neural Processing Letters, 5:185-192.
  • Yin and Allinson (1999) Interpolating self-organizing map (iSOM), Electronics Letters, 35(19):1649-1650.
  • Kawano, Orii, Shiraishi and Maeda (2010) A Method for Multiple Image Interpolation Employing Self-Organizing Map, Proceedings of the SMC'2010, pages 4035-4040, 10-13 October, Istanbul.
Now, applying this to RL so that we can construct an attractor landscape is a very interesting idea. 

Wednesday, February 15, 2012

Incremental exploration

Learn from demonstration typically means learning from training data that are in the form of a relatively small number of complex sequences of observations and potentially actions. The strength of this learning paradigm is that the data provided is related to the crucial areas of the problem space. In the case of reinforcement learning, this would involve the key reward states and effective paths to these states from relevant starting states. However, due to the restricted part of a problem space that can be covered using this form of learning, it typically leads to brittle behaviours that are not able to compensate for perturbations that place the robot outside the known area. One solution to this problem is to use the training data to learn a policy that generalises across large areas of the problem space, such as the Nonlinear Dynamical Systems presented by Ijspeert et al. [3]. Another approach is to hard code a mechanism for returing to the known area such as the extension to Gaussian Mixture Models presented by Calinon [2]. Abbeel and Ng [1], argued, from their experience in the domain of autonomous helicopter control, that an explicit exploration policy is not required in order to improve performance up to or beyond that of the teacher. Instead, the natural perturbations would provide sufficient exploration.

Bill Smart and Leslie Kaelbling [5] developed the JAQL (Joystick and Q-learning?) algorithm to overcome this problem. The JAQL algorithms has two different learning phases. In the first phase, the robot is driven through the "interesting" parts of the problem space by a hand coded controller or by a human controller using a joystick. In the second phase, the policy learned was in control and responsible for further exploration, running in a more standard reinforcement learning mode. The JAQL algorithm has an explicit exploration policy designed to work with policies learnt from demonstration.

The JAQL exploration policy creates slight deviations from the greedy action by adding a small amount of Gaussian noise [4]. This policy creates actions that are "similar to, but different from", the greedy action.

Our RLSOM algorithm has so far been applied only to learning by demonstration, but should be capable of handling learning from exploration without other modifications that a reasonable exploration policy. This is one of the most exciting direction in which to take our research.

[1] Pieter Abbeel and Andrew Y. Ng, Exploration and apprenticeship learning in reinforcement learning. In the Proceedings of the 22nd International Conference on Machine Learning (ICML'05), pp1-8, August 7-11, Bonn, Germany, 2005.

[2] Sylvain Calinon, Robot Programming by Demonstration: A Probabilistic Approach. EPFL/CRC Press, 2009.

[3] Auke J. Ijspeert, Jun Nakanishi and Stefan Schaal, Movement imitation with nonlinear dynamical systems in humanoid robots. In the Proceedings of the International Conference on Robotics and Automation (ICRA'02), pp1398-1403, May 11 - 15, Washington, DC, 2002.

[4] William D. Smart, Making Reinforcement Learning Work on Real Robots. Ph.D. thesis, Department of Computer Science, Brown University, 2002.

[5] William D. Smart and Leslie Pack Kaelbling, Reinforcement Learning for Robot Control. In Mobile Robots XVI (Proceedings of the SPIE 4573), pp92-103, Douglas W. Gage and Howie M. Choset (eds.), Boston, Massachusetts, 2001.

Sunday, February 05, 2012

Dimensions of cognition: A Generalist Manifesto

The article below builds on some old ideas discussed in my old ECAL 2001 paper, Evolution, Adaptation, and Behavioural Holism in Artificial Intelligence and also included in my PhD dissertation Behaviour-Based Learning: Evolution-Inspired Development of Adaptive Robot Behaviours. They were originally a comment on behavior-based robotics as presented in Rodney Brooks's papers A Robust Layered Control System for a Mobile Robot and Elephants Don't Play Chess, and in Ronald C. Arkin's book Behavior-based Robotics.

These ideas gained new relevance recently when I participated in a the Challenges for Artificial Cognitive System II workshop arranged by the European Network for the Advancement of Artificial Cognitive Systems, Interaction and Robotics (EUCogII). Some ideas from that workshop were captures in the workshop wiki. Below I summarize what I took from the workshop.

There is a history, in the science of AI and its related fields, to look a problems in isolation and to simplify problems to the extent where the solutions do contribute significantly to our knowledge of cognition. Three examples are symbolic problem solving, Computer Vision and Speech Recognition. While each of these areas have produced valuable technologies, in their own right, they have also, by narrowing their focus, removed themselves so far from the problems faced by animals, including humans, that their solutions have not, to any great extent, helped our understanding of cognitive systems. One cannot criticize this work for simplifying and narrowing their studies, as this is a necessary part of developing working solutions to given problems. The specificity of their solutions however, raises the question of whether it is possible to develop systems that both solve a specific problem, and also provide key insights into cognition. Cognition is, arguably, the ability to apply a general understanding of the world to new problems and situations, and, if so, cognition is generalization rather than specialization.
A generalist approach to modelling cognition raises two fundamental questions:
  1. Generalize across what?
  2. What kind of problems can provide tractable challenges for generalist cognitive systems?
By challenges being tractable we mean that it is possible to imagine solutions based on the incremental development or integration of existing technologies. There are many examples of interesting but intractable challenges for generalist cognitive systems. Most challenges requiring human-like cognitive abilities such as autonomous robot workers or companions are clearly still intractable, but many challenges which have, arguably, lower cognitive requirements, e.g., robotic sheep dogs, guard dogs, steeds or even pack animals, also look intractable w.r.t. many of the sub-problems they contain.
Beyond finding tractable challenges it is also interesting to consider whether we can identify a sequence of challenges that could form milestones along a path of increasingly high levels of generalist cognitive abilities. Following from the second question, we can also ask whether any such problems would be practical, by which we mean that solving them would provide a technology that could be useful to society.
Physiology
Recognizing that cognition is dependent on physiology introduces a number of physiological dimensions of cognition:
  • Sensors; Vision (stereo, colour), audition (stereo), proprioceptive, tactile
  • Actuators; Muscles, Legs, arms, hands, thumbs
  • Nervous system; Spine, limbic system, cortical architecture
Neil R. Carlson's Physiology of Behavior is a good introduction to sensors, actuators and related physiology. The architecture of the complete brain is described in Larry W. Swanson's book Brain Architecture: Understanding the Basic Plan. The cortical architecture is discussed in Joaquin M Fuster's book Cortex and Mind: Unifying Cognition.
Environment
The environment is also a crucial actor in enabling or prohibiting intelligent behaviour. I have divided this into:
  • Physical environment
  • Social environment
A book that discusses some issues in this respect is John Alcock's Animal Behavior: An Evolutionary Approach.
Cognition
Anette Karmiloff-Smith, in her book Beyond Modularity: A developmental perspective on cognitive science builds on traditional developmental approaches such as those of Piaget and Fodor to suggest that a child's cognitive development takes place within five different domains before the process of representational redescription produces domain generic knowledge from the previous domain specific knowledge. Anette Karmiloff-Smith considers the following domains:
  1. The child as a linguist
  2. The child as a physicist
  3. The child as a mathematician
  4. The child as a psychologist
  5. The child as a notator
Looking for dimensions of cognition, Stephen Mithen, in his book The prehistory of the mind, 1996, suggests a number of specialized intelligences that act as a foundation, or 'chapels' around an initial general intelligence. On top of these chapels, the 'superchapel' of meta-representation is then built, to provide the cognitive capabilities of modern humans. The specialized intelligences suggested by Steven Mithen are:
  1. Technical intelligence
  2. Natural history intelligence
  3. Social intelligence
  4. Linguistic intelligence
Finally, in the book A Roadmap for Cognitive Development in Humanoid Robots, David Vernon, Claes von Hofsten and Luciano Fadiga have used knowledge from human cognitive development to define a road map for cognitive development in humanoid robots. This work is very much in the spirit of what I suggest here, but I would like to consider a wider scientific that will give us a better chance of identifying realistic milestones.
Examples
The Kizmet robot was developed at MIT and used in a wide range of research activities.
The iCub robot is a popular humanoid research robot that has also been used for a wide range of research activities.  As a result, it also has a well developed cognitive architecture.

Monday, January 23, 2012

Prioritized sweeping

In considering how we can make our RLSOM algorithm more efficient, the idea of focusing the activation spreading around the area of maximum activity was raised (by my PhD student Georgios Pierris). This concept was formalised as 'prioritized sweeping' by Moore and Atkeson in 1993.
Like growing SOMs, this is a concept I would like to explore further.

Reference:
Moore, A.W., Atkeson, C.G.: Prioritized sweeping: Reinforcement learning with less data and less time. Machine Learning, 13:103–130, 1993.

Monday, October 31, 2011

Matching landscapes

Georgios Pierris, one of my PhD students had a very good idea today. In combining SOMs and RL we are defining two landscapes, the SOM landscape and the reward landscape. It would be very interesting to see if it would be possible to use the SOM landscape to approximate the RL landscape and, thus, calculate utilities for unexplored areas of the RL landscape. It brings to mind Ulrich Nehmzow's work on sub-symbolic planning.
  • John Pisokas and Ulrich Nehmzow, Performance Comparison of Three Subsymbolic Action Planners for Mobile Robots, Robotics and Autonomous Systems, 51(1):55-67, 2005
Must develop this idea further...

Monday, October 17, 2011

Dimensions of RLSOM

At the Cognitive Robotics Research Centre at the University of Wales, Newport, we have been working on a reinforcement learning self-organizing map (RLSOM).
Without going into the details of the RLSOM algorithm, I'd liek to list some potential extensions:
  • From fixed length memory to primitives
  • STM Length and Learning Frequency
  • Representing space using decaying activation
  • Sparse representation of connections
  • One-shot and incremental learning
  • Pre-structuring hierarchies
  • Random connections across hierarchical levels

Monday, September 27, 2010

Stability

I have remembered another feature that is required of efficient robot learning. This is stability, i.e., the ability to correct perturbed trajectory. Bugmann et al. (2006) demonstrate this property in an autonomous wheelchair. When perturbations occur, blind, i.e., open loop, reproduction of a learned trajectory will not lead the robot to a target. In order to reach the goal, the robot must be able to receive feedback and to correct the learned trajectory so that the final target is reached.

This feature has been described by Ijspeert et al. (2002) as an attractor landscape where the the learned policy will produce trajectories that move towards the ideal trajectory. This can be achieved through designing reward functions that promote the ideal trajectory. Algorithms that explore the landscape around this will produce an attractor landscape through the discounted rewards.

Thursday, September 23, 2010

Features Necessary for Efficient Robot Learning

I have been collecting features I think are necessary for efficient robot learning algorithms. Much of my research lately has been related to integrating these features into a single system.

The features so far are:
  1. Sequential - Sequences of events in time. This is something artificial neural networks (ANNs) are typically not very good at.

  2. Hierarchical - Hierarchical structures for long term memory (LTM). This enables reuse of low level behaviours and efficient encoding of observations. It is a well studied problem in reinforcement learning (RL) and Barto and Mahadevan (2003) have published a good review paper.

  3. Incremental - Observations are processed in the sequence they are made without storing and revisiting old observations, improving the existing system before further observations are made. Reinke and Michalski (1988) first introduced this concept w.r.t. their incremental AQ algorithm for learning concept descriptions. Many ANNs are statistic, i.e., they need to repeatedly pass through batches of observations.

  4. One-shot - Able to learn from a single, or a small number of, observations. Bayesian approaches to this problem have been published by Fei-Fei et al. (2003) and Maas and Kemp (2009). Instance-based learning algorithms such as the Nearest Sequence Memory (NSM) algorithm presented by McCallum (1996) are an extreme form of one-shot learning. Wu and Demiris (2010) recently published an algorithm that is both hierarchical and one-shot.

  5. Auto-associative - Retrieving memory content using the content itself as a reference, e.g., retieving a stored image of a person's face using another image or a partially obscured image of that face. This is what neural networks are good at and it makes them able to handle noisy real world date. This is something RL algorithms are typically not so good at.

  6. Future reward prediction - Predicting what actions will optimise future rewards in what world states. This is a core feature of RL algorithms (Sutton & Barto 1998).

  7. Hidden state identification - Different world states can produce identical observations, e.g., two different corridors in a building can look exactly the same. The only way to tell these states apart is to remember observations from the past, e.g., you know what corridor you're in because you remember what floor you took the elevator to. Some, but not all, RL algorithms have this feature.

Cohen et al. (2002) have presented the Constructivist Learning Architecture (CLA) that was both sequential, hierarchical and auto-associative, using decaying node activities in a hierarchical self-organising map (SOM) or Kohonen network (Kohonen 2001).

My paper (2010) presented a hierarchical extension of McCallum's NSM algorithm that was both sequential, one-shot and hierarchical with future reward prediction and hidden state identification capabilities.

Pierris and I (2010) published an algorithm that used a version of the Cohen's CLA algorithm, the Compressed Sparse Code (CoSCo) SOM, to reproduce a humanoid robot motion demonstrated by a human teacher. This work went beyond the original work of Cohen et al. in that it repeatedly used the learned SOM for action selection in order to reproduce the motion.

In the future I aim to use the mechanisms I developed for the hierarchical NSM algorithm to extend Chaput's decaying activity hierarchical SOMs so that it can support future reward discount and hidden state identification. This will create an algorithm suitable for RL problems such as reaching.