State-augmented reinforcement learning
Inventors
Chen, Pin-Yu • Zhu, Yada • Xiong, JinJun • Bhaskaran, Kumar • Ye, Yunan • Li, Bo
Assignees
International Business Machines Corp • University of Illinois at Urbana Champaign
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A processor training a reinforcement learning model can include receiving a first dataset representing an observable state in reinforcement learning to train a machine to perform an action. The processor receives a second dataset. Using the second dataset, the processor trains a machine learning classifier to make a prediction about an entity related to the action. The processor extracts an embedding from the trained machine learning classifier, and augments the observable state with the embedding to create an augmented state. Based on the augmented state, the processor trains a reinforcement learning model to learn a policy for performing the action, the policy including a mapping from state space to action space.
Core Innovation
The invention provides a state-augmented reinforcement learning framework in which a reinforcement learning observable state is augmented with an embedding extracted from a machine learning classifier. A first dataset representing an observable state in reinforcement learning is used to train a machine to perform an action, while a second dataset is used to train the classifier to make a prediction about an entity related to the action. The trained classifier outputs an embedding that converts heterogeneous information into a usable vector representation for a reinforcement learning model.
The augmented state is used to train a reinforcement learning model to learn a policy for performing the action, where the policy includes a mapping from state space to action space. The embedding is defined as a vector representation of heterogeneous information converted by the machine learning classifier into a usable form by the reinforcement learning model. This pipeline connects classifier-based prediction features with reinforcement learning policy learning through the augmented state.
In the described examples and supporting discussion, the entity prediction is asset movement and the action is asset allocation for portfolio management. Encoders/classifiers are used to predict asset movement from prices and/or financial news articles, and the resulting movement embedding is incorporated into the reinforcement learning model to support asset allocation. The framework is positioned for handling noisy, imbalanced, and non-stationary data, with experiments/simulations indicating improved results.
Claims Coverage
The document contains three independent claims, each centered on the same core inventive pipeline: train a classifier on a second dataset, extract an embedding, augment a reinforcement learning observable state with the embedding, and train a reinforcement learning model to learn a policy mapping state space to action space. Across the independent claims, the inventive focus is on converting heterogeneous information into a usable vector embedding that conditions reinforcement learning policy learning.
Augmenting reinforcement learning observable state using classifier embedding
Receiving a first dataset representing an observable state in reinforcement learning to train a machine to perform an action; receiving a second dataset; training a machine learning classifier using the second dataset to make a prediction about an entity related to the action; extracting an embedding from the trained machine learning classifier; augmenting the observable state with the embedding to create an augmented state; and based on the augmented state, training a reinforcement learning model to learn a policy for performing the action, the policy including a mapping from state space to action space, wherein the embedding includes a vector representation of information in heterogeneous form converted by the machine learning classifier into a usable form by the reinforcement learning model.
Machine learning classifier prediction embedding conditioned reinforcement learning policy learning
Receive a hardware processor configured to receive a first dataset representing an observable state in reinforcement learning to train a machine to perform an action; receive a second dataset; train a machine learning classifier using the second dataset to make a prediction about an entity related to the action; extract an embedding from the trained machine learning classifier; augment the observable state with the embedding to create an augmented state; and based on the augmented state, train a reinforcement learning model to learn a policy for performing the action, the policy including a mapping from state space to action space, wherein the embedding includes a vector representation of information in heterogeneous form converted by the machine learning classifier into a usable form by the reinforcement learning model.
Computer program product implementing embedding-augmented reinforcement learning policy mapping
A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to cause the device to: receive a first dataset representing an observable state in reinforcement learning to train a machine to perform an action; receive a second dataset; train a machine learning classifier using the second dataset to make a prediction about an entity related to the action; extract an embedding from the trained machine learning classifier; augment the observable state with the embedding to create an augmented state; and based on the augmented state, train a reinforcement learning model to learn a policy for performing the action, the policy including a mapping from state space to action space, wherein the embedding includes a vector representation of information in heterogeneous form converted by the machine learning classifier into a usable form by the reinforcement learning model.
Across the independent claims, the shared claim coverage is the use of a classifier-trained embedding from heterogeneous information to augment a reinforcement learning observable state, thereby enabling reinforcement learning policy learning as a state-space-to-action-space mapping for performing the action.
Stated Advantages
Improved results in experiments/simulations.
Handling noisy, imbalanced, and non-stationary data.
Documented Applications
Portfolio management, where predicting asset movement and performing asset allocation uses the embedding-augmented reinforcement learning framework.
Interested in licensing this patent?