Javascript must be enabled to continue!
Representation Discovery for MDPs Using Bisimulation Metrics
View through CrossRef
We provide a novel, flexible, iterative refinement algorithm to automatically construct an approximate statespace representation for Markov Decision Processes (MDPs). Our approach leverages bisimulation metrics, which have been used in prior work to generate features to represent the state space of MDPs. We address a drawback of this approach, which is the expensive computation of the bisimulation metrics. We propose an algorithm to generate an iteratively improving sequence of state space partitions. Partial metric computations guide the representation search and provide much lower space and computational complexity, while maintaining strong convergence properties. We provide theoretical results guaranteeing convergence as well as experimental illustrations of the accuracy and savings (in time and memory usage) of the new algorithm, compared to traditional bisimulation metric computation.
Association for the Advancement of Artificial Intelligence (AAAI)
Title: Representation Discovery for MDPs Using Bisimulation Metrics
Description:
We provide a novel, flexible, iterative refinement algorithm to automatically construct an approximate statespace representation for Markov Decision Processes (MDPs).
Our approach leverages bisimulation metrics, which have been used in prior work to generate features to represent the state space of MDPs.
We address a drawback of this approach, which is the expensive computation of the bisimulation metrics.
We propose an algorithm to generate an iteratively improving sequence of state space partitions.
Partial metric computations guide the representation search and provide much lower space and computational complexity, while maintaining strong convergence properties.
We provide theoretical results guaranteeing convergence as well as experimental illustrations of the accuracy and savings (in time and memory usage) of the new algorithm, compared to traditional bisimulation metric computation.
Related Results
Developing Optimal Decision Strategies with Markovian Decision Process
Developing Optimal Decision Strategies with Markovian Decision Process
The use of Markovian Decision Processes (MDPs) in creating the best possible decision strategies is examined in this research. When outcomes are partly controlled by a decision-mak...
Bisimulation for quantum processes
Bisimulation for quantum processes
Quantum cryptographic systems have been commercially available, with a striking advantage over classical systems that their security and ability to detect the presence of eavesdrop...
Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management
Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management
Abstract
We study the monotonicity properties of optimal policies for a class of fully/partially observable Markov Decision Processes (MDPs) motivated by renewable natural ...
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
O objetivo deste trabalho foi determinar a digestibilidade, usando óxido crômico (Cr2O3) e FDN indigestível, como indicadores, e a degradação de dietas compostas de milho desintegr...
The Glory of the Past and Geometrical Concurrency
The Glory of the Past and Geometrical Concurrency
This paper contributes to the general understanding of the "geometrical model of concurrency" that was named higher dimensional automata (HDAs) by Pratt and van Glabbeek. In partic...
An algebra of quantum processes
An algebra of quantum processes
We introduce an algebra qCCS of pure quantum processes in which communications by moving quantum states physically are allowed and computations are modeled by super-operators, but ...
Solving Finite-Horizon Discounted Non-Stationary MDPS
Solving Finite-Horizon Discounted Non-Stationary MDPS
Abstract
Research background
Markov Decision Processes (
MDPs
...
Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
This paper addresses vectorial form of Markov Decision Processes (MDPs) to solve MDPs with unknown rewards. Our method to find optimal strategies is based on reducing the computati...

