Javascript must be enabled to continue!
Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
View through CrossRef
This paper addresses vectorial form of Markov Decision Processes (MDPs) to solve MDPs with unknown rewards. Our method to find optimal strategies is based on reducing the computation to the determination of two separate polytopes. The first one is the set of admissible vector-valued functions and the second is the set of admissible weight vectors. Unknown weight vectors are discovered according to an agent with a set of preferences. Contrary to most existing algorithms for reward-uncertain MDPs, our approach does not require interactions with user during optimal policies generation. Instead, we use a variant of approximate value iteration on vectorial value MDPs based on classifying advantages, that allows us to approximate the set of non-dominated policies regardless of user preferences. Since any agent's optimal policy comes from this set, we propose an algorithm for discovering in this set an approximated optimal policy according to user priorities while narrowing interactively the weight polytope.
Title: Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
Description:
This paper addresses vectorial form of Markov Decision Processes (MDPs) to solve MDPs with unknown rewards.
Our method to find optimal strategies is based on reducing the computation to the determination of two separate polytopes.
The first one is the set of admissible vector-valued functions and the second is the set of admissible weight vectors.
Unknown weight vectors are discovered according to an agent with a set of preferences.
Contrary to most existing algorithms for reward-uncertain MDPs, our approach does not require interactions with user during optimal policies generation.
Instead, we use a variant of approximate value iteration on vectorial value MDPs based on classifying advantages, that allows us to approximate the set of non-dominated policies regardless of user preferences.
Since any agent's optimal policy comes from this set, we propose an algorithm for discovering in this set an approximated optimal policy according to user priorities while narrowing interactively the weight polytope.
Related Results
Developing Optimal Decision Strategies with Markovian Decision Process
Developing Optimal Decision Strategies with Markovian Decision Process
The use of Markovian Decision Processes (MDPs) in creating the best possible decision strategies is examined in this research. When outcomes are partly controlled by a decision-mak...
Solving Finite-Horizon Discounted Non-Stationary MDPS
Solving Finite-Horizon Discounted Non-Stationary MDPS
Abstract
Research background
Markov Decision Processes (
MDPs
...
Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management
Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management
Abstract
We study the monotonicity properties of optimal policies for a class of fully/partially observable Markov Decision Processes (MDPs) motivated by renewable natural ...
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
O objetivo deste trabalho foi determinar a digestibilidade, usando óxido crômico (Cr2O3) e FDN indigestível, como indicadores, e a degradação de dietas compostas de milho desintegr...
Single-Valued Neutrosophic Ideal Approximation Spaces
Single-Valued Neutrosophic Ideal Approximation Spaces
In this paper, we defined the basic idea of the single-valued neutrosophic upper (αn)δ, single-valued neutrosophic lower (αn)δ and single-valued neutrosophic boundary sets (αn)B of...
Multiobjective Immune Algorithm with Nondominated Neighbor-Based Selection
Multiobjective Immune Algorithm with Nondominated Neighbor-Based Selection
Nondominated Neighbor Immune Algorithm (NNIA) is proposed for multiobjective optimization by using a novel nondominated neighbor-based selection technique, an immune inspired opera...
Role of NGOs in balancing environmental rights within Chinese-funded Mega Development Projects in Sri Lanka
Role of NGOs in balancing environmental rights within Chinese-funded Mega Development Projects in Sri Lanka
The ‘economic growth’ centered development discourse in Sri Lanka, driven by financial assistance from China, concentrates on Mega development projects (MDPs) entailing environment...
Analisis Kebutuhan Modul Matematika untuk Meningkatkan Kemampuan Pemecahan Masalah Siswa SMP N 4 Batang
Analisis Kebutuhan Modul Matematika untuk Meningkatkan Kemampuan Pemecahan Masalah Siswa SMP N 4 Batang
Pemecahan masalah merupakan suatu usaha untuk menyelesaikan masalah matematika menggunakan pemahaman yang telah dimilikinya. Siswa yang mempunyai kemampuan pemecahan masalah rendah...

