Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management

View through CrossRef
Abstract We study the monotonicity properties of optimal policies for a class of fully/partially observable Markov Decision Processes (MDPs) motivated by renewable natural resource management. We first introduce a new class of deterministic monotonic MDPs with continuous state and action spaces, and we prove that their optimal policies are monotonic in the state. These monotonic MDPs are particularly interesting, because they do not satisfy certain supermodularity assumptions on the reward function and the transition dynamics in previous monotonicity analyses for the optimal policies of MDPs, and our proof requires an original inductive analysis. Our theoretical results are numerically illustrated on a fully observable fishery management problem, using a modified value iteration algorithm that computes near-optimal monotonic policies with a max-min smoothing strategy. We then introduce a generalization of our monotonic MDPs to handle partial observability and stochas-ticity, and we conjecture the existence of a near-optimal policy that is monotonic in the expected state. Our conjecture is supported by numerical results on a partially observable fishery management problem, using algorithms for optimizing over a class of monotonic policies called multi-threshold policies. Moreover, when the environment is unknown and needs to be learned through interactions, multi-threshold policies appear to be more robust to model learning error than general policies in a model-based offline reinforcement learning approach.
Springer Science and Business Media LLC
Title: Optimal policy analysis for monotonic (PO)MDPs and an application to fishery management
Description:
Abstract We study the monotonicity properties of optimal policies for a class of fully/partially observable Markov Decision Processes (MDPs) motivated by renewable natural resource management.
We first introduce a new class of deterministic monotonic MDPs with continuous state and action spaces, and we prove that their optimal policies are monotonic in the state.
These monotonic MDPs are particularly interesting, because they do not satisfy certain supermodularity assumptions on the reward function and the transition dynamics in previous monotonicity analyses for the optimal policies of MDPs, and our proof requires an original inductive analysis.
Our theoretical results are numerically illustrated on a fully observable fishery management problem, using a modified value iteration algorithm that computes near-optimal monotonic policies with a max-min smoothing strategy.
We then introduce a generalization of our monotonic MDPs to handle partial observability and stochas-ticity, and we conjecture the existence of a near-optimal policy that is monotonic in the expected state.
Our conjecture is supported by numerical results on a partially observable fishery management problem, using algorithms for optimizing over a class of monotonic policies called multi-threshold policies.
Moreover, when the environment is unknown and needs to be learned through interactions, multi-threshold policies appear to be more robust to model learning error than general policies in a model-based offline reinforcement learning approach.

Related Results

Developing Optimal Decision Strategies with Markovian Decision Process
Developing Optimal Decision Strategies with Markovian Decision Process
The use of Markovian Decision Processes (MDPs) in creating the best possible decision strategies is examined in this research. When outcomes are partly controlled by a decision-mak...
Piece by piece: Collaborative mosaic-making for inclusive policy development
Piece by piece: Collaborative mosaic-making for inclusive policy development
This report sets out the findings from one of four projects commissioned by Wellcome Policy Lab to pilot creative approaches to policy development. In this project, Scientia Script...
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
Digestibilidade e degradabilidade de rações à base de milho desintegrado com palha e sabugo em diferentes graus de moagem
O objetivo deste trabalho foi determinar a digestibilidade, usando óxido crômico (Cr2O3) e FDN indigestível, como indicadores, e a degradação de dietas compostas de milho desintegr...
A bioeconomic analysis of the coastal fishery of Pernambuco State, North-eastern Brazil
A bioeconomic analysis of the coastal fishery of Pernambuco State, North-eastern Brazil
Los conceptos bioeconómicos modernos invocan la importancia de definir los derechos de propiedad posibles de ser llevados a cabo dentro del contexto de gestión. La investigación pe...
Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
Solving MDPs with Unknown Rewards Using Nondominated Vector-Valued Functions
This paper addresses vectorial form of Markov Decision Processes (MDPs) to solve MDPs with unknown rewards. Our method to find optimal strategies is based on reducing the computati...
Impact of fishery association management on fishery community
Impact of fishery association management on fishery community
Tam Giang lagoon system has got an area size of around 22.100 ha; a length of 68 km along the coast of central Vietnam. In which, the animal system is abundant, including many kind...
Responsibilised Resilience? Reworking Neoliberal Social Policy Texts
Responsibilised Resilience? Reworking Neoliberal Social Policy Texts
Introduction This essay begins with the premise that resilience, broadly defined as positive adaptation despite adversity (Garmezy and Rutter), and resilience building are importa...

Back to Top