Abstract
Reinforcement learning (RL) is a powerful and flexible paradigm for learning an optimal sequence of actions from trial-and-error interactions with an environment. It has demonstrated tremendous success in various applications. However, it has yet to be applied to a wide range of real-world safety-critical tasks. The main obstacle to the broader application of RL is the safety concern. This thesis tackles the RL safety and robustness issue from three approaches: capturing uncertainty, learning from human feedback and offline RL.Capturing uncertainty is the first step towards achieving safe and robust RL. It tells what is known and what is not known, and we can improve safety by avoiding unknown situations. For our first approach, we investigated a method capturing the uncertainties and utilising it to avoid actions whose outcomes are highly uncertain.
Capturing uncertainty alone cannot help safety in the early stage of the training because everything is highly uncertain at that stage. For our second approach, we employ human feedback to obtain extra information to guide the agent in the training stage and improve its safety. However, human feedback poses issues such as inconsistency, infrequency, and high costs. We deal with these problems by obtaining human feedback from a crowd and estimating each trainer's reliability to avoid negative impact from unreliable trainers.
For our third approach, we investigate offline RL that eliminates the risk of failure in the early training stages by learning the task from static datasets. Several algorithms have been proposed for the offline RL. They have different characteristics, and no single algorithm performs well across all the environments. We propose a novel offline RL approach that improves the robustness and works well across different types of environments.
Our work explored these three lines of approaches and showed how they enhance RL's safety and robustness. We hope this thesis contributes to various real-world RL applications.
| Date of Award | 10 Dec 2024 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Raul Santos-Rodriguez (Supervisor) |
Keywords
- Reinforcement Learning
- Human in the loop
- Machine Learning
- Markov Decision Process
- Uncertainty
Cite this
- Standard