The thesis studies human-in-the-loop methods for off-policy reinforcement learning, focusing on safer and more sample-efficient robot learning.
The work evaluates learning from human interventions experimentally and forms the basis for continued research toward a scientific publication at FZI.