Skip to main navigation Skip to search Skip to main content

Improvements in learning to control perched landings

Research output: Contribution to journalArticle (Academic Journal)peer-review

7 Citations (Scopus)
118 Downloads (Pure)

Abstract

Reinforcement learning has previously been applied to the problem of controlling a perched landing manoeuvre for a custom sweep-wing aircraft. Previous work showed that the use of domain randomisation to train with atmospheric disturbances improved the real-world performance of the controllers, leading to increased reward. This paper builds on the previous project, investigating enhancements and modifications to the learning process to further improve performance, and reduce final state error. These changes include modifying the observation by adding information about the airspeed to the standard aircraft state vector, employing further domain randomisation of the simulator, optimising the underlying RL algorithm and network structure, and changing to a continuous action space. Simulated investigations identified hyperparameter optimisation as achieving the most significant increase in reward performance. Several test cases were explored to identify the best combination of enhancements. Flight testing was performed, comparing a baseline model against some of the best performing test cases from simulation. Generally, test cases that performed better than the baseline in simulation also performed better in the real world. However, flight tests also identified limitations with the current numerical model. For some models, the chosen policy performs well in simulation yet stalls prematurely in reality, a problem known as the reality gap.
Original languageEnglish
Pages (from-to)1101-1123
Number of pages23
JournalAeronautical Journal
Volume126
Issue number1301
DOIs
Publication statusPublished - 4 May 2022

Bibliographical note

Funding Information:
This work was part-funded by the Defence Science and Technology Laboratory (DSTL), Ministry of Defence. This work was partially supported by the EPSRC Centre for Doctoral Training in Future Autonomous and Robotic Systems (FARSCOPE) at the Bristol Robotics Laboratory. The authors would like to thank all contributors to ArduPilot [], and the pymavlink and stable-baselines [] libraries, all of which have enabled this work. This work was carried out using the computational facilities of the Advanced Computing Research Centre, University of Bristol – http://www.bris.ac.uk/acrc/

Publisher Copyright:
©

Fingerprint

Dive into the research topics of 'Improvements in learning to control perched landings'. Together they form a unique fingerprint.

Cite this