Skip to main navigation Skip to search Skip to main content

View-invariant human movement assessment

  • Faegheh Sardari

    Student thesis: Doctoral ThesisDoctor of Philosophy (PhD)

    Abstract

    In computer vision, human action or movement assessment is the task of evaluating the quality of a person’s movements when they perform specific actions through observing their video. Typically, human movement assessment approaches are view-specific and are not able to assess the quality of human movement when they are applied to camera viewpoints different to their training data, i.e. unseen viewpoints. This thesis explores view-invariance in human movement assessment, with a particular focus on the healthcare domain. Furthermore, the current approaches in the field of healthcare are based on 3D skeleton data as the features derived from 3D data are rich and can be leveraged to assess a wide range of movements. However, acquiring 3D skeleton data can be cumbersome, if not impractical, in in-the-wild scenarios. This thesis instead focuses on assessing the quality of human movement from RGB data. As all existing action quality assessment datasets are single view, this thesis also introduces two multi-view human movement assessment datasets, SMAD and QMAR, to demonstrate the superior performance of the proposed methods.

    To deal with view-invariance, one solution is to develop a method that is trained on data from multiple views. In this scenario, it is important the method’s complexity does not increase with the number of training views and the developed approach maintains a high performance on the single views. To achieve this, a pose estimation approach is proposed that estimates high-level pose in a canonical manifold space from RGB images, toward human movement assessment under a multi-view learning scenario.

    As capturing a dataset including numerous viewpoints is cumbersome and rare, ideally, a view-invariant approach should be trained on data from as few views as possible while it can operate on arbitrary viewpoints at inference. Thus, this thesis develops an RGB-based approach that learns view-invariant spatio-temporal features by training on only one or two viewpoints and is able to analyse the quality of human movement on novel viewpoints.

    This thesis also presents an unsupervised method that learns view-invariant 3D human posture representation from 2D RGB data for unseen view downstream tasks, e.g. action recognition and assessment, such that the pose features can be transferred into other domains. The proposed method is particularly helpful in applications where the use of multi-view data is essential and recording 3D skeletons is challenging, e.g. action quality assessment in rehabilitation exercises.

    This thesis includes results on SMAD, QMAR, KIMORE, and NTU RGB+D, and obtains comparative evaluation results against the state-of-the-art approaches where it is possible.
    Date of Award21 Jun 2022
    Original languageEnglish
    Awarding Institution
    • University of Bristol
    SupervisorMajid Mirmehdi (Supervisor) & Adeline T M Paiement (Supervisor)

    Cite this

    '