Skip to main navigation Skip to search Skip to main content

Combining CNN streams of RGB-D and skeletal data for human activity recognition

  • Pushpajit Khaire*
  • , Praveen Kumar
  • , Javed Imran
  • *Corresponding author for this work

    Research output: Contribution to journalArticle (Academic Journal)peer-review

    140 Citations (Scopus)

    Abstract

    Inspired by the success of deep learning methods, for human activity recognition based on individual vision cues, this paper presents a ConvNets based approach for activity recognition by combining multiple vision cues. Moreover, a new method of creating skeleton images, from skeleton joint sequences, representing motion information is presented in this paper. Motion representation images, namely, Motion History Image (MHI), Depth Motion Maps (DMMs) and skeleton images are constructed from RGB, depth and skeletal data of RGB-D sensor. These images are then separately trained on ConvNets and respective softmax scores are fused at the decision level. The combination of these distinct vision cues, leads to complete utilization of data, available from RGB-D sensor. To evaluate the effectiveness of the proposed 5-CNNs approach, we conduct our experiments on three well known and challenging RGB-D datasets, CAD-60, SBU Kinect interaction and UTD-MHAD. Results show that the proposed approach of combining multiple cues by means of decision level fusion is competitive with other state of the art methods.

    Original languageEnglish
    Pages (from-to)107-116
    Number of pages10
    JournalPattern Recognition Letters
    Volume115
    DOIs
    Publication statusPublished - 1 Nov 2018

    Bibliographical note

    Funding Information:
    This research was supported by Science and Engineering Research Board (SERB) under project no. ECR/2016/000387 , in cooperation with the Department of Science & Technology (DST), Government of India. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of DST-SERB or the Government of India. The DST-SERB or Government of India is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation thereon.

    Funding Information:
    This research was supported by Science and Engineering Research Board (SERB) under project no. ECR/2016/000387, in cooperation with the Department of Science & Technology (DST), Government of India. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of DST-SERB or the Government of India. The DST-SERB or Government of India is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation thereon.

    Publisher Copyright:
    © 2018 Elsevier B.V.

    Keywords

    • Convolutional neural networks
    • Deep learning
    • Depth motion map
    • Motion history image and fusion
    • RGB-D sensors
    • Skeleton
    • UTD-MHAD

    Fingerprint

    Dive into the research topics of 'Combining CNN streams of RGB-D and skeletal data for human activity recognition'. Together they form a unique fingerprint.

    Cite this