View article

[PDF] from researchgate.net

Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition

Authors

Dung Nguyen, Kien Nguyen, Sridha Sridharan, David Dean, Clinton Fookes

Publication date

2018/9/1

Journal

Computer vision and image understanding

Volume

174

Pages

33-42

Publisher

Academic Press

Description

Multimodal emotion recognition has attracted great interest recently and numerous methodologies have been successfully investigated. However, the task requires the effective fusion multimodal representations in audio and video domains, and existing approaches still perform poorly on such a challenging task. This paper proposes a novel framework for recognizing emotion from multiple sources including facial expression, pose, body movements, and voice. In this framework, we first introduce new deep spatio-temporal features by cascading 3-dimensional convolution neural networks (C3Ds) and deep belief networks (DBNs) to effectively model spatial and temporal information presented in video and audio for emotion recognition. We subsequently propose a new feature-level fusion approach based on a bilinear pooling theory to combine the visual and audio feature vectors. The proposed fusion strategy …

Total citations

Cited by 103

2019202020212022202320246 13 12 29 25 17

Scholar articles

Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition

D Nguyen, K Nguyen, S Sridharan, D Dean, C Fookes - Computer vision and image understanding, 2018