Multimodal Deep Learning for Group Activity Recognition in Smart Office Environments

George Albert Florea and Radu-Casian Mihailescu

Resumen

Deep learning (DL) models have emerged in recent years as the state-of-the-art technique across numerous machine learning application domains. In particular, image processing-related tasks have seen a significant improvement in terms of performance due to increased availability of large datasets and extensive growth of computing power. In this paper we investigate the problem of group activity recognition in office environments using a multimodal deep learning approach, by fusing audio and visual data from video. Group activity recognition is a complex classification task, given that it extends beyond identifying the activities of individuals, by focusing on the combinations of activities and the interactions between them. The proposed fusion network was trained based on the audio?visual stream from the AMI Corpus dataset. The procedure consists of two steps. First, we extract a joint audio?visual feature representation for activity recognition, and second, we account for the temporal dependencies in the video in order to complete the classification task. We provide a comprehensive set of experimental results showing that our proposed multimodal deep network architecture outperforms previous approaches, which have been designed for unimodal analysis, on the aforementioned AMI dataset.

Palabras claves

multimodal learning - deep learning - activity recognition

Acceso

PÁGINAS

pp. 0 - 0

NÚMERO

Volumen: 12 Parte: 8 (2020)

MATERIAS

INFRAESTRUCTURA

REVISTAS SIMILARES

Big Data and Cognitive Computing
ISPRS International Journal of Geo-Information

DOI

https://doi.org/10.3390/fi12080133

Artículos similares

Disaster Image Classification by Fusing Multimodal Social Media Data

Acceso

Zhiqiang Zou, Hongyu Gan, Qunying Huang, Tianhui Cai and Kai Cao

Social media datasets have been widely used in disaster assessment and management. When a disaster occurs, many users post messages in a variety of formats, e.g., image and text, on social media platforms. Useful information could be mined from these mul... ver más

Revista: ISPRS International Journal of Geo-Information

A Multi-Modality Deep Network for Cold-Start Recommendation

Acceso

Mingxuan Sun, Fei Li and Jian Zhang

Collaborative filtering (CF) approaches, which provide recommendations based on ratings or purchase history, perform well for users and items with sufficient interactions. However, CF approaches suffer from the cold-start problem for users and items with... ver más

Revista: Big Data and Cognitive Computing

Revistas destacadas

Acceso directo a los números publicados en la revista Infrastructures

Infrastructures

Acceso directo a los números publicados en la revista Informed Infraestructure

Informed Infraestructure

Acceso directo a los números publicados en la revista BiT

Acceso directo a los números publicados en la revista Revista de la Construcción

Revista de la Construcción

Ver todas las revistas