Abstract
No-audio Multimodal Speech Detection is one of the tasks in Media- Eval 2020, with the goal to automatically detect whether someone is speaking in social interaction on the basis of body movement signals. In this paper, a multimodal fusion method, combining signals obtained by an overhead camera and a wearable accelerometer, was proposed to determine whether someone was speaking. The proposed system directly takes the accelerometer signals as input, while using a pre-trained 3D convolutional network to extract the video features that work as input. Experiments on the No-audio Multimodal Speech Detection task show that our method outperforms all submissions of previous years.
| Original language | English |
|---|---|
| Journal | CEUR Workshop Proceedings |
| Volume | 2882 |
| State | Published - 2020 |
| Event | Multimedia Evaluation Benchmark Workshop 2020, MediaEval 2020 - Virtual, Online Duration: 14 Dec 2020 → 15 Dec 2020 |
Fingerprint
Dive into the research topics of 'Multimodal fusion of body movement signals for no-audio speech detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver