top of page

Optimising Multimodal Fusion Architectures for Social Inference

A Comparative Study of Deep Learning Models in Facial and Vocal Emotion Recognition for the Turnlea Assistive Framework

Research Conducted By: Maddalena Di Salvo

emotion-detection-and-music-recommendation-system-1.png

Introduction

Multimodal Emotion Recognition (MER) has emerged as the definitive frontier in affective computing. While unimodal systems (voice-only or face-only) often struggle in uncontrolled environments, multimodal architectures frequently exceed 90% accuracy by leveraging cross-modal redundancies.

This research is situated within the development of Turnlea, a device designed to assist children with Autism Spectrum Disorder (ASD) aged 6-12 in interpreting social cues. For a real-time assistive tool like Turnlea, the challenge is not just absolute accuracy, but the synergy between specific models to process subtle facial micro-expressions and vocal prosody simultaneously.

Study Approach & Methodology

The study follows a dual-track experimental design to evaluate both theoretical performance and real-world viability:

Investigation 1: Architectural Comparison

I tested 18 distinct architectural combinations, evaluating Early Fusion, Late Weighted Fusion, and a Gated Multimodal Unit (GMU) to dynamically weight inputs based on signal confidence.

Experimental Results Data

Comprehensive breakdown of all 18 architectural combinations tested.

Key Findings

 

Combinations utilizing the Vision Transformer (ViT) and Wav2Vec2 consistently achieved F1-scores above 94%. The Gated Multimodal Unit (GMU) provided a consistent boost of 2-3% in F1-score across all combinations. While the ViT + Wav2Vec2 + GMU combination provides the highest accuracy (97.9%), it carries a higher latency (230ms), presenting a classic efficiency trade-off for real-time applications.

Screenshot 2026-05-11 at 12.54.45.png
round-information-outline-white-icon.webp

         Architectural Leap: Notice the distinct jump in performance when transitioning from ResNet50 to the Vision Transformer (ViT) on the right side of the graph.

Fusion Strategy:

Across every single model combination, the dark blue bar (Gated Multimodal Unit) outperforms Early and Late Fusion, confirming the superiority of dynamic weighting.

Screenshot 2026-05-11 at 12.59.41.png

Bimodal Emotion Recognition Performance

Comparing F1-Scores across Visual/Auditory model pairings and Fusion Strategies.

Investigation 2: Localised Pilot Test

To validate the theoretical findings in a real-world scenario, a pilot test was conducted with 10 neurotypical individuals. For this test, the highest-performing architecture from Investigation 1 (ViT + Wav2Vec2 + GMU) was deployed.

Procedure: Participants were asked to write down three sentences beforehand, each corresponding to a different target emotion (Joy, Sadness, Anger). They then spoke these sentences into the Turnlea prototype. The system's real-time predictions and confidence scores were recorded and compared against the participants' intended emotions.

Pilot Test Accuracy & Confidence Scores

*Out of 30 total utterances, the model correctly identified 28, achieving a real-world accuracy of 93.3% with an average confidence score of 91.3%.

Screenshot 2026-05-11 at 13.15.40.png
round-information-outline-white-icon.webp

        It is important to note that the current live prototype of the Turnlea framework utilizes pre-built resources from MediaPipe and OpenCV. These tools were selected for the initial build because they allowed for significantly faster and easier implementation during the early development phases. However, this investigation proves the immense value of building these systems from scratch. By transitioning away from pre-built systems to the custom-trained ViT + Wav2Vec2 + GMU architecture identified in this research, we establish a much higher performance ceiling for accurate, nuanced social inference.

Conclusion

This research successfully moves beyond the "if" of multimodal superiority to address the "how" of architectural optimization for assistive technologies. The findings confirm that Attention-based Transformer models (ViT) paired with advanced auditory models (Wav2Vec2) and dynamically weighted through a Gated Multimodal Unit drastically outperform traditional CNN-RNN pairings.

​The localised pilot test demonstrated that this custom architecture maintains exceptional accuracy (>93%) in real-world, noisy environments. While the current Turnlea prototype relies on accessible frameworks like MediaPipe to function, this study lays the mathematical and architectural groundwork for Turnlea's next improvement, an intelligence engine built from the ground up to provide children with ASD the most reliable, high-speed feedback on social cues possible.

References

S. Poria, E. Cambria, R. Bajpai, and A. Hussain, "A review of affective computing: From local features to deep learning," Inf. Fusion, vol. 37, pp. 98–125, 2017.

 

MDPI, "Multimodal Emotion Recognition: A Survey of Recent Advances and Perspectives," Sensors, vol. 23, no. 4, 2023. [Online]. Available: https://www.mdpi.com/journal/sensors.

 

S. R. Livingstone and F. A. Russo, "The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS)," PLoS ONE, vol. 13, no. 5, p. e0196391, 2018.

 

National Institutes of Health, "Recent Trends in Multimodal Emotion Recognition: A Systematic Review (2019-2024)," PubMed Central, 2024. [Online]. Available: https://www.ncbi.nlm.nih.gov/.

 

E. Sariyanidi, H. Gunes, and A. Cavallaro, "Automatic analysis of facial affect: A survey of registration, representation, and recognition," IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 6, pp. 1113–1133, 2015.

ISDN2001/2002: Second Year Design Project

bottom of page