Development and Integration of a Multimodal Context-Aware Chatbot for Emotion Recognition and Mental Health Support

  • Unique Paper ID: 198652
  • Volume: 12
  • Issue: 11
  • PageNo: 11553-11563
  • Abstract:
  • Mental health chatbots that rely only on text often miss the rich non-verbal signals that clinicians naturally interpret in real-world settings. This paper presents a real-time, CPU-resident multimodal emotion recognition system that integrates three heterogeneous input channels: natural language text (DistilBERT, 7-class), live facial expression analysis (CNN trained on FER2013, 5-class), and heart-rate variability signals (Random Forest trained on WESAD, 3-class). These modalities are fused into a unified seven-class emotion distribution using a novel confidence-weighted dynamic late fusion strategy. In this approach, each modality’s contribution is adjusted at runtime based on its prediction confidence, allowing the system to remain robust and degrade gracefully when one or more inputs are missing or unreliable. On top of this, a context-aware therapeutic response generation pipeline combining Retrieval-Augmented Generation over a local mental-health knowledge base, intent classification, active listening-based reflection, and longitudinal emotion trend tracking generates personalized and non-repetitive empathic responses without relying on external large language model APIs. Evaluation on 30 labelled test scenarios shows a fusion accuracy of 81.4%, which is a 7.2 percentage point improvement over the best single-modality baseline (text: 74.2%) and a 2.8 percentage point improvement over naive equal-weight fusion (78.6%). The system achieves an average response relevance score of 0.53 (0–1 scale) across 120 user interactions. Importantly, the entire pipeline operates on standard CPU hardware and can be deployed offline, making it suitable for privacy-sensitive mental health applications where data security and local processing are essential.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{198652,
        author = {Rahul Jayesh Wala and Prof. Anuradha Deokar},
        title = {Development and Integration of a Multimodal Context-Aware Chatbot for Emotion Recognition and Mental Health Support},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {11},
        pages = {11553-11563},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=198652},
        abstract = {Mental health chatbots that rely only on text often miss the rich non-verbal signals that clinicians naturally interpret in real-world settings. This paper presents a real-time, CPU-resident multimodal emotion recognition system that integrates three heterogeneous input channels: natural language text (DistilBERT, 7-class), live facial expression analysis (CNN trained on FER2013, 5-class), and heart-rate variability signals (Random Forest trained on WESAD, 3-class). These modalities are fused into a unified seven-class emotion distribution using a novel confidence-weighted dynamic late fusion strategy. In this approach, each modality’s contribution is adjusted at runtime based on its prediction confidence, allowing the system to remain robust and degrade gracefully when one or more inputs are missing or unreliable. On top of this, a context-aware therapeutic response generation pipeline combining Retrieval-Augmented Generation over a local mental-health knowledge base, intent classification, active listening-based reflection, and longitudinal emotion trend tracking generates personalized and non-repetitive empathic responses without relying on external large language model APIs. Evaluation on 30 labelled test scenarios shows a fusion accuracy of 81.4%, which is a 7.2 percentage point improvement over the best single-modality baseline (text: 74.2%) and a 2.8 percentage point improvement over naive equal-weight fusion (78.6%). The system achieves an average response relevance score of 0.53 (0–1 scale) across 120 user interactions. Importantly, the entire pipeline operates on standard CPU hardware and can be deployed offline, making it suitable for privacy-sensitive mental health applications where data security and local processing are essential.},
        keywords = {Mental Health Chatbot, Multimodal emotion recognition, affective computing, heart-rate variability, facial expression recognition, DistilBERT, confidence-weighted fusion, retrieval-augmented generation.},
        month = {April},
        }

Cite This Article

Wala, R. J., & Deokar, P. A. (2026). Development and Integration of a Multimodal Context-Aware Chatbot for Emotion Recognition and Mental Health Support. International Journal of Innovative Research in Technology (IJIRT), 12(11), 11553–11563.

Related Articles