MoodElevate: A Multimodal AI-Driven Mental Health Companion Using Text and Audio-Based Emotion Recognition

  • Unique Paper ID: 202944
  • Volume: 12
  • Issue: 12
  • PageNo: 12510-12516
  • Abstract:
  • MoodElevate is an AI-powered mental and emotional health application that helps users identify their emotional state, track mood patterns, and receive personalised support. The system takes text or voice as input, then applies machine learning models which classify emotions, and delivers customised content such as music, videos, and articles, to help users regulate their mood. A digital journalling feature analyses entries and gives mood-tracking reports. It uses DistilBERT model fine-tuned on the GoEmotions dataset achieved 85–90% accuracy (F1=0.86) for text-based emotion detection, a CNN–LSTM hybrid model trained on CREMA-D dataset showed 79% accuracy for speech-based detection. Multimodal fusion improved overall accuracy by 15%, proving the benefit of combining text and audio inputs for real-world affective computing.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{202944,
        author = {Akanksha Kuvhare and Mansi Konde and Suchet Mahamuni and Piyush Gawali},
        title = {MoodElevate: A Multimodal AI-Driven Mental Health Companion Using Text and Audio-Based Emotion Recognition},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {12},
        pages = {12510-12516},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=202944},
        abstract = {MoodElevate is an AI-powered mental and emotional health application that helps users identify their emotional state, track mood patterns, and receive personalised support. The system takes text or voice as input, then applies machine learning models which classify emotions, and delivers customised content such as music, videos, and articles, to help users regulate their mood. A digital journalling feature analyses entries and gives mood-tracking reports. It uses DistilBERT model fine-tuned on the GoEmotions dataset achieved 85–90% accuracy (F1=0.86) for text-based emotion detection, a CNN–LSTM hybrid model trained on CREMA-D dataset showed 79% accuracy for speech-based detection. Multimodal fusion improved overall accuracy by 15%, proving the benefit of combining text and audio inputs for real-world affective computing.},
        keywords = {Emotion Detection, Emotional Wellness, Artificial Intelligence, Sentiment Analysis, Journalling System, Mood Tracking, DistilBERT, CNN–LSTM, Multimodal Fusion},
        month = {May},
        }

Cite This Article

Kuvhare, A., & Konde, M., & Mahamuni, S., & Gawali, P. (2026). MoodElevate: A Multimodal AI-Driven Mental Health Companion Using Text and Audio-Based Emotion Recognition. International Journal of Innovative Research in Technology (IJIRT), 12(12), 12510–12516.

Related Articles