AI BASED MULTIMODAL LIE DETECTION VIDEO AND AUDIO CALLS

  • Unique Paper ID: 200065
  • Volume: 12
  • Issue: 12
  • PageNo: 4069-4079
  • Abstract:
  • The rapid growth of online communication platforms such as video conferencing and voice-based interactions has increased the need for reliable and automated deception detection systems. Traditional lie detection techniques, including polygraph tests, rely on physiological signals and often suffer from limited accuracy, lack of scalability, and susceptibility to manipulation. To address these challenges, this research proposes a Multimodal AI-based framework for detecting deception in real-time video and audio calls. The proposed system integrates three primary modalities: visual, acoustic, and linguistic features. The visual module analyzes facial expressions, micro-expressions, and eye movements using computer vision techniques and deep learning models. The audio module extracts vocal features such as pitch, tone, speech rate, and stress patterns to identify anomalies associated with deceptive behavior. In addition, the language module processes transcribed speech using natural language processing techniques to evaluate sentence structure, sentiment, and semantic inconsistencies. A combination of deep learning architectures, including Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNNs), and transformer-based models, is employed to capture spatial and temporal dependencies in the data. The output of individual modalities is fused using a multimodal fusion strategy to improve the robustness and accuracy of the system. The proposed approach is evaluated on both publicly available datasets and customcollected data to ensure generalizability across different scenarios. Experimental results demonstrate that the multimodal framework significantly outperforms unimodal approaches in terms of accuracy, precision, recall, and F1-score. The system shows promising potential for applications in online interviews, fraud detection, security monitoring, and remote assessments. However, challenges such as data variability, cultural differences, and realtime processing constraints remain. Furthermore, this study highlights important ethical considerations, including privacy concerns, the risk of false accusations, and the need for informed user consent. Future work will focus on improving model explainability, expanding datasets, and deploying the system in real-world environments for enhanced reliability.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{200065,
        author = {KISHAN KANNAUJIYA and AJAY KUMAR PAL and SANDEEP SINGH and ASHUTOSH KUMAR RATHOUR and SAKSHI RASTOGI},
        title = {AI BASED MULTIMODAL LIE DETECTION VIDEO AND AUDIO CALLS},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {12},
        pages = {4069-4079},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=200065},
        abstract = {The rapid growth of online communication platforms such as video conferencing and voice-based interactions has increased the need for reliable and automated deception detection systems. Traditional lie detection techniques, including polygraph tests, rely on physiological signals and often suffer from limited accuracy, lack of scalability, and susceptibility to manipulation. To address these challenges, this research proposes a Multimodal AI-based framework for detecting deception in real-time video and audio calls. The proposed system integrates three primary modalities: visual, acoustic, and linguistic features. The visual module analyzes facial expressions, micro-expressions, and eye movements using computer vision techniques and deep learning models. The audio module extracts vocal features such as pitch, tone, speech rate, and stress patterns to identify anomalies associated with deceptive behavior. In addition, the language module processes transcribed speech using natural language processing techniques to evaluate sentence structure, sentiment, and semantic inconsistencies. A combination of deep learning architectures, including Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNNs), and transformer-based models, is employed to capture spatial and temporal dependencies in the data. The output of individual modalities is fused using a multimodal fusion strategy to improve the robustness and accuracy of the system. The proposed approach is evaluated on both publicly available datasets and customcollected data to ensure generalizability across different scenarios. Experimental results demonstrate that the multimodal framework significantly outperforms unimodal approaches in terms of accuracy, precision, recall, and F1-score. The system shows promising potential for applications in online interviews, fraud detection, security monitoring, and remote assessments. However, challenges such as data variability, cultural differences, and realtime processing constraints remain. Furthermore, this study highlights important ethical considerations, including privacy concerns, the risk of false accusations, and the need for informed user consent. Future work will focus on improving model explainability, expanding datasets, and deploying the system in real-world environments for enhanced reliability.},
        keywords = {Deception Detection, Lie Detection, Multimodal Artificial Intelligence, Video Analysis, Audio Signal Processing, Facial Expression Recognition, Voice Stress Analysis, Natural Language Processing, Deep Learning, Computer Vision, Speech Analysis, Real-Time Detection, Behavioral Analysis, Human-Computer Interaction.},
        month = {May},
        }

Cite This Article

KANNAUJIYA, K., & PAL, A. K., & SINGH, S., & RATHOUR, A. K., & RASTOGI, S. (2026). AI BASED MULTIMODAL LIE DETECTION VIDEO AND AUDIO CALLS. International Journal of Innovative Research in Technology (IJIRT), 12(12), 4069–4079.

Related Articles