Detecting Deepfake Videos: A Deep Learning and Neural Network Approach

  • Unique Paper ID: 201747
  • Volume: 12
  • Issue: 12
  • PageNo: 6928-6940
  • Abstract:
  • Recent progress in GenAI has enabled the creation of deepfake videos that are strikingly realistic, capable of reconstructing visuals with minimal distortions, modifying voices, and seamlessly swapping identities. These manipulated videos introduce serious risks to public trust, digital security, political stability, and individual privacy. A major problem for law enforcement and cybersecurity systems is deepfake-based video fraud, which includes identity theft, impersonation schemes, and deceptive political propaganda. With a focus on deep learning architectures, multimodal detection techniques, transformer-based systems, and adversarially resilient frameworks, this review paper investigates cutting-edge deepfake detection techniques. Currently, existing methods such as hybrid CNN-LSTM models [1], Swin-BiLSTM architectures [6], and multimodal audio-visual fusion networks [10], [18] demonstrate considerable capabilities in identifying manipulated frames across spatial and temporal domains. Recent works additionally focus on transformer-based systems like EffiSwinNet [3] and lightweight models such as FakeFormer [17] to improve generalization and reduce computational load for real-time deployments. Furthermore, new research indicates that detectors are vulnerable to adversarial attacks and quality variations, prompting the development of robust systems such as identity masking frameworks and gray-box-resilient detectors. This review summarizes theoretical foundations, technology modeling, empirical performance insights, and highlights key research gaps necessary for reliable deepfake detection systems, addressing issues like cross-dataset robustness and real-time inference constraints.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{201747,
        author = {Dinesh Choudhary and Lalit Hinduja and Parth Tilekar and Omkar Pansare and Dr. Shilpa Khedkar},
        title = {Detecting Deepfake Videos: A Deep Learning and Neural Network Approach},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {12},
        pages = {6928-6940},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=201747},
        abstract = {Recent progress in GenAI has enabled the creation of deepfake videos that are strikingly realistic, capable of reconstructing visuals with minimal distortions, modifying voices, and seamlessly swapping identities. These manipulated videos introduce serious risks to public trust, digital security, political stability, and individual privacy. A major problem for law enforcement and cybersecurity systems is deepfake-based video fraud, which includes identity theft, impersonation schemes, and deceptive political propaganda. With a focus on deep learning architectures, multimodal detection techniques, transformer-based systems, and adversarially resilient frameworks, this review paper investigates cutting-edge deepfake detection techniques. Currently, existing methods such as hybrid CNN-LSTM models [1], Swin-BiLSTM architectures [6], and multimodal audio-visual fusion networks [10], [18] demonstrate considerable capabilities in identifying manipulated frames across spatial and temporal domains. Recent works additionally focus on transformer-based systems like EffiSwinNet [3] and lightweight models such as FakeFormer [17] to improve generalization and reduce computational load for real-time deployments.
Furthermore, new research indicates that detectors are vulnerable to adversarial attacks and quality variations, prompting the development of robust systems such as identity masking frameworks and gray-box-resilient detectors. This review summarizes theoretical foundations, technology modeling, empirical performance insights, and highlights key research gaps necessary for reliable deepfake detection systems, addressing issues like cross-dataset robustness and real-time inference constraints.},
        keywords = {Deepfake detection, deepfake video, generative adversarial networks (GANs), transformer models, computer vision, video forgery detection, multimodal analysis, CNN-LSTM architectures, temporal consistency, spatial feature extraction, adversarial robustness, perceptual fidelity.},
        month = {May},
        }

Cite This Article

Choudhary, D., & Hinduja, L., & Tilekar, P., & Pansare, O., & Khedkar, D. S. (2026). Detecting Deepfake Videos: A Deep Learning and Neural Network Approach. International Journal of Innovative Research in Technology (IJIRT), 12(12), 6928–6940.

Related Articles