Hybrid CNN–Transformer Framework for Robust Deepfake Detection

  • Unique Paper ID: 198095
  • Volume: 12
  • Issue: 11
  • PageNo: 14748-14754
  • Abstract:
  • The rapid advancement of generative artificial intelligence has intensified the challenge of distinguishing authentic media from synthetic deepfakes. Building upon prior CNN-based frameworks, this continuation research explores advanced detection strategies that address real-world deployment issues such as adversarial robustness, multimodal manipulations, and interpretability. A hybrid architecture integrating convolutional neural networks with transformer-based attention mechanisms is proposed, alongside multimodal fusion of image, audio, and metadata features. Experimental evaluation across benchmark datasets demonstrates improved generalization, resilience against compression, and enhanced transparency through explainable AI techniques. The findings highlight that scalable, interpretable, and fairness-aware detection systems are essential for practical applications in journalism, law enforcement, and social media moderation, thereby contributing to the broader goal of safeguarding digital trust in the era of synthetic media.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{198095,
        author = {Shreya Gupta and Krish Lal Srivastava and Diksha Asnora and Harshit Mehra},
        title = {Hybrid CNN–Transformer Framework for Robust Deepfake Detection},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {11},
        pages = {14748-14754},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=198095},
        abstract = {The rapid advancement of generative artificial intelligence has intensified the challenge of distinguishing authentic media from synthetic deepfakes. Building upon prior CNN-based frameworks, this continuation research explores advanced detection strategies that address real-world deployment issues such as adversarial robustness, multimodal manipulations, and interpretability. A hybrid architecture integrating convolutional neural networks with transformer-based attention mechanisms is proposed, alongside multimodal fusion of image, audio, and metadata features. Experimental evaluation across benchmark datasets demonstrates improved generalization, resilience against compression, and enhanced transparency through explainable AI techniques. The findings highlight that scalable, interpretable, and fairness-aware detection systems are essential for practical applications in journalism, law enforcement, and social media moderation, thereby contributing to the broader goal of safeguarding digital trust in the era of synthetic media.},
        keywords = {Deepfake Detection, CNN, Transformer, Explainable AI, Multimodal Fusion, Adversarial Robustness},
        month = {April},
        }

Cite This Article

Gupta, S., & Srivastava, K. L., & Asnora, D., & Mehra, H. (2026). Hybrid CNN–Transformer Framework for Robust Deepfake Detection. International Journal of Innovative Research in Technology (IJIRT), 12(11), 14748–14754.

Related Articles