VISION TRANSFORMERS FOR CROSS-DOMAIN PRESENTATION ATTACK DETECTION

  • Unique Paper ID: 204419
  • Volume: 13
  • Issue: 1
  • PageNo: 2727-2734
  • Abstract:
  • Facial biometric systems are essential for modern security and authentication, but they remain highly vulnerable to physical spoofing methods, such as printed photos and screen replays. Because deepfake detection and physical presentation attack detection (PAD) both focus on identifying fabricated faces, recent research has hypothesized that deepfake models might naturally generalize to physical PAD without needing task-specific retraining. This study tests that hypothesis using the Cross-Domain Generalization Framework for Presentation Attack Detection (CDGF-PAD). Within this framework, Vision Transformer (ViT) and Convolutional Neural Network (CNN) architectures were evaluated across both in-domain and cross-domain scenarios. ISO/IEC 30107-3 metrics (APCER, BPCER, ACER, and AUC) were used to evaluate the performance of models. These models were trained on a subset of CelebA-Spoof and tested on the CASIA-FASD dataset. Our results show that ViT models trained solely on deepfakes produce near-random accuracy at zero-shot physical PAD. On the other hand, ViTs show significantly greater cross-domain generalization than CNNs when exposed to dataset shifts when explicitly trained on PAD data. In the conclusion, this study shows that physical PAD and deepfake detection are really different problems that require separate training.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{204419,
        author = {Raghupatruni Venkata Harsha Vardhan},
        title = {VISION TRANSFORMERS FOR CROSS-DOMAIN PRESENTATION ATTACK DETECTION},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {1},
        pages = {2727-2734},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=204419},
        abstract = {Facial biometric systems are essential for modern security and authentication, but they remain highly vulnerable to physical spoofing methods, such as printed photos and screen replays. Because deepfake detection and physical presentation attack detection (PAD) both focus on identifying fabricated faces, recent research has hypothesized that deepfake models might naturally generalize to physical PAD without needing task-specific retraining. This study tests that hypothesis using the Cross-Domain Generalization Framework for Presentation Attack Detection (CDGF-PAD). Within this framework, Vision Transformer (ViT) and Convolutional Neural Network (CNN) architectures were evaluated across both in-domain and cross-domain scenarios. ISO/IEC 30107-3 metrics (APCER, BPCER, ACER, and AUC) were used to evaluate the performance of models. These models were trained on a subset of CelebA-Spoof and tested on the CASIA-FASD dataset. Our results show that ViT models trained solely on deepfakes produce near-random accuracy at zero-shot physical PAD. On the other hand, ViTs show significantly greater cross-domain generalization than CNNs when exposed to dataset shifts when explicitly trained on PAD data. In the conclusion, this study shows that physical PAD and deepfake detection are really different problems that require separate training.},
        keywords = {Biometric Security, Cross-Domain Generalization, Deepfake Detection, Face Anti-Spoofing, Presentation Attack Detection, Vision Transformers.},
        month = {June},
        }

Cite This Article

Vardhan, R. V. H. (2026). VISION TRANSFORMERS FOR CROSS-DOMAIN PRESENTATION ATTACK DETECTION. International Journal of Innovative Research in Technology (IJIRT), 13(1), 2727–2734.

Related Articles