A Hybrid Deep Learning Framework for AI-Generated Image Detection Using CNN and Vision Transformer

  • Unique Paper ID: 206938
  • Volume: 13
  • Issue: 2
  • PageNo: 4118-4123
  • Abstract:
  • Artificial Intelligence has made a lot of progress in the few years. This means it can now create pictures that look very real. You can find these Artificial Intelligence generated images on media and other online platforms. While this technology is useful it also causes some prob-lems like information and copyright issues. So it is very important to be able to tell if a picture is real or made by Artificial Intelligence. This is a challenge for people who work with computers and investigate digital crimes. To solve this problem we are suggesting a way of using deep learning that combines two methods called Efficient-Net and Vision Transformer for detecting Artificial Intel-ligence generated images. EfficientNet looks at the details in a picture like edges and textures. Vision Transformer looks at the picture and understands how everything fits together. We put these two things together to make our system better at telling if a picture is real or not. We also use something called Grad-CAM to help us understand why our system makes decisions. This makes our system more transparent and trustworthy. We started by testing a version of our system using EfficientNet. We used a set of pictures that were balanced between Artificial Intelligence generated images. Our simple system was able to identify the pictures about 75.5% of the time. This shows that deep learning can be very useful for this task. Now we will test our system on a larger set of pictures to make it even better. Our system could be used in areas, like digital forensics checking social media content, cybersecurity and detecting fake content. Artificial Intelligence generated images are a problem and our system can help solve it.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{206938,
        author = {Vivek Kumar Khatana},
        title = {A Hybrid Deep Learning Framework for AI-Generated Image Detection Using CNN and Vision Transformer},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {2},
        pages = {4118-4123},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=206938},
        abstract = {Artificial Intelligence has made a lot of progress in the few years. This means it can now create pictures that look very real. You can find these Artificial Intelligence generated images on media and other online platforms. While this technology is useful it also causes some prob-lems like information and copyright issues. So it is very important to be able to tell if a picture is real or made by Artificial Intelligence. This is a challenge for people who work with computers and investigate digital crimes. To solve this problem we are suggesting a way of using deep learning that combines two methods called Efficient-Net and Vision Transformer for detecting Artificial Intel-ligence generated images. EfficientNet looks at the details in a picture like edges and textures. Vision Transformer looks at the picture and understands how everything fits together. We put these two things together to make our system better at telling if a picture is real or not. We also use something called Grad-CAM to help us understand why our system makes decisions. This makes our system
more transparent and trustworthy.
We started by testing a version of our system using EfficientNet. We used a set of pictures that were balanced between Artificial Intelligence generated images. Our simple system was able to identify the pictures about 75.5% of the time. This shows that deep learning can be very useful for this task. Now we will test our system on a larger set of pictures to make it even better. Our system could be used in areas, like digital forensics checking social media content, cybersecurity and detecting fake content. Artificial Intelligence generated images are a problem and our system can help solve it.},
        keywords = {},
        month = {July},
        }

Cite This Article

Khatana, V. K. (2026). A Hybrid Deep Learning Framework for AI-Generated Image Detection Using CNN and Vision Transformer. International Journal of Innovative Research in Technology (IJIRT), 13(2), 4118–4123.

Related Articles