AI-Generated Fake Voice Detection Using CNN and Spectrogram Analysis

  • Unique Paper ID: 194298
  • Volume: 12
  • Issue: 11
  • PageNo: 8585-8591
  • Abstract:
  • The rapid advancement of AI-based speech synthesis technologies has enabled the generation of highly realistic synthetic voices, creating critical security vulnerabilities through audio deepfakes. These AI-generated fake voices are increasingly exploited for identity impersonation, financial fraud, and voice-based authentication bypass, making robust detection an urgent cybersecurity necessity. However, existing detection approaches relying on traditional signal processing pipelines, handcrafted acoustic features such as MFCCs and spectral centroid, or recurrent architectures struggle to generalize across unseen spoofing attacks and fail to capture subtle spectral artifacts embedded within synthesized speech signals. Furthermore, prior works underutilize high-resolution visual representations of audio, limiting the exploitation of rich time-frequency information inherent in synthetic speech. To address these limitations, this paper proposes a deep learning framework for detecting AI-generated fake voices using spectrogram analysis combined with a Convolutional Neural Network (CNN) architecture. Raw audio signals are transformed into Mel-spectrogram representations via Short-Time Fourier Transform (STFT), converting temporal speech signals into two-dimensional time-frequency visual representations. These spectrogram images are processed by a custom CNN incorporating multiple stacked convolutional layers, batch normalization, max pooling, and fully connected classification layers to hierarchically learn fine-grained spectral dependencies and spatial frequency patterns. Evaluated on a balanced dataset of 14,000 audio samples (7,000 real and 7,000 fake), the proposed model achieves 99% accuracy, precision, recall, and F1-score, alongside an outstanding ROC-AUC score of 0.99. The framework demonstrates strong generalization capability, effectively capturing subtle spectral irregularities and artifact patterns characteristic of synthesized speech across multiple TTS and voice conversion systems. These results highlight the suitability of CNN-based spectrogram classification as a reliable and computationally efficient approach, thereby contributing to strengthened security in voice-based authentication and digital communication systems.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{194298,
        author = {B. Hem sundar sai reddy and B. Gowtham naidu and B. Prasad and B. Gopi chandu},
        title = {AI-Generated Fake Voice Detection Using CNN and Spectrogram Analysis},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {11},
        pages = {8585-8591},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=194298},
        abstract = {The rapid advancement of AI-based speech synthesis technologies has enabled the generation of highly realistic synthetic voices, creating critical security vulnerabilities through audio deepfakes. These AI-generated fake voices are increasingly exploited for identity impersonation, financial fraud, and voice-based authentication bypass, making robust detection an urgent cybersecurity necessity. However, existing detection approaches relying on traditional signal processing pipelines, handcrafted acoustic features such as MFCCs and spectral centroid, or recurrent architectures struggle to generalize across unseen spoofing attacks and fail to capture subtle spectral artifacts embedded within synthesized speech signals. Furthermore, prior works underutilize high-resolution visual representations of audio, limiting the exploitation of rich time-frequency information inherent in synthetic speech. To address these limitations, this paper proposes a deep learning framework for detecting AI-generated fake voices using spectrogram analysis combined with a Convolutional Neural Network (CNN) architecture. Raw audio signals are transformed into Mel-spectrogram representations via Short-Time Fourier Transform (STFT), converting temporal speech signals into two-dimensional time-frequency visual representations. These spectrogram images are processed by a custom CNN incorporating multiple stacked convolutional layers, batch normalization, max pooling, and fully connected classification layers to hierarchically learn fine-grained spectral dependencies and spatial frequency patterns. Evaluated on a balanced dataset of 14,000 audio samples (7,000 real and 7,000 fake), the proposed model achieves 99% accuracy, precision, recall, and F1-score, alongside an outstanding ROC-AUC score of 0.99. The framework demonstrates strong generalization capability, effectively capturing subtle spectral irregularities and artifact patterns characteristic of synthesized speech across multiple TTS and voice conversion systems. These results highlight the suitability of CNN-based spectrogram classification as a reliable and computationally efficient approach, thereby contributing to strengthened security in voice-based authentication and digital communication systems.},
        keywords = {Classification, CNN, Deep Convolutional, Deepfake, Neural Networks, Spectrogram.},
        month = {April},
        }

Cite This Article

reddy, B. H. S. S., & naidu, B. G., & Prasad, B., & chandu, B. G. (2026). AI-Generated Fake Voice Detection Using CNN and Spectrogram Analysis. International Journal of Innovative Research in Technology (IJIRT), 12(11), 8585–8591.

Related Articles