Explainable Multimodal Hate Speech Detection Using Cross-Attention Transformer Fusion Networks

  • Unique Paper ID: 205454
  • Volume: 13
  • Issue: 1
  • PageNo: 6846-6854
  • Abstract:
  • The widespread adoption of social media platforms has led to a significant increase in the dissemination of hateful and offensive content through text, images, memes, and other multimodal formats. Conventional hate speech detection systems predominantly rely on textual analysis and often fail to capture the contextual relationships between visual and linguistic information. To address these challenges, this work presents an explainable multimodal hate speech detection framework that combines advanced transformer-based models for comprehensive content understanding. The proposed architecture employs XLM-RoBERTa to extract contextual multilingual textual representations and CLIP Vision Transformer (ViT-B/32) to learn semantic visual features from images. A Cross-Attention Fusion Network is introduced to establish meaningful interactions between textual tokens and visual regions, enabling effective identification of implicit, symbolic, and context-dependent hate expressions. To improve model transparency and support trustworthy decision-making, explainable artificial intelligence techniques, including SHAP and Grad-CAM, are integrated to highlight influential textual elements and image regions contributing to the classification outcome. The framework is evaluated on the MultiOFF benchmark dataset using standard performance metrics. Experimental results demonstrate strong classification capability, achieving high accuracy, precision, recall, F1-score, and Area Under the Curve (AUC), while maintaining interpretability of predictions. The findings indicate that the proposed transformer-based multimodal fusion approach provides a reliable and scalable solution for automated hate speech detection and intelligent content moderation across diverse online environments.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{205454,
        author = {Yaseen Malik and Dr. Meera Alphy},
        title = {Explainable Multimodal Hate Speech Detection Using Cross-Attention Transformer Fusion Networks},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {1},
        pages = {6846-6854},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=205454},
        abstract = {The widespread adoption of social media platforms has led to a significant increase in the dissemination of hateful and offensive content through text, images, memes, and other multimodal formats. Conventional hate speech detection systems predominantly rely on textual analysis and often fail to capture the contextual relationships between visual and linguistic information. To address these challenges, this work presents an explainable multimodal hate speech detection framework that combines advanced transformer-based models for comprehensive content understanding. The proposed architecture employs XLM-RoBERTa to extract contextual multilingual textual representations and CLIP Vision Transformer (ViT-B/32) to learn semantic visual features from images. A Cross-Attention Fusion Network is introduced to establish meaningful interactions between textual tokens and visual regions, enabling effective identification of implicit, symbolic, and context-dependent hate expressions. To improve model transparency and support trustworthy decision-making, explainable artificial intelligence techniques, including SHAP and Grad-CAM, are integrated to highlight influential textual elements and image regions contributing to the classification outcome. The framework is evaluated on the MultiOFF benchmark dataset using standard performance metrics. Experimental results demonstrate strong classification capability, achieving high accuracy, precision, recall, F1-score, and Area Under the Curve (AUC), while maintaining interpretability of predictions. The findings indicate that the proposed transformer-based multimodal fusion approach provides a reliable and scalable solution for automated hate speech detection and intelligent content moderation across diverse online environments.},
        keywords = {Multimodal Hate Speech Detection, Explainable Artificial Intelligence (XAI), Cross-Attention Fusion Network, XLM-RoBERTa, CLIP Vision Transformer, Social Media Content Moderation, Transformer-Based Learning, Deep Learning.},
        month = {June},
        }

Cite This Article

Malik, Y., & Alphy, D. M. (2026). Explainable Multimodal Hate Speech Detection Using Cross-Attention Transformer Fusion Networks. International Journal of Innovative Research in Technology (IJIRT), 13(1), 6846–6854.

Related Articles