Multimodal Analysis for Hate Speech Detection in Memes Using MemeSentinel-XT

  • Unique Paper ID: 204565
  • Volume: 13
  • Issue: 1
  • PageNo: 3129-3135
  • Abstract:
  • The rapid growth of social media platforms has signifi-cantly increased the spread of hateful and offensive content through multimodal memes. Unlike traditional hate speech, modern hateful content combines textual, visual, audio, and video elements, making detection extremely difficult for conventional Natural Language Processing (NLP) systems. Multimodal hate speech detection therefore requires simultaneous understand-ing of images, text, speech signals, and video context. This research paper proposes a multimodal hate speech detection framework named MemeSentinel-XT, designed to analyze image-text relationships within memes using transformer-based multimodal learning architectures. The proposed framework integrates Optical Character Recognition (OCR), contextual text embeddings, visual feature extraction, audio feature analysis, video understanding modules, cross-modal attention mechanisms, and classification layers for accurate hateful content detection. The study evaluates the effectiveness of MemeSentinel-XT using benchmark datasets such as the Facebook Hateful Memes Dataset. Experimental analysis demonstrates that multimodal fusion significantly improves detection accuracy com-pared to unimodal approaches. Furthermore, the pa-per discusses challenges such as contextual ambiguity, sarcasm, bias, ethical concerns, and dataset imbalance in hateful meme detection. This research contributes toward building safer online communities through in-telligent AI-driven moderation systems.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{204565,
        author = {Dikshita Yesambare and Srushti Wele and Sakshi Ragit and Darshan Bhatarkar and Vaijanath Sapkal and Dr. Vanita Buradkar},
        title = {Multimodal Analysis for Hate Speech Detection in Memes Using MemeSentinel-XT},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {1},
        pages = {3129-3135},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=204565},
        abstract = {The rapid growth of social media platforms has signifi-cantly increased the spread of hateful and offensive content through multimodal memes. Unlike traditional hate speech, modern hateful content combines textual, visual, audio, and video elements, making detection extremely difficult for conventional Natural Language Processing (NLP) systems. Multimodal hate speech detection therefore requires simultaneous understand-ing of images, text, speech signals, and video context. This research paper proposes a multimodal hate speech detection framework named MemeSentinel-XT, designed to analyze image-text relationships within memes using transformer-based multimodal learning architectures. The proposed framework integrates Optical Character Recognition (OCR), contextual text embeddings, visual feature extraction, audio feature analysis, video understanding modules, cross-modal attention mechanisms, and classification layers for accurate hateful content detection. The study evaluates the effectiveness of MemeSentinel-XT using benchmark datasets such as the Facebook Hateful Memes Dataset. Experimental analysis demonstrates that multimodal fusion significantly improves detection accuracy com-pared to unimodal approaches. Furthermore, the pa-per discusses challenges such as contextual ambiguity, sarcasm, bias, ethical concerns, and dataset imbalance in hateful meme detection. This research contributes toward building safer online communities through in-telligent AI-driven moderation systems.},
        keywords = {Hate Speech Detection, Multimodal Learning, Memes, Deep Learning, NLP, Computer Vision, Transformer Models, MemeSentinel-XT.},
        month = {June},
        }

Cite This Article

Yesambare, D., & Wele, S., & Ragit, S., & Bhatarkar, D., & Sapkal, V., & Buradkar, D. V. (2026). Multimodal Analysis for Hate Speech Detection in Memes Using MemeSentinel-XT. International Journal of Innovative Research in Technology (IJIRT), 13(1), 3129–3135.

Related Articles