AI-Powered Video-RAG for Interactive E-Learning - Project Implementation

  • Unique Paper ID: 204027
  • Volume: 13
  • Issue: 1
  • PageNo: 1060-1068
  • Abstract:
  • Artificial Intelligence (AI) is transforming digital education by creating smarter and more engaging learning environments. Despite this progress, many online learning platforms still rely heavily on conventional video lectures, which often lack interactivity, personalized guidance, and instant academic support for learners. To address these limitations, this study presents a Personalized AI-Driven Video Retrieval-Augmented Generation (Video-RAG) framework designed to improve the e-learning experience. The proposed framework integrates speech recognition, multimodal retrieval methods, and large language models to enable efficient interaction with educational video content. The Whisper speech recognition model is used to transcribe lecture audio into text, while CLIP-based retrieval mechanisms establish connections between visual and textual data for more accurate information retrieval. Furthermore, a Large Language Model (LLM) generates context-aware and meaningful responses to learner queries. The system also incorporates personalization capabilities that examine learner interactions and adapt responses according to individual learning patterns and preferences. The framework was evaluated using educational video datasets related to fields such as computer science and artificial intelligence. Experimental findings indicate notable improvements in retrieval precision, response quality, and learner engagement compared to traditional video-based learning approaches. Feedback from users further highlights enhanced concept understanding, quicker access to relevant information, and a more effective learning experience through interactive question-answering and adaptive support features. This study demonstrates the potential of combining multimodal AI techniques with retrieval-augmented generation to develop more intelligent, adaptive, and learner-focused e-learning platforms. The proposed approach also provides opportunities for future advancements in areas such as personalized education, multilingual learning systems, and AI-assisted tutoring technologies.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{204027,
        author = {RUSHIKA  PAGA and Pragati Waghmare and Prajwal Berad and Siddhesh Bhor},
        title = {AI-Powered Video-RAG for Interactive E-Learning - Project Implementation},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {1},
        pages = {1060-1068},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=204027},
        abstract = {Artificial Intelligence (AI) is transforming digital education by creating smarter and more engaging learning environments. Despite this progress, many online learning platforms still rely heavily on conventional video lectures, which often lack interactivity, personalized guidance, and instant academic support for learners. To address these limitations, this study presents a Personalized AI-Driven Video Retrieval-Augmented Generation (Video-RAG) framework designed to improve the e-learning experience.
The proposed framework integrates speech recognition, multimodal retrieval methods, and large language models to enable efficient interaction with educational video content. The Whisper speech recognition model is used to transcribe lecture audio into text, while CLIP-based retrieval mechanisms establish connections between visual and textual data for more accurate information retrieval. Furthermore, a Large Language Model (LLM) generates context-aware and meaningful responses to learner queries. The system also incorporates personalization capabilities that examine learner interactions and adapt responses according to individual learning patterns and preferences.
The framework was evaluated using educational video datasets related to fields such as computer science and artificial intelligence. Experimental findings indicate notable improvements in retrieval precision, response quality, and learner engagement compared to traditional video-based learning approaches. Feedback from users further highlights enhanced concept understanding, quicker access to relevant information, and a more effective learning experience through interactive question-answering and adaptive support features.
This study demonstrates the potential of combining multimodal AI techniques with retrieval-augmented generation to develop more intelligent, adaptive, and learner-focused e-learning platforms. The proposed approach also provides opportunities for future advancements in areas such as personalized education, multilingual learning systems, and AI-assisted tutoring technologies.},
        keywords = {},
        month = {June},
        }

Cite This Article

PAGA, R. ., & Waghmare, P., & Berad, P., & Bhor, S. (2026). AI-Powered Video-RAG for Interactive E-Learning - Project Implementation. International Journal of Innovative Research in Technology (IJIRT), 13(1), 1060–1068.

Related Articles