Plagiarism Detection System with Source Identification and Text Highlighting

  • Unique Paper ID: 209075
  • Volume: 13
  • Issue: 5
  • PageNo: 560-567
  • Abstract:
  • Copying text from the web has become effortless, and teachers, journal editors and publishers now receive a steady flow of submissions whose originality is hard to judge by eye. Most checkers return a lone percentage, leaving the reader to guess which lines were borrowed, from where, and whether the borrowing is verbatim or reworded. This paper describes a text plagiarism checker built to answer those three questions. A submitted file (PDF, DOCX or TXT) or pasted text is cleaned and split into sentences with NLTK; every sentence is converted into a semantic vector with Sentence Transformers; and the vectors are matched against a SQLite store of reference documents and against relevant web pages. Sentences whose best match crosses a similarity threshold are credited to the closest source, painted in distinct colors inside the submitted text, and counted toward an overall percentage. The findings are packaged in a downloadable report. The prototype is written in Python and served through Streamlit or Flask. The paper documents the design, the output format, the expected behavior and a plan for evaluation; measured benchmark figures are reserved for future work.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{209075,
        author = {Sreelal P Surendran and Abhinav P and KM Vignesh and Shivakrishna M and Shruthi D},
        title = {Plagiarism Detection System with Source Identification and Text Highlighting},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {5},
        pages = {560-567},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=209075},
        abstract = {Copying text from the web has become effortless, and teachers, journal editors and publishers now receive a steady flow of submissions whose originality is hard to judge by eye. Most checkers return a lone percentage, leaving the reader to guess which lines were borrowed, from where, and whether the borrowing is verbatim or reworded. This paper describes a text plagiarism checker built to answer those three questions. A submitted file (PDF, DOCX or TXT) or pasted text is cleaned and split into sentences with NLTK; every sentence is converted into a semantic vector with Sentence Transformers; and the vectors are matched against a SQLite store of reference documents and against relevant web pages. Sentences whose best match crosses a similarity threshold are credited to the closest source, painted in distinct colors inside the submitted text, and counted toward an overall percentage. The findings are packaged in a downloadable report. The prototype is written in Python and served through Streamlit or Flask. The paper documents the design, the output format, the expected behavior and a plan for evaluation; measured benchmark figures are reserved for future work.},
        keywords = {Plagiarism Detection, Source Identification, Text Highlighting, Natural Language Processing, Sentence Transformers, Semantic Similarity, Report Generation.},
        month = {October},
        }

Cite This Article

Surendran, S. P., & P, A., & Vignesh, K., & M, S., & D, S. (2026). Plagiarism Detection System with Source Identification and Text Highlighting. International Journal of Innovative Research in Technology (IJIRT), 13(5), 560–567.

Related Articles