Assessing the Performance of Large Language Models for Coreference Resolution

  • Unique Paper ID: 201356
  • Volume: 12
  • Issue: 12
  • PageNo: 3514-3520
  • Abstract:
  • Coreference resolution is a critical task in natural language processing, determining when multiple expressions in text refer to the same entity. This ca-pability is particularly vital in question answering (QA) systems, that require pre-cise interpretation of textual context. This paper shows usage of SpanBERT and LLaMA model to enhance reference resolution accuracy and processing efficiency. A hybrid approach involving LLaMA 3.1 with Prompt Engineering achieves significant gains over traditional models like SpanBERT, reaching up to 80.48% accuracy on the GAP dataset. These improvements highlight the po-tential of integrating advanced LLMs in real-world information retrieval applications. The modular architecture of the proposed system allows easy integration into existing QA pipelines. Moreover, by combining zero-shot capabilities of LLMs with deterministic fallback mechanisms, the framework can offer adaptability to domain-specific scenarios such as biomedical QA, conversational data, and multilingual document understanding.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{201356,
        author = {Shivam Tiwari and Pradnya Gotmare and Manish potey},
        title = {Assessing the Performance of Large Language Models for Coreference Resolution},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {12},
        pages = {3514-3520},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=201356},
        abstract = {Coreference resolution is a critical task in natural language processing, determining when multiple expressions in text refer to the same entity. This ca-pability is particularly vital in question answering (QA) systems, that require pre-cise interpretation of textual context. This paper shows usage of SpanBERT and LLaMA model to enhance reference resolution accuracy and processing efficiency. A hybrid approach involving LLaMA 3.1 with Prompt Engineering achieves significant gains over traditional models like SpanBERT, reaching up to 80.48% accuracy on the GAP dataset. These improvements highlight the po-tential of integrating advanced LLMs in real-world information retrieval applications. The modular architecture of the proposed system allows easy integration into existing QA pipelines. Moreover, by combining zero-shot capabilities of LLMs with deterministic fallback mechanisms, the framework can offer adaptability to domain-specific scenarios such as biomedical QA, conversational data, and multilingual document understanding.},
        keywords = {Natural Language Processing, Coreference Resolution, Question Answering, Large Language Models, GAP Dataset.},
        month = {May},
        }

Cite This Article

Tiwari, S., & Gotmare, P., & potey, M. (2026). Assessing the Performance of Large Language Models for Coreference Resolution. International Journal of Innovative Research in Technology (IJIRT), 12(12), 3514–3520.

Related Articles