OcuSight: Benchmark Performance vs. Clinical Deployment Reliability in Retinal Multi-Disease Classification

  • Unique Paper ID: 200619
  • Volume: 12
  • Issue: 12
  • PageNo: 2381-2390
  • Abstract:
  • Just having reliable benchmarking statistics do not provide confidence to physicians that an automated retinal screening system will perform predictably. Predictability of performance is crucial for physicians during evaluation of these systems. In this paper, we document our experience developing and iterating on OcuSight—a multi-label classifier for fundus disease on the Retinal Fundus Multi-Disease Image Dataset (RFMiD)—and outline the architectural choices made throughout the development/iteration processes. Two well-established CNN backbones (ResNet-50 and DenseNet-121) were trained in parallel on the RFMiD dataset and compared across three different types of imaging (retinal fundus, brain MRI, and chest x-ray). DenseNet-121 demonstrated superior performance compared to ResNet-50 (macro-F1=0.2023 for DenseNet-121 and 0.1769 for ResNet-50) based on controlled test sets. This result appears to be consistent with the results from numerous studies which have previously demonstrated DenseNet’s superiority in fine- grained and low-data problems. However, once the two models were integrated into the OcuSight pipeline and used to evaluate fundus images (i.e., images not used for training or testing), we found two different types of major errors in the performance of OcuSight based on using DenseNet-121—the first being that healthy retina (i.e., no confirmed pathology) were incorrectly identified as having pathology and second, that pathological retinas (i.e., confirmed pathology) were incorrectly identified as healthy by OcuSight. Based on the magnitude of these two different types of failures, neither of these two error types would be medically acceptable. A detailed investigation into both types of failures was undertaken, and as a result of that investigation, we changed the base architecture of OcuSight from DenseNet- 121 to ResNet-50. The result was that a stable set of outputs (i.e., well-calibrated) were produced; for example, one of the abnormal fundus produced a 86.11% probability of having general pathology and produced only a 12.34% probability of being normal. Supporting clinical findings was accomplished using a Comprehensive Architectural Comparison, A Survey of All 51 Disease Categories in The RFMiD 2.0 Database (Disease Classifications), And A Grad-CAM Saliency Analysis.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{200619,
        author = {Shiva Marath and Soubhik Sadhu and Adnan Nagdiwala and Vaddi Ranga Koushik},
        title = {OcuSight: Benchmark Performance vs. Clinical Deployment Reliability in Retinal Multi-Disease Classification},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {12},
        pages = {2381-2390},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=200619},
        abstract = {Just having reliable benchmarking statistics do not provide confidence to physicians that an automated retinal screening system will perform predictably. Predictability of performance is crucial for physicians during evaluation of these systems. In this paper, we document our experience developing and iterating on OcuSight—a multi-label classifier for fundus disease on the Retinal Fundus Multi-Disease Image Dataset (RFMiD)—and outline the architectural choices made throughout the development/iteration processes. Two well-established CNN backbones (ResNet-50 and DenseNet-121) were trained in parallel on the RFMiD dataset and compared across three different types of imaging (retinal fundus, brain MRI, and chest x-ray). DenseNet-121 demonstrated superior performance compared to ResNet-50 (macro-F1=0.2023 for DenseNet-121 and 0.1769 for ResNet-50) based on controlled test sets. This result appears to be consistent with the results from numerous studies which have previously demonstrated DenseNet’s superiority in fine- grained and low-data problems. However, once the two models were integrated into the OcuSight pipeline and used to evaluate fundus images (i.e., images not used for training or testing), we found two different types of major errors in the performance of OcuSight based on using DenseNet-121—the first being that healthy retina (i.e., no confirmed pathology) were incorrectly identified as having pathology and second, that pathological retinas (i.e., confirmed pathology) were incorrectly identified as healthy by OcuSight. Based on the magnitude of these two different types of failures, neither of these two error types would be medically acceptable. A detailed investigation into both types of failures was undertaken, and as a result of that investigation, we changed the base architecture of OcuSight from DenseNet- 121 to ResNet-50. The result was that a stable set of outputs (i.e., well-calibrated) were produced; for example, one of the abnormal fundus produced a 86.11% probability of having general pathology and produced only a 12.34% probability of being normal. Supporting clinical findings was accomplished using a Comprehensive Architectural Comparison, A Survey of All
51 Disease Categories in The RFMiD 2.0 Database (Disease Classifications), And A Grad-CAM Saliency Analysis.},
        keywords = {ResNet-50; DenseNet-121; Classification of Retinal Diseases; OcuSight; RFMiD; False Positive Rate; False Negative Rate; Multi- Label Classification/Transfer Learning; Diabetic Retinopathy; Grad-CAM; And Clinical Applicability.},
        month = {May},
        }

Cite This Article

Marath, S., & Sadhu, S., & Nagdiwala, A., & Koushik, V. R. (2026). OcuSight: Benchmark Performance vs. Clinical Deployment Reliability in Retinal Multi-Disease Classification. International Journal of Innovative Research in Technology (IJIRT), 12(12), 2381–2390.

Related Articles