Hybrid Model for Scene Text Detection: Integrating Object Detection and Segmentation-Based Approaches

  • Unique Paper ID: 207086
  • Volume: 13
  • Issue: 2
  • PageNo: 4483-4488
  • Abstract:
  • Detection of scene text in real images is still a challenging task in computer vision due to the wide variance of real-world texts in terms of orientation, scale, font types, and other factors. In one-vs-one systems, there is always a compromise between the flexibility of shapes and the speed of prediction. Neither purely regression-based methods nor segmentation-based methods alone can deal with the whole set of variances in real-world cases. This paper presents a novel approach by combining the output results of a YOLO-based object detection branch and a DBNet++-based segmentation branch into a unified pipeline. The proposed weighted NonMaximum Suppression fusion layer unifies the two candidate sets and delivers a final output. Our model has achieved F1scores of 88.0% and 85.9%, respectively, on the ICDAR 2015 and CTW1500 datasets while running at 24 fps. Index Terms—Scene text detection, YOLO, DBNet++, hybrid framework, deep learning, Non-Maximum Suppression, ICDAR 2015, CTW1500

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{207086,
        author = {Darshan M S and keerthi shankar V U and Priyanka Mohan},
        title = {Hybrid Model for Scene Text Detection: Integrating Object Detection and Segmentation-Based Approaches},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {2},
        pages = {4483-4488},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=207086},
        abstract = {Detection of scene text in real images is still a challenging task in computer vision due to the wide variance of real-world texts in terms of orientation, scale, font types, and other factors. In one-vs-one systems, there is always a compromise between the flexibility of shapes and the speed of prediction. Neither purely regression-based methods nor segmentation-based methods alone can deal with the whole set of variances in real-world cases. This paper presents a novel approach by combining the output results of a YOLO-based object detection branch and a DBNet++-based segmentation branch into a unified pipeline. The proposed weighted NonMaximum Suppression fusion layer unifies the two candidate sets and delivers a final output. Our model has achieved F1scores of 88.0% and 85.9%, respectively, on the ICDAR 2015 and CTW1500 datasets while running at 24 fps. Index Terms—Scene text detection, YOLO, DBNet++, hybrid framework, deep learning, Non-Maximum Suppression, ICDAR 2015, CTW1500},
        keywords = {},
        month = {July},
        }

Cite This Article

S, D. M., & U, K. S. V., & Mohan, P. (2026). Hybrid Model for Scene Text Detection: Integrating Object Detection and Segmentation-Based Approaches. International Journal of Innovative Research in Technology (IJIRT). https://doi.org/doi.org/10.64643/IJIRTV13I2-207086-459

Related Articles