Vision Mamba Network For Tuberculosis Detection From Chest X-Rays

  • Unique Paper ID: 208677
  • Volume: 13
  • Issue: 4
  • PageNo: 2478-2497
  • Abstract:
  • Detecting Tuberculosis (TB) in chest radiographs is not a simple problem for a model because it has to notice very small, localized lesions while also keeping the overall context of both lung fields, and the aim here is to address this local-plus-global challenge. In this paper, MSVMamba-TB is a hierarchically nested multi-scale Vision Mamba network classifier of 1 channel chest X-ray (CXR) into normal and TB defined in binary classification. We compared it with a deep learning technique, namely Vision Transformer (PaperViT), and two other approaches: ConvNeXtV2-Pico and EfficientNet-B3. Not only that, but all the models were also used in conjunction, fine-tuned, and verified on the same computer with the same specs and setup to keep the comparison for every model as similar as possible. In our proposed model, a multi-scale two-dimensional selective scan, a convolutional feed-forward block, squeeze-and-excitation channel weighting, global pooling, and a small binary head are put together in one architecture. Experiment A contained a total of 4, 200 patient-level records, of which 3, 500 were normal and 700 were tuberculosis. The data distribution was an 80/10/10 split, whereas Experiment B consisted of 7, 000 balanced records with a 70/15/15 data split. Both scenarios have the same preprocessing, augmentation, sampling, loss, optimizer, scheduler, early stopping, and other model features, and no model is given a separate advantage in terms of training. In Experiment A the model obtained accuracy of 99.05%, a sensitivity of 97.14%, and precision of 97.14%, with the F1-score of 97.14% being the experimental results. Experiment B was able to reach an accuracy of 98.38%, a precision of 99.42%, a sensitivity of 97.33%, an F1-score of 98.36%, and the highest score was on the AUC at 0.9992. The model proposed here ranked top in both accuracy and F1-score and made fewer mistakes. Therefore, our proposed model aims to have good prediction performance and a small size. Other than that, this is a very small model with only 7.012M trainable parameters, which, relative to PaperViT with 13.527 M, is only 48.17%. Detailed epoch records, confusion matrices, ROC curves, error analysis, activation maps, and cases where failures were made are presented for all four networks. Hence, the results are not only reported as a single accuracy number. Moreover, these results indicate that MSVMamba-TB is an encouraging candidate for fast tuberculosis screening, which is useful for computer-aided TB screening.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{208677,
        author = {GOLI HARSHAVARDHAN and I. Kullayamma},
        title = {Vision Mamba Network For Tuberculosis Detection From Chest X-Rays},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {4},
        pages = {2478-2497},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=208677},
        abstract = {Detecting Tuberculosis (TB) in chest radiographs is not a simple problem for a model because it has to notice very small, localized lesions while also keeping the overall context of both lung fields, and the aim here is to address this local-plus-global challenge. In this paper, MSVMamba-TB is a hierarchically nested multi-scale Vision Mamba network classifier of 1 channel chest X-ray (CXR) into normal and TB defined in binary classification. We compared it with a deep learning technique, namely Vision Transformer (PaperViT), and two other approaches: ConvNeXtV2-Pico and EfficientNet-B3. Not only that, but all the models were also used in conjunction, fine-tuned, and verified on the same computer with the same specs and setup to keep the comparison for every model as similar as possible. In our proposed model, a multi-scale two-dimensional selective scan, a convolutional feed-forward block, squeeze-and-excitation channel weighting, global pooling, and a small binary head are put together in one architecture. Experiment A contained a total of 4, 200 patient-level records, of which 3, 500 were normal and 700 were tuberculosis. The data distribution was an 80/10/10 split, whereas Experiment B consisted of 7, 000 balanced records with a 70/15/15 data split. Both scenarios have the same preprocessing, augmentation, sampling, loss, optimizer, scheduler, early stopping, and other model features, and no model is given a separate advantage in terms of training. In Experiment A the model obtained accuracy of 99.05%, a sensitivity of 97.14%, and precision of 97.14%, with the F1-score of 97.14% being the experimental results. Experiment B was able to reach an accuracy of 98.38%, a precision of 99.42%, a sensitivity of 97.33%, an F1-score of 98.36%, and the highest score was on the AUC at 0.9992. The model proposed here ranked top in both accuracy and F1-score and made fewer mistakes. Therefore, our proposed model aims to have good prediction performance and a small size. Other than that, this is a very small model with only 7.012M trainable parameters, which, relative to PaperViT with 13.527 M, is only 48.17%. Detailed epoch records, confusion matrices, ROC curves, error analysis, activation maps, and cases where failures were made are presented for all four networks. Hence, the results are not only reported as a single accuracy number. Moreover, these results indicate that MSVMamba-TB is an encouraging candidate for fast tuberculosis screening, which is useful for computer-aided TB screening.},
        keywords = {Chest X-ray, Computer-aided detection, ConvNextV2, Deep learning, EfficientNet, Multi-Scale Scanning, Selective State-Space Model, Tuberculosis, Vision Mamba, Vision Transformer.},
        month = {September},
        }

Cite This Article

HARSHAVARDHAN, G., & Kullayamma, I. (2026). Vision Mamba Network For Tuberculosis Detection From Chest X-Rays. International Journal of Innovative Research in Technology (IJIRT). https://doi.org/10.64643/IJIRTV13I4-208677-459

Related Articles