Predicting AI Specialization Domains from Practitioner Skill Profiles: A Ten-Class Multiclass Classification Study with Gradient Boosting and SHAP Explainability

  • Unique Paper ID: 205545
  • Volume: 13
  • Issue: 1
  • PageNo: 8520-8528
  • Abstract:
  • The rapid expansion of artificial intelligence as a professional discipline has created an urgent need for data-driven career guidance systems capable of matching practitioners to specialization domains based on measurable skill and experience profiles. This study frames AI domain recommendation as a supervised ten-class classification problem and evaluates five state-of-the-art classifiers Logistic Regression, Random Forest, XG Boost, Light GBM, and Cat Boost on a dataset of 30,000 AI practitioner profiles across 30 features spanning technical skill scores, experience metrics, productivity indicators, and tooling preferences. Rigorous experimental controls include explicit exclusion of target-leaking domain labels, stratified 60/20/20 train-validation-test splitting, median imputation for three low-missingness columns, and 5-fold stratified cross-validation. Light GBM achieves the highest performance with test accuracy of 83.67%, macro-F1 of 0.7050, weighted F1 of 0.8250, and macro-ROC-AUC of 0.9352. SHAP-based explainability identifies largest model parameters million, skill consistency score, and mathematics for ai score as the dominant global predictors. Severe class imbalance between Generative AI (21.4%) and Reinforcement Learning (1.4%) drives differential per-class performance, with Finance AI (F1 = 0.129) and Healthcare AI (F1 = 0.198) remaining severely under-predicted despite class-weighted training. This work establishes the first replicable, leakage-free methodology and performance baseline for AI specialization domain recommendation.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{205545,
        author = {Ranveer Singh and Aryan Somanna and Sidhesh Mishra and Shrey Jauhari and Abhijeet Pandey and Ridhuvarshini},
        title = {Predicting AI Specialization Domains from Practitioner Skill Profiles: A Ten-Class Multiclass Classification Study with Gradient Boosting and SHAP Explainability},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {13},
        number = {1},
        pages = {8520-8528},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=205545},
        abstract = {The rapid expansion of artificial intelligence as a professional discipline has created an urgent need for data-driven career guidance systems capable of matching practitioners to specialization domains based on measurable skill and experience profiles. This study frames AI domain recommendation as a supervised ten-class classification problem and evaluates five state-of-the-art classifiers Logistic Regression, Random Forest, XG Boost, Light GBM, and Cat Boost on a dataset of 30,000 AI practitioner profiles across 30 features spanning technical skill scores, experience metrics, productivity indicators, and tooling preferences. Rigorous experimental controls include explicit exclusion of target-leaking domain labels, stratified 60/20/20 train-validation-test splitting, median imputation for three low-missingness columns, and 5-fold stratified cross-validation. Light GBM achieves the highest performance with test accuracy of 83.67%, macro-F1 of 0.7050, weighted F1 of 0.8250, and macro-ROC-AUC of 0.9352. SHAP-based explainability identifies largest model parameters million, skill consistency score, and mathematics for ai score as the dominant global predictors. Severe class imbalance between Generative AI (21.4%) and Reinforcement Learning (1.4%) drives differential per-class performance, with Finance AI (F1 = 0.129) and Healthcare AI (F1 = 0.198) remaining severely under-predicted despite class-weighted training. This work establishes the first replicable, leakage-free methodology and performance baseline for AI specialization domain recommendation.},
        keywords = {AI specialization prediction; multiclass classification; LightGBM; SHAP explainability; skill profile analysis; gradient boosting; career recommendation; class imbalance},
        month = {June},
        }

Cite This Article

Singh, R., & Somanna, A., & Mishra, S., & Jauhari, S., & Pandey, A., & Ridhuvarshini, (2026). Predicting AI Specialization Domains from Practitioner Skill Profiles: A Ten-Class Multiclass Classification Study with Gradient Boosting and SHAP Explainability. International Journal of Innovative Research in Technology (IJIRT), 13(1), 8520–8528.

Related Articles