Suryavanshi et al. Res. Trends Int. J. Technol. Innov., April - June 2026, 1 (2) : 54-61
1School of Business Management, NMIMS University, Mumbai, India; 2Department of Computer Applications, SRM Institute of Science and Technology, Chennai, India
Article History
Accepted : 03 Jun 2026
Published : 29 Jun 2026
Publication Issue
Volume 1, Issue 2
April - June 2026
Page Number54–61
Microfinance institutions serving thin-file borrowers need default-prediction models that perform well with limited traditional credit history. This paper benchmarks XGBoost, LightGBM and CatBoost against logistic regression on a dataset of 38,500 microloan records incorporating alternative data such as mobile recharge frequency and utility payment regularity. CatBoost achieved the highest AUC of 0.879, followed closely by LightGBM at 0.874, both substantially outperforming logistic regression at 0.792, while categorical-feature handling in CatBoost reduced preprocessing time by an estimated 35 percent.
Keywords - credit scoring, gradient boosting, microfinance, alternative data, default prediction
Thin-file borrowers in microfinance settings often lack the credit bureau history that underpins conventional scorecards, prompting growing interest in alternative behavioural data sources and in gradient boosting methods that handle heterogeneous, sparse feature sets well.
Three gradient boosting variants, XGBoost, LightGBM and CatBoost, were trained on 38,500 microloan records with 41 features spanning demographic, transactional and alternative behavioural data, and benchmarked against an L2-regularised logistic regression baseline using five-fold stratified cross-validation with AUC as the primary metric.
CatBoost achieved the highest AUC at 0.879, followed by LightGBM at 0.874 and XGBoost at 0.869, with all three gradient boosting variants outperforming logistic regression at 0.792. CatBoost native handling of categorical features reduced feature-engineering and preprocessing time by an estimated 35 percent relative to the one-hot encoding pipeline required for the other models.
Gradient boosting methods, particularly CatBoost, offer meaningful accuracy gains over linear baselines for microfinance default prediction while reducing preprocessing overhead. Future work will examine fairness across demographic subgroups.
[1] Prokhorenkova L. et al., CatBoost: Unbiased boosting with categorical features, NeurIPS, 2018. [2] Ke G. et al., LightGBM, NeurIPS, 2017. [3] Bjorkegren D. and Grissen D., Behavior revealed in mobile phone usage predicts credit repayment, World Bank Economic Review, 2020.
© 2026 The Author(s). Published by IJEIA Editorial Office. This is an open access article under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Nitin Suryavanshi, Bhavya Shah (2026). Comparative Analysis of Gradient Boosting Algorithms for Credit Default Prediction in Microfinance. IJEIA, 1(2), 54-61.
Varun Kapoor, Snehal Joshi
Machine LearningKunal Bhatt, Alisha Fernandes
Machine LearningSanjay Kulkarni, Meera Iyer, Deepak Rathi