Machine Learning
International Journal of Engineering Innovation and Advancement An International Peer-Reviewed, Refereed & Open-Access Journal
ISSN : ISSN (Online): Applied For
Available Online at : https://bgmiliteofficial.com/
doi : https://doi.org/10.5555/ijeia.2026.v1i2.017

Suryavanshi et al. Res. Trends Int. J. Technol. Innov., April - June 2026, 1 (2) : 54-61

Comparative Analysis of Gradient Boosting Algorithms for Credit Default Prediction in Microfinance

Nitin Suryavanshi1, Bhavya Shah2

1School of Business Management, NMIMS University, Mumbai, India; 2Department of Computer Applications, SRM Institute of Science and Technology, Chennai, India

Article Info

Article History Accepted : 03 Jun 2026
Published : 29 Jun 2026

Publication Issue Volume 1, Issue 2
April - June 2026

Page Number54–61

Abstract

Microfinance institutions serving thin-file borrowers need default-prediction models that perform well with limited traditional credit history. This paper benchmarks XGBoost, LightGBM and CatBoost against logistic regression on a dataset of 38,500 microloan records incorporating alternative data such as mobile recharge frequency and utility payment regularity. CatBoost achieved the highest AUC of 0.879, followed closely by LightGBM at 0.874, both substantially outperforming logistic regression at 0.792, while categorical-feature handling in CatBoost reduced preprocessing time by an estimated 35 percent.

Keywords - credit scoring, gradient boosting, microfinance, alternative data, default prediction

I. INTRODUCTION

Thin-file borrowers in microfinance settings often lack the credit bureau history that underpins conventional scorecards, prompting growing interest in alternative behavioural data sources and in gradient boosting methods that handle heterogeneous, sparse feature sets well.

II. METHODOLOGY

Three gradient boosting variants, XGBoost, LightGBM and CatBoost, were trained on 38,500 microloan records with 41 features spanning demographic, transactional and alternative behavioural data, and benchmarked against an L2-regularised logistic regression baseline using five-fold stratified cross-validation with AUC as the primary metric.

III. RESULTS AND EVALUATION

CatBoost achieved the highest AUC at 0.879, followed by LightGBM at 0.874 and XGBoost at 0.869, with all three gradient boosting variants outperforming logistic regression at 0.792. CatBoost native handling of categorical features reduced feature-engineering and preprocessing time by an estimated 35 percent relative to the one-hot encoding pipeline required for the other models.

IV. CONCLUSION

Gradient boosting methods, particularly CatBoost, offer meaningful accuracy gains over linear baselines for microfinance default prediction while reducing preprocessing overhead. Future work will examine fairness across demographic subgroups.

V. REFERENCES

[1] Prokhorenkova L. et al., CatBoost: Unbiased boosting with categorical features, NeurIPS, 2018. [2] Ke G. et al., LightGBM, NeurIPS, 2017. [3] Bjorkegren D. and Grissen D., Behavior revealed in mobile phone usage predicts credit repayment, World Bank Economic Review, 2020.

© 2026 The Author(s). Published by IJEIA Editorial Office. This is an open access article under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

Cite this article

Nitin Suryavanshi, Bhavya Shah (2026). Comparative Analysis of Gradient Boosting Algorithms for Credit Default Prediction in Microfinance. IJEIA, 1(2), 54-61.

Related Papers