Journal of Business Management

Journal of Business Management

Predicting Bank Customer Value Using Machine Learning: A Novel Approach Based on Adaptive Synthetic Sampling and Feature Importance

Document Type : Research Paper

Authors
1 Prof., Department of Industrial Management, Faculty of Industrial Management and Technology, College of Management, University of Tehran, Tehran, Iran.
2 Ph.D. Candidate, Department of Industrial Management, Faculty of Industrial Management and Technology, College of Management, University of Tehran, Tehran, Iran.
3 Ph.D. Candidate, Department of Industrial Management, Kish International Campus, University of Tehran, Tehran, Iran.
Abstract
Objective
Accurate prediction of customer value in the banking industry is one of the fundamental challenges that can contribute to optimal decision-making in customer management and resource allocation. This study aims to develop a comprehensive approach for predicting the value of banking customers. The primary focus of this research is on addressing the challenge of imbalanced data, improving the performance of machine learning models, and selecting key features that are effective in predicting customer value for real-world applications in banking environments.
Methodology
This study used a dataset from a bank comprising 2,000 customers and 14 features related to customer transactions and activities. Data preprocessing was first performed, followed by feature selection, data imbalance handling, and the application of the Adaptive Synthetic Sampling (ADASYN) technique. Correlation analysis and feature importance based on the Random Forest algorithm were subsequently employed to refine the feature selection process. Through this procedure, highly correlated variables were identified, and the final set of relevant features was selected. Subsequently, 11 machine learning algorithms, including CatBoost, XGBoost, Random Forest, LightGBM, and both linear and nonlinear models, were applied to predict customer value. To improve model performance, Optuna was used for hyperparameter tuning, and five-fold cross-validation was employed to ensure robust model evaluation. Model performance was assessed using four evaluation metrics: accuracy, precision, recall, and the F1-score.

Findings
The results showed that ensemble learning-based algorithms provided the best performance in predicting customer value. The CatBoost model, with an F1 Score of 0.9324 and an accuracy of 0.909, was identified as the best-performing model. This model achieved a proper balance between precision and recall, with a precision of 0.9677 and a recall of 0.8998 in predicting valuable customers. The XGBoost and Random Forest models also demonstrated similar performance to CatBoost, with F1 Scores of 0.9322 and 0.932, respectively. The use of a combined approach for feature selection and the application of the ADASYN method for data balancing played a significant role in improving the performance of these models.

Conclusion
These results show that a different approach to data preprocessing with the help of the ADASYN algorithm in combination with modern machine learning methods can positively affect the effectiveness of models predicting customer value. The selection of correlated variables and the feature importance based on Random Forest was important in improving the general performance of the models. This revolution allowed strengthening the work of models through the elimination of features and information that had less impact on the final decisions, making the latter more precise. Based on the results of its evaluation, it can be concluded that ensemble learning models, including CatBoost, XGBoost, and Random Forest, are the most appropriate for banking settings because of their efficiency and effectiveness in dealing with large-scale, complex, and imbalanced datasets. Thus, this study extends previous research by addressing the challenges of imbalanced data and feature selection to enhance customer value prediction and customer management in the banking sector, thereby contributing to the development of an effective approach to this problem. The findings provide useful insights for identifying high-value bank customers and for developing more effective customer retention and service strategies.
Keywords
Subjects

References
Aldi, F., Nozomi, I., Sentosa, R. B., & Junaidi, A. (2023). Machine learning to identify monkey pox disease. Sinkron: jurnal dan penelitian teknik informatika, 7(3), 1335-1347.
Amirhasankhan, H., Toloie Eshlaghy, A., Radfar, R. & Pourebrahimi, A. (2024). Presenting a Hybrid Model based on the Machine Learning for the Classification of Banking and Insurance Industry Common Customers, Journal of Productivity Management, 18(68), 53-80. magiran.com/p2710297 (in Persian)
Ashraf, R. (2024). Bank Customer Churn Prediction Using Machine Learning Framework. Journal of Applied Finance & Banking, 14(4), 65-109, 65–109. https://doi.org/10.47260/jafb/1445
Awad, M. & Khanna, R. (2015). Support vector machines for classification. In Apress eBooks (pp. 39–66). https: //doi.org/10.1007/978-1-4302-5990-9_3
Bentéjac, C., Csörgő, A. & Martínez-Muñoz, G. (2020). A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3), 1937–1967. https: //doi.org/10.1007/s10462-020-09896-5
Chehreh, S. & Sarabadani, A. (2024). A model based on random forest algorithm and Jaya optimization to predict bank customer churn. Journal of Engineering Management and Soft Computing, 9(2), 132-148. doi: 10.22091/jemsc.2024.9541.1174 (in Persian)
Conn, D., Ngun, T., Li, G. & Ramirez, C. M. (2019). Fuzzy Forests: Extending random Forest feature selection for Correlated, High-Dimensional Data. Journal of Statistical Software, 91(9). https: //doi.org/10.18637/jss.v091.i09
Dewi, O. I. P., Santiko, V. N., Safa, S. W. P., Gunawan, A. A. S. & Setiawan, K. E. (2024). Machine Learning for Imbalanced Data in Telecom Churn Classification. 2024 International Conference on Information Technology Research and Innovation (ICITRI), 30–35. https: //doi.org/10.1109/icitri62858.2024.10699226
Dias, J., Godinho, P. & Torres, P. (2020). Machine learning for customer churn prediction in retail banking. In Lecture notes in computer science (pp. 576–589). https: //doi.org/10.1007/978-3-030-58808-3_42
Dorogush, A. V., Ershov, V. & Gulin, A. (2018). CatBoost: gradient boosting with categorical features support. arXiv (Cornell University). https: //doi.org/10.48550/arxiv.1810.11363
Galal, M., Rady, S. & Aref, M. (2022). Enhancing Customer Churn Prediction in Digital Banking using Ensemble Modeling. 2022 4th Novel Intelligent and Leading Emerging Sciences Conference (NILES), 1, 21–25. https: //doi.org/10.1109/niles56402.2022.9942408
Gholamian, M. & Mozafari, A. (2016). Prediction of Bank Customer Value based on RFM Model Using Improved Decision Tree to Reduce the Maximum Required Memory. Business Intelligence Management Studies, 5(17), 93-121. https://doi.org/10.22054/ims.2016.6993 (in Persian)
Guido, R., Ferrisi, S., Lofaro, D. & Conforti, D. (2024). An Overview on the advancements of support vector machine models in healthcare Applications: a review. Information, 15(4), 235. https: //doi.org/10.3390/info15040235
Halder, R. K., Uddin, M. N., Uddin, M. A., Aryal, S. & Khraisat, A. (2024). Enhancing K-nearest neighbor algorithm: a comprehensive review and performance analysis of modifications. Journal of Big Data, 11(1). https: //doi.org/10.1186/s40537-024-00973-y
Hodge, V. & Austin, J. (2004). A survey of outlier detection methodologies. Artificial intelligence review, 22(2), 85-126.
Hu, N. W., Hu, N. W. & Maybank, S. (2008). ADABoost-Based Algorithm for network Intrusion Detection. IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics), 38(2), 577–583. https: //doi.org/10.1109/tsmcb.2007.914695
Ilyas, S., Zia, S., Muneer, U., Letchmunan, S. & Un, Z. (2020). Predicting the Future Transaction from Large and Imbalanced Banking Dataset. International Journal of Advanced Computer Science and Applications, 11(1). https: //doi.org/10.14569/ijacsa.2020.0110134
Iranzad, R. & Liu, X. (2024). A review of random forest-based feature selection methods for data science education and applications. International Journal of Data Science and Analytics. https: //doi.org/10.1007/s41060-024-00509-w
Jafari Eskandari, M. & Rohii, M. (2017). Credit Risk Management of Banking Customers Using Support Vector Machine Optimized by Genetic Algorithm with Data Mining Approach. Journal of Asset Management and Financing, 5(4), 17-32. doi: 10.22108/amf.2017.21191 (in Persian)
Jafarnejad, A., Rezasoltani, A. & Khani, A.M. (2024). Unleashing the Power of Ensemble Learning: Predicting National Ranks in Iran’s University Entrance Examination. Industrial Management Journal, 16(3), 457-481. https: //doi.org/10.22059/imj.2024.381521.1008178 (in Persian)
Kaisar, S. & Sifat, S. T. (2023). Explainable Machine Learning Models for Credit Risk Analysis: A Survey. In Data Analytics for Management, Banking and Finance (pp. 51–72). https: //doi.org/10.1007/978-3-031-36570-6_2
Khani, A. M., Kazazi, A.  Taqhavi Fard, M. T. (2022). Evaluating the quality of services of the cultural and social deputy of Tehran municipality in the field of culture and art. Social Development & Welfare Planning, 13(50), 205-250. https://doi.org/10.22054/qjsd.2021.58035.2110 (in Persian)
Khanlari, A., Ahrari, M. & Mirpoor, S. (2017). Predicting Customer Lifetime Value Based on Financial and Demographic Characteristics Using GMDH Neural Network Case Study: Individual Customers of a Private Bank of Iran. Journal of Business Management, 8(4), 833-860. https://doi.org/10.22059/jibm.2017.61302 (in Persian)
Khashei Varnamkhasti, V.  & Farsi, S. (2024). A Model for Creating and Implementing Ambidextrous Innovation in Iranian Banking. Journal of Business Management, 16(4), 1002-1028. https://doi.org/10.22059/jibm.2023.364184.4642 (in Persian)
Kotsiantis, S. B. (2011). Decision trees: a recent overview. Artificial Intelligence Review, 39(4), 261–283. https: //doi.org/10.1007/s10462-011-9272-4
Kruse, R., Mostaghim, S., Borgelt, C., Braune, C. & Steinbrecher, M. (2022). Multi-layer perceptrons. In Texts in computer science (pp. 53–124). https: //doi.org/10.1007/978-3-030-42227-1_5
Leevy, J. L., Khoshgoftaar, T. M., Bauder, R. A. & Seliya, N. (2018). A survey on addressing high-class imbalance in big data. Journal of Big Data, 5(1). https: //doi.org/10.1186/s40537-018-0151-6
Luong, H. H., Tran, T. T., Van Nguyen, N., Le, A. D., Nguyen, H. T. T., Nguyen, K. D., Tran, N. C. & Nguyen, H. T. (2022). Feature Selection Using Correlation Matrix on Metagenomic Data with Pearson Enhancing Inflammatory Bowel Disease Prediction. In Lecture notes in electrical engineering (pp. 1073–1084). https: //doi.org/10.1007/978-981-16-2183-3_102
Mastelini, S. M., Nakano, F. K., Vens, C. & De Leon Ferreira De Carvalho, A. C. P. (2022). Online Extra Trees Regressor. IEEE Transactions on Neural Networks and Learning Systems, 34(10), 6755–6767. https: //doi.org/10.1109/tnnls.2022.3212859
Mehregan, M. R. & Khani, A. M. (2024). Improving organizational performance: the role of supply chain 4.0 and financing in reducing supply chain risk. Journal of International Businesses Administration, 7(3), 39-59. doi: 10.22034/jiba.2024.60005.2164 (in Persian)
Mienye, I. D. & Jere, N. (2024). A survey of Decision trees: Concepts, algorithms, and applications. IEEE Access, 12, 86716–86727. https: //doi.org/10.1109/access.2024.3416838
Mousavi, S. M.  & Amiri Aghdaie, S. F. (2021). Identifying the Constructive Elements of “Value Proposition” and their Impact on Customers’ Satisfaction using Sentiment Analysis based on Text Mining. Journal of Business Management, 12(4), 1092-1116. https://doi.org/10.22059/jibm.2020.302987.3847 (in Persian)
Mujahid, M., Kına, E., Rustam, F., Villar, M. G., Alvarado, E. S., De La Torre Diez, I. & Ashraf, I. (2024). Data oversampling and imbalanced datasets: an investigation of performance for machine learning and feature engineering. Journal of Big Data, 11(1). https: //doi.org/10.1186/s40537-024-00943-4
Muneer, A., Ali, R. F., Alghamdi, A., Taib, S. M., Almaghthawi, A. & Ghaleb, E. a. A. (2022). Predicting customers churning in banking industry: A machine learning approach. Indonesian Journal of Electrical Engineering and Computer Science, 26(1), 539. https: //doi.org/10.11591/ijeecs.v26.i1.pp539-549
Nasiroleslami, E., Saniee, E., Abbasian, E., Fathpour Kashani, R. & Gheysari, N. (2024). Prediction of Bank Deposits by Machine Learning Method. Journal of Applied Economics Studies in Iran, 13(50), 137-167. doi: 10.22084/aes.2023.27444.3564
(in Persian)
Nguyen, Q., Nguyen, H., Le, D. T. & Bui, Q. (2022). Fine-Tuning LightGBM using an artificial Ecosystem-Based optimizer for forest fire analysis. Forest Science, 69(1), 73–82. https: //doi.org/10.1093/forsci/fxac039
Peng, K., Peng, Y. & Li, W. (2023). Research on customer churn prediction and model interpretability analysis. PLoS ONE, 18(12), e0289724. https: //doi.org/10.1371/journal.pone.0289724
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V. & Gulin, A. (2018). CatBoost: unbiased boosting with categorical features. Neural Information Processing Systems, 31, 6639–6649. https: //papers.nips.cc/paper/7898-catboost-unbiased-boosting-with-categorical-features.pdf
Przybyła-Kasperek, M. & Marfo, K. F. (2024). A multi-layer perceptron neural network for varied conditional attributes in tabular dispersed data. PLoS ONE, 19(12), e0311041. https: //doi.org/10.1371/journal.pone.0311041
Sanmorino, A., Marnisah, L. & Sunardi, H. (2023). Feature selection using Extra Trees Classifier for Research Productivity Framework in Indonesia. In Lecture notes in electrical engineering (pp. 13–21). https: //doi.org/10.1007/978-981-99-0248-4_2
Schonlau, M. & Zou, R. Y. (2020). The random forest algorithm for statistical learning. The Stata Journal Promoting Communications on Statistics and Stata, 20(1), 3–29. https: //doi.org/10.1177/1536867x20909688
Shingi, G. (2020). A federated learning based approach for loan defaults prediction. 2021 International Conference on Data Mining Workshops (ICDMW), 362–368. https: //doi.org/10.1109/icdmw51313.2020.00057
Siddiqui, N., Haque, M. A., Khan, S. M. S., Adil, M. & Shoaib, H. (2024). Different ML-based strategies for customer churn prediction in banking sector. Journal of Data Information and Management, 6(3), 217–234. https: //doi.org/10.1007/s42488-024-00126-z
Syriopoulos, P. K., Kalampalikis, N. G., Kotsiantis, S. B. & Vrahatis, M. N. (2023). kNN Classification: a review. Annals of Mathematics and Artificial Intelligence. https: //doi.org/10.1007/s10472-023-09882-x
Tékouabou, S. C. K., Gherghina, Ș. C., Toulni, H., Mata, P. N. & Martins, J. M. (2022). Towards explainable machine learning for bank churn prediction using data balancing and Ensemble-Based methods. Mathematics, 10(14), 2379. https: //doi.org/10.3390/math10142379
Tolles, J. & Meurer, W. J. (2016). Logistic regression. JAMA, 316(5), 533. https: //doi.org/10.1001/jama.2016.7653
Tran, H. D., Le, N. & Nguyen, V. (2023). Customer churn prediction in the banking sector using Machine Learning-Based classification models. Interdisciplinary Journal of Information Knowledge and Management, 18, 087–105. https: //doi.org/10.28945/5086
Vittinghoff, E. & McCulloch, C. E. (2006). Relaxing the rule of ten events per variable in logistic and Cox regression. American Journal of Epidemiology, 165(6), 710–718. https: //doi.org/10.1093/aje/kwk052
Vu, V. (2024). An efficient customer churn prediction technique using combined machine learning in commercial banks. Operations Research Forum, 5(3). https: //doi.org/10.1007/s43069-024-00345-5
Wyner, A. J., Olson, M., Bleich, J. & Mease, D. (2017). Explaining the success of adaboost and random forests as interpolating classifiers. Journal of Machine Learning Research, 18(1), 1558–1590. https: //doi.org/10.5555/3122009.3153004
Yaghoubpour, M., Khodadad Hosseini, S. H., Janatifar, H.  Sanavifard, R. (2024). Designing a Digital Marketing Excellence Model (A Case Study of Commercial Banking). Journal of Business Management, 16(3), 596-617. https://doi.org/10.22059/jibm.2024.376605.4785 (in Persian)
Yang, H., Chen, Z., Yang, H. & Tian, M. (2023). Predicting coronary heart disease using an improved LightGBM Model: Performance analysis and comparison. IEEE Access, 11, 23366–23380. https: //doi.org/10.1109/access.2023.3253885