Prediksi Keberhasilan Akademik Siswa Berbasis Fitur Kategorikal Na-tive dengan Explainable AI (SHAP) menggunakan CatBoost vs LightGBM

Authors

  • Bayu Anugerah Putra Universitas Muhammadiyah Riau
  • Soni Soni Universitas Muhammadiyah Riau
  • Rahmad Firdaus Universitas Muhammadiyah Riau
  • Ayunda Putri Universitas Muhammadiyah Riau
  • Anisa Dwi Sanggar Wati Universitas Muhammadiyah Riau

DOI:

https://doi.org/10.37859/jf.v16i2.12360
Keywords: catboost, lightGBM, academic prediction, k-fold cross-validation, SHAP

Abstract

Prediction of students academic success is important to support decision-making in education. Educational datasets are generally dominated by categorical variables that require encoding before modeling, which may cause information loss and reduce accuracy. This study applies the CatBoost algorithm, which processes categorical variables natively without additional encoding, to predict students' Exam Score on the Student Performance Factors dataset from Kaggle, with LightGBM used as a comparison model. Evaluation was carried out under three data-split schemes (70:30, 80:20, 90:10) using k-fold cross-validation and three regression metrics (R², MAE, RMSE), followed by model interpretation using Shapley Additive Explanations (SHAP). The results show that CatBoost consistently outperforms LightGBM across all schemes, with the best performance obtained under the 90:10 scheme (CatBoost: R² = 0.851, MAE = 0.475, RMSE = 1.414; LightGBM: R² = 0.809, MAE = 0.758, RMSE = 1.599). SHAP analysis identifies Attendance, Hours_Studied, and Previous_Scores as the most influential features in the prediction. These findings confirm that combining CatBoost with SHAP produces an academic prediction model that is both accurate and transparen.

Downloads

Download data is not yet available.

References

B. Santana-Perera, C. García-Barceló, M. González Arcas, and D. Gil, “Exploring Predictive Insights on Student Success Using Explainable Machine Learning: A Synthetic Data Study,” Information, vol. 16, no. 9, p. 763, Sep. 2025, doi: 10.3390/info16090763.

M. Credé, S. G. Roch, and U. M. Kieszczynka, “Class At-tendance in College,” Rev. Educ. Res., vol. 80, no. 2, pp. 272–295, Jun. 2010, doi: 10.3102/0034654310362998.

F. Aprilia, R. A. Anggraini, and Y. D. Putri, “Prediksi Kelulu-san Siswa dengan Algoritma Pembelajaran Mesin: Aplikasi Regresi Linear dan Logistik pada Faktor-Faktor Pendidikan,” ROUTERS: Jurnal Sistem dan Teknologi Informasi, pp. 55–64, Feb. 2025, doi: 10.25181/rt.v3i1.3897.

A. Joshi, P. Saggar, R. Jain, M. Sharma, D. Gupta, and A. Khanna, “CatBoost — An Ensemble Machine Learning Model for Prediction and Classification of Student Academic Performance,” Advances in Data Science and Adaptive Analysis, vol. 13, no. 03n04, Jul. 2021, doi: 10.1142/S2424922X21410023.

Zulwisli, Ambiyar, Muhammad Anwar, and Andhika Herayono, “Student’s Digital Intentions Prediction Using CatBoost,” JTP - Jurnal Teknologi Pendidikan, vol. 27, no. 1, pp. 277–290, Apr. 2025, doi: 10.21009/jtp.v27i1.54035.

N. Levi Sabili, F. Rakhmat Umbara, and M. Melina, “KLAS-IFIKASI PENYAKIT DIABETES MENGGUNAKAN AL-GORITMA CATEGORICAL BOOSTING DENGAN FAKTOR RISIKO DIABETES,” JATI (Jurnal Mahasiswa Teknik Informatika), vol. 8, no. 6, pp. 11391–11398, Nov. 2024, doi: 10.36040/jati.v8i6.11447.

M. T. Syamkalla, S. Khomsah, and Y. S. R. Nur, “Implemen-tasi Algoritma Catboost Dan Shapley Additive Explanations (SHAP) Dalam Memprediksi Popularitas Game Indie Pada Platform Steam,” Jurnal Teknologi Informasi dan Ilmu Kom-puter, vol. 11, no. 4, pp. 777–786, Aug. 2024, doi: 10.25126/jtiik.1148503.

A. B. Mawardi, R. S. Pradini, and M. S. Haris, “KOMPA-RASI ALGORITMA BOOSTING UNTUK PREDIKSI GANGGUAN TIDUR,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 3, Jul. 2025, doi: 10.23960/jitet.v13i3.7281.

Tangkas Surya Wibawa, N. K. Ningrum, and Ahmad Syahreza, “Comparison of CatBoost and LightGBM Models for Air Humidity Prediction,” Journal of Applied Informatics and Computing, vol. 9, no. 3, pp. 803–809, Jun. 2025, doi: 10.30871/jaic.v9i3.9570.

J. T. Hancock and T. M. Khoshgoftaar, “CatBoost for big data: an interdisciplinary review,” J. Big Data, vol. 7, no. 1, p. 94, Dec. 2020, doi: 10.1186/s40537-020-00369-8.

T. Gori, A. Sunyoto, and H. Al Fatta, “Preprocessing Data dan Klasifikasi untuk Prediksi Kinerja Akademik Siswa,” Jurnal Teknologi Informasi dan Ilmu Komputer, vol. 11, no. 1, pp. 215–224, Feb. 2024, doi: 10.25126/jtiik.20241118074.

S. Bates, T. Hastie, and R. Tibshirani, “Cross-Validation: What Does It Estimate and How Well Does It Do It?,” J. Am. Stat. Assoc., vol. 119, no. 546, pp. 1434–1445, Apr. 2024, doi: 10.1080/01621459.2023.2197686.

T. Z. Jasman, M. A. Fadhlullah, A. L. Pratama, and R. Ris-mayani, “Analisis Algoritma Gradient Boosting, Adaboost dan Catboost dalam Klasifikasi Kualitas Air,” Jurnal Teknik Informatika dan Sistem Informasi, vol. 8, no. 2, Aug. 2022, doi: 10.28932/jutisi.v8i2.4906.

Z. Fan, J. Gou, and S. Weng, “A Feature Importance-Based Multi-Layer CatBoost for Student Performance Prediction,” IEEE Trans. Knowl. Data Eng., vol. 36, no. 11, pp. 5495–5507, Nov. 2024, doi: 10.1109/TKDE.2024.3393472.

D. Chicco, M. J. Warrens, and G. Jurman, “The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analy-sis evaluation,” PeerJ Comput. Sci., vol. 7, p. e623, Jul. 2021, doi: 10.7717/peerj-cs.623.

S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nat. Mach. In-tell., vol. 2, no. 1, pp. 56–67, Jan. 2020, doi: 10.1038/s42256-019-0138-9.

Downloads

Published

2026-08-30