Pengembangan Model Klasifikasi Code Smells Pada Backend Python Menggunakan Algoritma Random Forest (Studi Kasus Proyek Open Source Github)
DOI:
https://doi.org/10.37859/jf.v16i2.12107
Abstract
Deadline pressure in software development often drives coding shortcuts, leading to internal quality degradation known as code smells. These structural anomalies contribute to technical debt accumulation and complicate system maintenance over time. This study develops an automated classification model to detect code smell contamination in the Python backend ecosystem. The methodology uses the Random Forest ensemble algorithm integrated with the Synthetic Minority Over-sampling Technique (SMOTE) for class balancing. Data mining on GitHub with high-reputation criteria extracted 137,728 code samples from 7 large-scale repositories using the Radon multi-metric tool. To simulate human error in real-world scenarios, 5% random noise was inserted into the labeling data. Testing using the confusion matrix shows the proposed model achieves highly stable and balanced performance, with average precision, recall, and f1-score of 0.95 in both macro and weighted averages. Ablation study analysis proves that SMOTE intervention effectively maintains detection consistency in minority class categories. Feature importance ranking identifies the Logical Lines of Code (LLOC) metric as the most crucial indicator with 37.65% influence weight, followed by LOC and Blank metrics. This research provides an automated quality assurance system for developers to detect code refactoring opportunities at an early stage.
Downloads
References
T. Lewowski and L. Madeyski, “How far are we from reproducible research on code smell detection? A systematic literature review,” Inf. Softw. Technol., vol. 144, no. November 2021, p. 106783, 2022, doi: 10.1016/j.infsof.2021.106783.
P. S. Yadav, R. S. Rao, A. Mishra, and M. Gupta, “Machine Learning-Based Methods for Code Smell Detection: A Survey,” Appl. Sci., vol. 14, no. 14, 2024, doi: 10.3390/app14146149.
M. Hadj-Kacem and N. Bouassida, "Application of Deep Learning for Code Smell Detection: Challenges and Opportunities," SN Comput. Sci., vol. 5, p. 553, 2024, doi: 10.1007/s42979-024-02956-5.
P. S. Thakur, S. S. Chouhan, and S. S. Rathore, "Systematic literature review on software code smell detection approaches," J. Syst. Softw., vol. 222, p. 112345, 2026, doi: 10.1016/j.jss.2026.112784.
A. Abdou and N. Darwish, "Severity classification of software code smells using machine learning techniques: A comparative study," J. Softw. Evol. Process, vol. 36, no. 1, p. e2454, 2024, doi: 10.1002/smr.2454.
A. Alazba and H. Aljamaan, “Code smell detection using feature selection and stacking ensemble: An empirical investigation,” Inf. Softw. Technol., vol. 138, no. May, p. 106648, 2021, doi: 10.1016/j.infsof.2021.106648.
S. Dewangan, R. S. Rao, A. Mishra, and M. Gupta, “A novel approach for code smell detection: An empirical study,” IEEE Access, vol. 9, pp. 162869–162883, 2021, doi: 10.1109/ACCESS.2021.3133810.
S. Dewangan, R. S. Rao, A. Mishra, and M. Gupta, "Code Smell Detection Using Ensemble Machine Learning Algorithms," Appl. Sci., vol. 12, no. 20, p. 10321, 2022, doi: 10.3390/app122010321.
R. S. Rao, S. Dewangan, A. Mishra, and M. Gupta, “A study of dealing class imbalance problem with machine learning methods for code smell severity detection using PCA-based feature selection technique,” Sci. Rep., vol. 13, no. 1, pp. 1–18, 2023, doi: 10.1038/s41598-023-43380-8.
T. Guggulothu and S. A. Moiz, “Code smell detection using multi-label classification approach,” Softw. Qual. J., vol. 28, no. 3, pp. 1063–1086, 2020, doi: 10.1007/s11219-020-09498-y.
F. do R. Santos and R. Choren, “Data preprocessing for machine learning based code smell detection: A systematic literature review,” Inf. Softw. Technol., vol. 184, p. 107752, 2025, doi: 10.1016/j.infsof.2025.107752.
K. Alkharabsheh, S. Alawadi, Y. Crespo, and J. A. Taboada, “Exploring the role of project status information in effective code smell detection,” Cluster Comput., vol. 28, no. 1, 2025, doi: 10.1007/s10586-024-04724-9.
M. Zakeri-Nasrabadi, S. Parsa, E. Esmaili, and F. Palomba, “A Systematic Literature Review on the Code Smells Datasets and Validation Mechanisms,” ACM Comput. Surv., vol. 55, no. 13 s, 2023, doi: 10.1145/3596908.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Aqilla Ar-Hammar, Mira Maisura

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Copyright Notice
An author who publishes in the Jurnal FASILKOM (teknologi inFormASi dan ILmu KOMputer) agrees to the following terms:
- Author retains the copyright and grants the journal the right of first publication of the work simultaneously licensed under the Creative Commons Attribution-ShareAlike 4.0 License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal
- Author is able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book) with the acknowledgement of its initial publication in this journal.
- Author is permitted and encouraged to post his/her work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of the published work (See The Effect of Open Access).
Read more about the Creative Commons Attribution-ShareAlike 4.0 Licence here: https://creativecommons.org/licenses/by-sa/4.0/.


