MITIGASI BIAS LARGE LANGUAGE MODEL MELALUI HUMAN-IN-THE-LOOP PADA AUTOMATED ESSAY SCORING BERBASIS ESAI IELTS-LIKE
DOI:
https://doi.org/10.37859/seis.v6i2.11478
Abstract
This study evaluates the reliability, algorithmic bias, and feasibility of implementing Human-in-the-Loop (HITL) in Large Language Model (LLM)-based Automated Essay Scoring (AES). Two architectures were compared: the Mixture-of-Experts architecture represented by GPT OSS 120B and the Dense architecture represented by Qwen3-32B, using a multiple-run scoring approach on 150 IELTS-like essays written by 50 respondents. Each essay was evaluated five times across four criteria: Grammar, Lexical Resource, Coherence, and Task Achievement. The evaluation employed the Intraclass Correlation Coefficient (ICC), Coefficient of Variation (CV), score range, paired t-test, and confidence-based routing simulation.The results show that GPT OSS 120B achieved higher reliability, with an ICC of 0.94, an average range of 2.40 points, and a CV of 4.69%, while Qwen3-32B obtained an ICC of 0.84, an average range of 5.56 points, and a CV of 9.17%. However, GPT OSS 120B experienced a parsing failure rate of 25.4%, whereas Qwen3-32B demonstrated full format compliance. Comparative analysis also showed that GPT OSS 120B tended to score more strictly, while Qwen3-32B was more lenient. In the Data Report task without visual input, both models exhibited conservatism bias due to contextual limitations. The HITL simulation showed that GPT OSS 120B could automatically approve 55.4% of essays, compared with only 7.8% for Qwen3-32B. These findings highlight the importance of HITL in maintaining the reliability, fairness, and integrity of academic assessment.
Downloads
References
Andersen, N., Mang, J., Goldhammer, F., & Zehner, F. (2025). Algorithmic fairness in automatic short answer scoring. International Journal of Artificial Intelligence in Education, 35, 3128–3165 (2025). https://doi.org/10.1007/s40593-025-00495-5
Caton, S., & Haas, C. (2024). Fairness in machine learning: A survey. ACM Computing Surveys, 56(7), Article 166. https://doi.org/10.1145/3616865
Che Mat, A., Zulkornain, L. H., & Abdul Rahman, N. A. (2024). Automated writing evaluation: Users' perception and expectations. International Journal of Information and Education Technology, 14(2), 183–192. https://doi.org/10.18178/ijiet.2024.14.2.2039
Du, J., & Nordin, N. R. M. (2025). A systematic review of automated writing evaluation (AWE) systems on university students’ English writing performance (2020 – 2024). Forum for Linguistic Studies, 7(11), 615–632. https://doi.org/10.30564/fls.v7i11.11764
Du, X., Gunter, T., Kong, X., Lee, M., Wang, Z., Zhang, A., Du, N., & Pang, R. (2024). Revisiting MoE and dense speed-accuracy comparisons for LLM training. arXiv. https://doi.org/10.48550/arXiv.2405.15052
Laclau, C., Largeron, C., & Choudhary, M. (2022). A survey on fairness for machine learning on graphs. arXiv. https://arxiv.org/abs/2205.05396v2
Li, H., Lo, K. M., Wang, Z., Wang, Z., Zheng, W., Zhou, S., Zhang, X., & Jiang, D. (2025). Can mixture-of-experts surpass dense LLMs under strictly equal resources? ArXiv.org. https://arxiv.org/abs/2506.12119v1
Nuong Deri, M., Singh, A., Zaazie, P., & Anandene, D. (2024). Leveraging artificial intelligence in higher educational institutions: A comprehensive overview. Revista De Educación Y Derecho, (30). https://doi.org/10.1344/REYD2024.30.45777
Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234
Palomino, A., Fischer, A., Buschhüter, D., Roller, R., Pinkwart, N., & Paassen, B. (2025). Mitigating bias in item retrieval for enhancing exam assembly in vocational education services. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), 183–193. https://doi.org/10.18653/v1/2025.naacl-industry.16
Patwardhan, N., Marrone, S., & Sansone, C. (2023). Transformers in the real world: A survey on NLP applications. Information, 14(4), 242. https://doi.org/10.3390/info14040242
Pratama A. R. (2025). The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication. PeerJ Computer science, 11, e2953. https://doi.org/10.7717/peerj-cs.2953
Selvam, M., & González Vallejo, R. (2025). human-in-the-loop models for ethical AI grading: Combining AI speed with human ethical oversight. EthAIca, 4, 413. https://doi.org/10.56294/ai2025413
Shao, M., Li, D., Zhao, C., Wu, X., Lin, Y., & Tian, Q. (2024). Supervised algorithmic fairness in distribution shifts: A survey. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI) (pp. 8225–8233). https://doi.org/10.24963/ijcai.2024/909
Siddique, S., Haque, M. A., George, R., Gupta, K. D., Gupta, D., & Faruk, M. J. H. (2024). Survey on machine learning biases and mitigation techniques. Digital, 4(1), 1-68. https://doi.org/10.3390/digital4010001
Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay grading: An examination of reliability and validity in rubric-based assessments. British Journal of Educational Technology, 56, 150–166. https://doi.org/10.1111/bjet.13494
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhammad Abdiel Al Hafiz, M. Syaiful Amin

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Copyright Notice
An author who publishes in the Journal of Software Engineering and Information System (SEIS) agrees to the following terms:
- Author retains the copyright and grants the journal the right of first publication of the work simultaneously licensed under the Creative Commons Attribution-ShareAlike 4.0 License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal
- Author is able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book) with the acknowledgement of its initial publication in this journal.
- Author is permitted and encouraged to post his/her work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of the published work (See The Effect of Open Access).
Read more about the Creative Commons Attribution-ShareAlike 4.0 Licence here: https://creativecommons.org/licenses/by-sa/4.0/.
