MITIGASI BIAS LARGE LANGUAGE MODEL MELALUI HUMAN-IN-THE-LOOP PADA AUTOMATED ESSAY SCORING BERBASIS ESAI IELTS-LIKE

Authors

  • Muhammad Abdiel Al Hafiz Teknik Informatika, Fakultas Ilmu Komputer, Universitas Amikom Purwokerto
  • M. Syaiful Amin Teknik Informatika, Fakultas Ilmu Komputer, Universitas Amikom Purwokerto

DOI:

https://doi.org/10.37859/seis.v6i2.11478
Keywords: Automated Essay Scoring, Algorithmic Bias, Human-in-the-Loop, Large Language Model

Abstract

This study evaluates the reliability, algorithmic bias, and feasibility of implementing Human-in-the-Loop (HITL) in Large Language Model (LLM)-based Automated Essay Scoring (AES). Two architectures were compared: the Mixture-of-Experts architecture represented by GPT OSS 120B and the Dense architecture represented by Qwen3-32B, using a multiple-run scoring approach on 150 IELTS-like essays written by 50 respondents. Each essay was evaluated five times across four criteria: Grammar, Lexical Resource, Coherence, and Task Achievement. The evaluation employed the Intraclass Correlation Coefficient (ICC), Coefficient of Variation (CV), score range, paired t-test, and confidence-based routing simulation.The results show that GPT OSS 120B achieved higher reliability, with an ICC of 0.94, an average range of 2.40 points, and a CV of 4.69%, while Qwen3-32B obtained an ICC of 0.84, an average range of 5.56 points, and a CV of 9.17%. However, GPT OSS 120B experienced a parsing failure rate of 25.4%, whereas Qwen3-32B demonstrated full format compliance. Comparative analysis also showed that GPT OSS 120B tended to score more strictly, while Qwen3-32B was more lenient. In the Data Report task without visual input, both models exhibited conservatism bias due to contextual limitations. The HITL simulation showed that GPT OSS 120B could automatically approve 55.4% of essays, compared with only 7.8% for Qwen3-32B. These findings highlight the importance of HITL in maintaining the reliability, fairness, and integrity of academic assessment.

Downloads

Download data is not yet available.

References

Andersen, N., Mang, J., Goldhammer, F., & Zehner, F. (2025). Algorithmic fairness in automatic short answer scoring. International Journal of Artificial Intelligence in Education, 35, 3128–3165 (2025). https://doi.org/10.1007/s40593-025-00495-5

Caton, S., & Haas, C. (2024). Fairness in machine learning: A survey. ACM Computing Surveys, 56(7), Article 166. https://doi.org/10.1145/3616865

Che Mat, A., Zulkornain, L. H., & Abdul Rahman, N. A. (2024). Automated writing evaluation: Users' perception and expectations. International Journal of Information and Education Technology, 14(2), 183–192. https://doi.org/10.18178/ijiet.2024.14.2.2039

Du, J., & Nordin, N. R. M. (2025). A systematic review of automated writing evaluation (AWE) systems on university students’ English writing performance (2020 – 2024). Forum for Linguistic Studies, 7(11), 615–632. https://doi.org/10.30564/fls.v7i11.11764

Du, X., Gunter, T., Kong, X., Lee, M., Wang, Z., Zhang, A., Du, N., & Pang, R. (2024). Revisiting MoE and dense speed-accuracy comparisons for LLM training. arXiv. https://doi.org/10.48550/arXiv.2405.15052

Laclau, C., Largeron, C., & Choudhary, M. (2022). A survey on fairness for machine learning on graphs. arXiv. https://arxiv.org/abs/2205.05396v2

Li, H., Lo, K. M., Wang, Z., Wang, Z., Zheng, W., Zhou, S., Zhang, X., & Jiang, D. (2025). Can mixture-of-experts surpass dense LLMs under strictly equal resources? ArXiv.org. https://arxiv.org/abs/2506.12119v1

Nuong Deri, M., Singh, A., Zaazie, P., & Anandene, D. (2024). Leveraging artificial intelligence in higher educational institutions: A comprehensive overview. Revista De Educación Y Derecho, (30). https://doi.org/10.1344/REYD2024.30.45777

Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234

Palomino, A., Fischer, A., Buschhüter, D., Roller, R., Pinkwart, N., & Paassen, B. (2025). Mitigating bias in item retrieval for enhancing exam assembly in vocational education services. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), 183–193. https://doi.org/10.18653/v1/2025.naacl-industry.16

Patwardhan, N., Marrone, S., & Sansone, C. (2023). Transformers in the real world: A survey on NLP applications. Information, 14(4), 242. https://doi.org/10.3390/info14040242

Pratama A. R. (2025). The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication. PeerJ Computer science, 11, e2953. https://doi.org/10.7717/peerj-cs.2953

Selvam, M., & González Vallejo, R. (2025). human-in-the-loop models for ethical AI grading: Combining AI speed with human ethical oversight. EthAIca, 4, 413. https://doi.org/10.56294/ai2025413

Shao, M., Li, D., Zhao, C., Wu, X., Lin, Y., & Tian, Q. (2024). Supervised algorithmic fairness in distribution shifts: A survey. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI) (pp. 8225–8233). https://doi.org/10.24963/ijcai.2024/909

Siddique, S., Haque, M. A., George, R., Gupta, K. D., Gupta, D., & Faruk, M. J. H. (2024). Survey on machine learning biases and mitigation techniques. Digital, 4(1), 1-68. https://doi.org/10.3390/digital4010001

Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay grading: An examination of reliability and validity in rubric-based assessments. British Journal of Educational Technology, 56, 150–166. https://doi.org/10.1111/bjet.13494

Downloads

Published

2026-08-31

How to Cite

Al Hafiz, M. A. ., & Amin, M. S. . (2026). MITIGASI BIAS LARGE LANGUAGE MODEL MELALUI HUMAN-IN-THE-LOOP PADA AUTOMATED ESSAY SCORING BERBASIS ESAI IELTS-LIKE. Journal of Software Engineering and Information System (SEIS), 6(2), 124–132. https://doi.org/10.37859/seis.v6i2.11478

Issue

Section

Articles