Model Deteksi Teks Generatif menggunakan Stacking Ensemble Berbasis DistilBERT dan Few-Shot Adaptation
DOI:
https://doi.org/10.37859/coscitech.v7i2.11770
Abstract
The rapid development of Generative Artificial Intelligence (AI) technologies such as ChatGPT, Claude, and Gemini has brought new challenges to academic integrity in higher education. Detection of AI-generated written works has become crucial to ensure the originality and intellectual honesty of students. However, conventional AI detectors are often static and have low accuracy against dynamic variations in writing styles. This research aims to develop a more adaptive and accurate generative text detection system by applying the Stacking Ensemble Learning method. This system integrates four basic classification models (Logistic Regression, SVM, MLP, and Random Forest) enhanced with high-level semantic features using DistilBERT Embeddings. In addition, this system utilizes the Google Gemini API as a feature augmentation instrument to obtain real-time qualitative risk assessments. The main innovation in this research lies in the application of the Few-Shot Adaptation technique. This feature allows lecturers or users in higher education to instantly recalibrate the Meta-Classifier by simply uploading a few examples of original documents (human-written) and AI-generated documents. The system was developed web-based using the Gradio framework, which can process various academic document formats such as .pdf, .docx, and .txt. The results of this research are expected to provide a tangible contribution to higher education institutions in the form of a digital instrument capable of precisely validating the authenticity of academic documents. In addition to providing a probability score, the system also includes a Text Highlighting feature to mark suspicious text sections, thus simplifying the evaluation process by educators.
Keywords: AI Detection, Stacking Ensemble, DistilBERT, Few-Shot Adaptation, LLM.
Downloads
References
Brown, T., et al. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems.
Radford, A., et al. (2019). Language Models are Unsupervised Multi-Task Learners. OpenAI Blog.
Mitchell, E., et al. (2023). DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. arXiv preprint.
Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing. Stanford University.
Vaswani, A., et al. (2017). Attention is All You Need. Advances in Neural Information Processing Systems.
Sanh, V., et al. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint.
Devlin, J., et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT.
Zhou, Z. H. (2012). Ensemble Methods: Foundations and Algorithms. CRC Press.
Wolpert, D. H. (1992). Stacked Generalization. Neural Networks.
Breiman, L. (1996). Stacked Regressions. Machine Learning.
Wang, Y., et al. (2020). Generalizing from a Few Examples: A Survey on Few-Shot Learning. ACM Computing Surveys.
Finn, C., et al. (2017). Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. ICML.
Jawahar, G., et al. (2020). Automatic Detection of Machine Generated Text: A Survey. Proceedings of the 28th COLING.
Eaton, S. E. (2021). Plagiarism in Higher Education: Tackling Academic Integrity with a Global Perspective. ABC-CLIO.
Crothers, E., et al. (2023). Machine-Generated Text Detection: A Comprehensive Survey. arXiv preprint.










