Pembangunan Multi-Domain Human-Labelled Dataset untuk Konversi Bahasa Alami ke SQL pada Bahasa Indonesia
DOI:
https://doi.org/10.37859/jf.v16i2.12020
Abstract
This study develops a multi-domain, human-labelled Text-to-SQL dataset in Indonesian to address the absence of such resources for the language. The dataset was constructed through a corpus-based approach comprising four stages: corpus planning, corpus construction, corpus validation, and model evaluation. Database schemas represented as Entity Relationship Diagrams (ERD) were used as the structural foundation for generating pairs of Indonesian natural language questions and SELECT-type SQL queries, all created through manual annotation. Validation was performed through four sequential mechanisms: SQL syntax checking, schema conformity verification, query execution testing, and semantic alignment assessment between questions and SQL queries. The resulting dataset is Spider-compatible in JSON format, covering 40 databases, 27 domains, 1,524 questions, and 1,355 unique SQL queries distributed across four difficulty levels. Preliminary evaluation using SQLNet and TypeSQL baseline models under example split and database split scenarios confirms that the dataset provides a representative and challenging evaluation environment for Indonesian Text-to-SQL experiments, though model performance remains limited on complex queries and previously unseen database schemas. The dataset is publicly available and is intended to support future development and evaluation of Indonesian Text-to-SQL models..
Downloads
References
Quamar A., Efthymiou V., Lei C., and Özcan F., 2022. Natural Language Interfaces to Data. Foundations and Trends in Databases, 11(4), pp. 319-414. doi: 10.1561/1900000078.
Hong Z., Yuan Z., Zhang Q., Chen H., and Dong J., 2025. Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL. arXiv preprint arXiv:2406.08426, pp. 1-20. doi: 10.48550/arXiv.2406.08426.
Deng N., Chen Y., and Zhang Y., 2022. Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect. Proceedings of the International Conference on Computational Linguistics (COLING), 29(1), pp. 2166-2187.
Qin B., Hui B., Wang L., Yang M., Li J., Li B., Geng R., Cao R., Sun J., Si L., Huang F., and Li Y., 2022. A Survey on Text-to-SQL Parsing: Concepts, Methods, and Future Directions. arXiv preprint arXiv:2208.13629, pp. 1-19. Tersedia di: http://arxiv.org/abs/2208.13629 [Accessed 5 September 2025].
Yu T., Zhang R., Yang K., Yasunaga M., Wang D., Li Z., Ma J., Li I., Yao Q., Roman S., Zhang Z., and Radev D., 2018. Spider: A Large-scale Human-labeled Dataset for Complex and Cross-domain Semantic Parsing and Text-to-SQL Task. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP 2018). pp. 3911-3921. doi: 10.18653/v1/d18-1425.
Hemphill C. T., Godfrey J. J., and Doddington G. R., 1990. The ATIS Spoken Language Systems Pilot Corpus. In: Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania. [Online] Tersedia di: https://aclanthology.org/H90-1021/ [Accessed 5 September 2025].
Zelle J. M., and Mooney R. J., 1996. Learning to Parse Database Queries Using Inductive Logic Programming. In: Proceedings of the Thirteenth National Conference on Artificial Intelligence (AAAI-96), Vol. 2. Portland, OR, August 1996. AAAI Press, pp. 1050-1055.
Mitsopoulou A., and Koutrika G., 2025. Analysis of Text-to-SQL Benchmarks: Limitations, Challenges and Opportunities. In: Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025).
Zhong V., Xiong C., and Socher R., 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. arXiv preprint arXiv:1709.00103, pp. 1-12. doi: 10.48550/arXiv.1709.00103.
Katsogiannis-Meimarakis G., and Koutrika G., 2023. A Survey on Deep Learning Approaches for Text-to-SQL. VLDB Journal, 32(4), pp. 905-936. doi: 10.1007/s00778-022-00776-8.
Dou L., Gao Y., Pan M., Wang D., Che W., Zhan D., and Lou J.-G., 2022. MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing. AAAI Technical Track on Speech and Natural Language Processing, 37(11), pp. 12745-12753. doi: 10.48550/arXiv.2212.13492.
Pisceldo F., Mahendra R., Manurung R., and Arka I. W., 2008. A Two-Level Morphological Analyser for the Indonesian Language.
Koto F., Rahimi A., Lau J. H., and Baldwin T., 2020. IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. In: Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020). Barcelona, Spain (Online), pp. 757-770. doi: 10.18653/v1/2020.coling-main.66.
Cahyawijaya S., et al., 2023. NusaCrowd: Open Source Initiative for Indonesian NLP Resources. In: Findings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada, pp. 13745-13818.
Mahendra R., Fikri A., Samuel A., Rahman F., and Vania C., 2021. IndoNLI: A Natural Language Inference Dataset for Indonesian. In: Proceedings of the Association for Computational Linguistics, pp. 10511-10527.
Klie J.-C., de Castilho R. E., and Gurevych I., 2024. Analyzing Dataset Annotation Quality Management in the Wild. Computational Linguistics, 50(3). doi: 10.1162/coli_a_00516.
Prasetya A., Sari Y. K., Iskandar J., and Ansor M. K., 2024. Identifikasi Jenis Operasi Data Manipulation Language Berbasis BiLSTM pada Kalimat Berbahasa Indonesia. JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), 9(4), pp. 2552-2557. doi: 10.29100/jipi.v9i4.8695.
Xu X., Liu C., and Song D., 2017. SQLNet: Generating Structured Queries from Natural Language without Reinforcement Learning. arXiv preprint arXiv:1711.04436, pp. 1-13. doi: 10.48550/arXiv.1711.04436.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Fiorella Asyfa Firlanda, Agung Prasetya

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Copyright Notice
An author who publishes in the Jurnal FASILKOM (teknologi inFormASi dan ILmu KOMputer) agrees to the following terms:
- Author retains the copyright and grants the journal the right of first publication of the work simultaneously licensed under the Creative Commons Attribution-ShareAlike 4.0 License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal
- Author is able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book) with the acknowledgement of its initial publication in this journal.
- Author is permitted and encouraged to post his/her work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of the published work (See The Effect of Open Access).
Read more about the Creative Commons Attribution-ShareAlike 4.0 Licence here: https://creativecommons.org/licenses/by-sa/4.0/.


