Pembangunan Multi-Domain Human-Labelled Dataset untuk Konversi Bahasa Alami ke SQL pada Bahasa Indonesia

Authors

  • Fiorella Asyfa Firlanda Universitas Bhinneka PGRI
  • Agung Prasetya Universitas Bhinneka PGRI Tulungagung

DOI:

https://doi.org/10.37859/jf.v16i2.12020
Keywords: human-labelled dataset, Indonesian language, multi-domain, NLP, text-to-SQL

Abstract

This study develops a multi-domain, human-labelled Text-to-SQL dataset in Indonesian to address the absence of such resources for the language. The dataset was constructed through a corpus-based approach comprising four stages: corpus planning, corpus construction, corpus validation, and model evaluation. Database schemas represented as Entity Relationship Diagrams (ERD) were used as the structural foundation for generating pairs of Indonesian natural language questions and SELECT-type SQL queries, all created through manual annotation. Validation was performed through four sequential mechanisms: SQL syntax checking, schema conformity verification, query execution testing, and semantic alignment assessment between questions and SQL queries. The resulting dataset is Spider-compatible in JSON format, covering 40 databases, 27 domains, 1,524 questions, and 1,355 unique SQL queries distributed across four difficulty levels. Preliminary evaluation using SQLNet and TypeSQL baseline models under example split and database split scenarios confirms that the dataset provides a representative and challenging evaluation environment for Indonesian Text-to-SQL experiments, though model performance remains limited on complex queries and previously unseen database schemas. The dataset is publicly available and is intended to support future development and evaluation of Indonesian Text-to-SQL models..

Downloads

Download data is not yet available.

References

Quamar A., Efthymiou V., Lei C., and Özcan F., 2022. Natural Language Interfaces to Data. Foundations and Trends in Databases, 11(4), pp. 319-414. doi: 10.1561/1900000078.

Hong Z., Yuan Z., Zhang Q., Chen H., and Dong J., 2025. Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL. arXiv preprint arXiv:2406.08426, pp. 1-20. doi: 10.48550/arXiv.2406.08426.

Deng N., Chen Y., and Zhang Y., 2022. Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect. Proceedings of the International Conference on Computational Linguistics (COLING), 29(1), pp. 2166-2187.

Qin B., Hui B., Wang L., Yang M., Li J., Li B., Geng R., Cao R., Sun J., Si L., Huang F., and Li Y., 2022. A Survey on Text-to-SQL Parsing: Concepts, Methods, and Future Directions. arXiv preprint arXiv:2208.13629, pp. 1-19. Tersedia di: http://arxiv.org/abs/2208.13629 [Accessed 5 September 2025].

Yu T., Zhang R., Yang K., Yasunaga M., Wang D., Li Z., Ma J., Li I., Yao Q., Roman S., Zhang Z., and Radev D., 2018. Spider: A Large-scale Human-labeled Dataset for Complex and Cross-domain Semantic Parsing and Text-to-SQL Task. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP 2018). pp. 3911-3921. doi: 10.18653/v1/d18-1425.

Hemphill C. T., Godfrey J. J., and Doddington G. R., 1990. The ATIS Spoken Language Systems Pilot Corpus. In: Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania. [Online] Tersedia di: https://aclanthology.org/H90-1021/ [Accessed 5 September 2025].

Zelle J. M., and Mooney R. J., 1996. Learning to Parse Database Queries Using Inductive Logic Programming. In: Proceedings of the Thirteenth National Conference on Artificial Intelligence (AAAI-96), Vol. 2. Portland, OR, August 1996. AAAI Press, pp. 1050-1055.

Mitsopoulou A., and Koutrika G., 2025. Analysis of Text-to-SQL Benchmarks: Limitations, Challenges and Opportunities. In: Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025).

Zhong V., Xiong C., and Socher R., 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. arXiv preprint arXiv:1709.00103, pp. 1-12. doi: 10.48550/arXiv.1709.00103.

Katsogiannis-Meimarakis G., and Koutrika G., 2023. A Survey on Deep Learning Approaches for Text-to-SQL. VLDB Journal, 32(4), pp. 905-936. doi: 10.1007/s00778-022-00776-8.

Dou L., Gao Y., Pan M., Wang D., Che W., Zhan D., and Lou J.-G., 2022. MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing. AAAI Technical Track on Speech and Natural Language Processing, 37(11), pp. 12745-12753. doi: 10.48550/arXiv.2212.13492.

Pisceldo F., Mahendra R., Manurung R., and Arka I. W., 2008. A Two-Level Morphological Analyser for the Indonesian Language.

Koto F., Rahimi A., Lau J. H., and Baldwin T., 2020. IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. In: Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020). Barcelona, Spain (Online), pp. 757-770. doi: 10.18653/v1/2020.coling-main.66.

Cahyawijaya S., et al., 2023. NusaCrowd: Open Source Initiative for Indonesian NLP Resources. In: Findings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada, pp. 13745-13818.

Mahendra R., Fikri A., Samuel A., Rahman F., and Vania C., 2021. IndoNLI: A Natural Language Inference Dataset for Indonesian. In: Proceedings of the Association for Computational Linguistics, pp. 10511-10527.

Klie J.-C., de Castilho R. E., and Gurevych I., 2024. Analyzing Dataset Annotation Quality Management in the Wild. Computational Linguistics, 50(3). doi: 10.1162/coli_a_00516.

Prasetya A., Sari Y. K., Iskandar J., and Ansor M. K., 2024. Identifikasi Jenis Operasi Data Manipulation Language Berbasis BiLSTM pada Kalimat Berbahasa Indonesia. JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), 9(4), pp. 2552-2557. doi: 10.29100/jipi.v9i4.8695.

Xu X., Liu C., and Song D., 2017. SQLNet: Generating Structured Queries from Natural Language without Reinforcement Learning. arXiv preprint arXiv:1711.04436, pp. 1-13. doi: 10.48550/arXiv.1711.04436.

Downloads

Published

2026-08-30