SYSTEMATIC LITERATURE REVIEW TERHADAP SEMANTIC PARSING BERBASIS TRANSFORMER DALAM NATURAL LANGUAGE PROCESSING

Authors

  • Dwi Suci Ramadani Universitas Asahan, Indonesia Author
  • Fahira Hasanah Br Harahap Universitas Asahan, Indonesia Author
  • Suci Lestari Br Batu Bara Universitas Asahan, Indonesia Author
  • Winda Sari Universitas Asahan, Indonesia Author
  • Dicky Apdillah Universitas Asahan, Indonesia Author

DOI:

https://doi.org/10.54314/jssr.v9i3.6689

Keywords:

BERT, Natural Language Processing, Semantic Parsing, Systematic Literature Review, Transformer, T5, Text-to-SQL

Abstract

Abstract: The development of transformer architecture has brought significant changes in Natural Language Processing (NLP) research, particularly in semantic parsing tasks. This study aims to analyze the development of transformer-based semantic parsing through a Systematic Literature Review (SLR) approach. The research method uses the PRISMA guidelines with a process of identification, selection, and analysis of articles from the Scopus, ScienceDirect, IEEE Xplore, Springer, and Google Scholar databases for the period 2019–2026. The results of the study show that transformer models such as BERT, T5, GPT, and transformer encoder–decoder have become the dominant approaches in modern semantic parsing because they are able to improve contextual understanding, semantic representation, and reasoning more effectively than previous methods. The most widely used datasets include Spider, WikiSQL, GeoQuery, ATIS, and Overnight, with the main evaluation being Execution Accuracy and Exact Match Accuracy. This study also identifies various challenges, such as limited labeled data, high computational costs, low model interpretability, the risk of hallucinations, and limitations in low-resource languages. This research contributes to a literature review, identifying research gaps, and recommending future research directions for transformer-based semantic parsing.

Keywords: BERT, Natural Language Processing, Semantic Parsing, Systematic Literature Review, Transformer, T5, Text-to-SQL.

 

Abstrak: Perkembangan arsitektur transformer telah membawa perubahan signifikan dalam penelitian Natural Language Processing (NLP), khususnya pada tugas semantic parsing. Penelitian ini bertujuan untuk menganalisis perkembangan semantic parsing berbasis transformer melalui pendekatan Systematic Literature Review (SLR). Metode penelitian menggunakan pedoman PRISMA dengan proses identifikasi, seleksi, dan analisis artikel dari database Scopus, ScienceDirect, IEEE Xplore, Springer, dan Google Scholar pada periode 2019–2026. Hasil kajian menunjukkan bahwa model transformer seperti BERT, T5, GPT, dan encoder–decoder transformer menjadi pendekatan dominan dalam semantic parsing modern karena mampu meningkatkan contextual understanding, semantic representation, dan reasoning secara lebih efektif dibandingkan metode sebelumnya. Dataset yang paling banyak digunakan meliputi Spider, WikiSQL, GeoQuery, ATIS, dan Overnight, dengan metrik evaluasi utama berupa Execution Accuracy dan Exact Match Accuracy. Penelitian ini juga mengidentifikasi berbagai tantangan, seperti keterbatasan data berlabel, tingginya biaya komputasi, rendahnya interpretabilitas model, risiko hallucination, serta keterbatasan pada low-resource language. Penelitian ini berkontribusi dalam memberikan pemetaan literatur, identifikasi research gap, dan rekomendasi arah penelitian semantic parsing berbasis transformer pada masa mendatang.

Kata Kunci: BERT, Natural Language Processing, Semantic Parsing, Systematic Literature Review, Transformer, T5, Text-to-SQL.

Downloads

Download data is not yet available.

References

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://doi.org/10.48550/arXiv.2005.14165

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, 4171–4186. https://doi.org/10.18653/v1/N19-1423

Guo, D., Tang, N., Wang, G., Shin, Y. C., & Yin, D. (2019). Towards complex text-to-SQL in cross-domain database with intermediate representation. Proceedings of ACL, 4524–4535. https://doi.org/10.18653/v1/P19-1444

Jiang, P., & Cai, X. (2024). A survey of semantic parsing techniques. Symmetry, 16(9), 1201. https://doi.org/10.3390/sym16091201

Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., ... Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274

Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., ... Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. Proceedings of ACL, 7871–7880. https://doi.org/10.18653/v1/2020.acl-main.703

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://doi.org/10.48550/arXiv.1907.11692

Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2021). Deep learning-based text classification: A comprehensive review. ACM Computing Surveys, 54(3), 1–40. https://doi.org/10.1145/3439726

OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

Qin, L., Xu, X., Che, W., Zhang, Y., & Liu, T. (2021). Dynamic fusion network for multi-domain end-to-end task-oriented dialog. Proceedings of ACL. https://doi.org/10.18653/v1/2021.acl-long.55

Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., ... Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67. https://doi.org/10.48550/arXiv.1910.10683

Schneider, P., Klettner, M., Jokinen, K., Simperl, E., & Matthes, F. (2024). Evaluating large language models in semantic parsing for conversational question answering over knowledge graphs. arXiv. https://doi.org/10.48550/arXiv.2401.13958

Shahin, N., & Ismail, L. (2024). From rule-based models to deep learning transformer architectures for natural language processing and sign language translation systems. Artificial Intelligence Review, 57, 271. https://doi.org/10.1007/s10462-024-10752-1

Shaw, P., Uszkoreit, J., & Vaswani, A. (2018). Self-attention with relative position representations. Proceedings of NAACL-HLT, 464–468. https://doi.org/10.18653/v1/N18-2074

Suhr, A., Iyer, S., & Artzi, Y. (2020). Exploring unexplored generalization challenges for cross-database semantic parsing. Proceedings of ACL, 8372–8388. https://doi.org/10.18653/v1/2020.acl-main.742

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. https://doi.org/10.48550/arXiv.1706.

Wang, Y., Xu, K., & Zhao, T. (2020). Enhancing text-to-SQL semantic parsing using BERT. Proceedings of COLING, 3345–3356. https://doi.org/10.18653/v1/2020.coling-main.301

Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., ... Rush, A. (2020). Transformers: State-of-the-art natural language processing. Proceedings of EMNLP: System Demonstrations, 38–45. https://doi.org/10.18653/v1/2020.emnlp-demos.6

Xu, K., Wang, Y., & Yu, G. (2021). Investigating schema linking for text-to-SQL parsing. Information Sciences, 559, 111–123. https://doi.org/10.1016/j.ins.2021.01.068

Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., ... Radev, D. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. Proceedings of EMNLP, 3911–3921. https://doi.org/10.18653/v1/D18-1425

Zhang, Y., Sun, S., Galley, M., Chen, Y. C., Brockett, C., Gao, X., ... Dolan, B. (2020). DialogPT: Large-scale generative pre-training for conversational response generation. Proceedings of ACL, 270–278. https://doi.org/10.18653/v1/2020.acl-demo.30

Zhong, V., Xiong, C., & Socher, R. (2017). Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv. https://doi.org/10.48550/arXiv.1709.00103

Zhou, V., Zhao, Y., & Wang, R. (2022). Cross-domain semantic parsing with transformer-based architectures. Expert Systems with Applications, 201, 117112. https://doi.org/10.1016/j.eswa.2022.117112

Zoph, B., & Le, Q. V. (2020). Neural architecture search with transformer models. arXiv. https://doi.org/10.48550/arXiv.2006.10270

Downloads

Published

2026-06-30

Issue

Section

Artikel

How to Cite

SYSTEMATIC LITERATURE REVIEW TERHADAP SEMANTIC PARSING BERBASIS TRANSFORMER DALAM NATURAL LANGUAGE PROCESSING. (2026). JOURNAL OF SCIENCE AND SOCIAL RESEARCH, 9(3), 4796-4804. https://doi.org/10.54314/jssr.v9i3.6689

Most read articles by the same author(s)