Document Type : Research Paper

Authors

1 Department of Information Technology Management, KI.C., Islamic Azad University, Kish, Iran

2 IDepartment of Industrial Management, CT.C., Islamic Azad University, Tehran, Iran Corresponding Author: moh.afsharkazemi@iauctb.ac.ir

3 Department of Industrial Management and Information Technology, Faculty of Management and Accounting, Shahid Beheshti University, Tehran, Iran

4 Department of Operations Management and Information Technology, Faculty of Management and Accounting, Allameh Tabataba’i University, Tehran, Iran

Abstract

Given the importance of the tourism industry in the country’s economy and culture, and the shortage of domain-specific datasets for training Persian language models, this research aims to fill the existing gap. The objective of this study is to improve the performance of Persian tourism-related question answering systems and to provide accurate responses to queries concerning Iranian attractions and destinations. To achieve this, a collection of reliable tourism texts was gathered and processed, and using advanced natural language processing approaches, including Prompt Engineering, approximately 20,000 question–answer records were generated. In addition, the application of the innovative Answer Window strategy enhanced the quality of the produced data. Subsequently, the ToKA-BERT model was fine-tuned with this dataset, and evaluation results revealed its significant superiority in specialized tourism question answering compared to general-purpose Persian QA models. Both the dataset and the final model have been published on the Hugging Face platform and can serve as valuable resources for developing intelligent systems in the tourism industry and enhancing user experience.

Keywords

Main Subjects

  1. Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., Rosenberg, M., Song, X., Stoica, A., Tiwary, S., & Wang, T. (2018). MS MARCO: A human generated machine reading comprehension dataset (No. arXiv:1611.09268). arXiv. https://doi.org/10.48550/arXiv.1611.09268
  2. Fajcik, M., Jon, J., & Smrz, P. (2021). Rethinking the objectives of extractive question answering (No. arXiv:2008.12804, Version 4). arXiv. https://doi.org/10.48550/arXiv.2008.12804
  3. Farahani, M., Gharachorloo, M., Farahani, M., & Manthouri, M. (2021). ParsBERT: Transformer-based model for Persian language understanding. Neural Processing Letters, 53(6), 3831–3847. https://doi.org/10.1007/s11063-021-10528-4
  4. Joshi, M., Choi, E., Weld, D., & Zettlemoyer, L. (2017). TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In R. Barzilay & M.-Y. Kan (Eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1601–1611).AssociationforComputationalLinguistics. https://doi.org/10.18653/v1/P17-1147
  5. Kazemi, A., Zojaji, Z., Malverdi, M., Mozafari, J., Ebrahimi, F., Abadani, N., Varasteh, M. R., & Nematbakhsh, M. A. (2023). FarsNewsQA: A deep learning-based question answering system for the Persian news articles. Information Retrieval Journal, 26(1), 3. https://doi.org/10.1007/s10791-023-09417-2
  6. Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova,
  7. Masumi, M., Majd, S. S., Shamsfard, M., & Beigy, H. (2024). FaBERT: Pre-training BERT on Persian blogs (No. arXiv:2402.06617).arXiv. https://doi.org/10.48550/arXiv.2402.06617
  8. Mozafari, J., Kazemi, A., Moradi, P., & Nematbakhsh, M. A. (2022). PerAnSel: A novel deep neural network-based system for Persian question answering. Computational Intelligence and Neuroscience,2022(1),3661286. https://doi.org/10.1155/2022/3661286
  9. Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ questions for machine comprehension of text(No.arXiv:1606.05250).arXiv. https://doi.org/10.48550/arXiv.1606.05250
  10. Rosenthal, S., Sil, A., Florian, R., & Roukos, S. (2024). CLAPNQ: Cohesive long-form answers from passages in natural questions for RAG systems (No. arXiv:2404.02103). arXiv. https://doi.org/10.48550/arXiv.2404.02103
  11. SadraeiJavaheri,M.,Moghaddaszadeh,A.,Molazadeh,M.,Naeiji, F.,Aghababaloo,F.,Rafiee,H.,Amirmahani,Z.,Abedini,T.,Sheikhi, F. Z., & Salehoof, A. (2024). TookaBERT: A step forward forPersianNLU(No.arXiv:2407.16382).arXiv. https://doi.org/10.48550/arXiv.2407.16382
  12. Soh, Y. J., & Zhao, J. (2024). A step towards mixture of grader: Statistical analysis of existing automatic evaluation metrics (No. arXiv:2410.10030,Version1).arXiv. https://doi.org/10.48550/arXiv.2410.10030
  13. Taghizadeh, N., Doostmohammadi, E., Seifossadat, E., Rabiee, H. R., & Tahaei, M. S. (2021). SINA-BERT: A pre-trained language model for analysis of medical texts in Persian (No. arXiv:2104.07613).arXiv. https://doi.org/10.48550/arXiv.2104.07613
  14. Tran, S. Q., Do, G.-H., Do, P. N.-T., Kretchmar, M., & Du, X. (2023). AGent: A novel pipeline for automatically creating unanswerablequestions(No.arXiv:2309.05103).arXiv. https://doi.org/10.48550/arXiv.2309.05103
  15. Wang, Z., Ng, P., Ma, X., Nallapati, R., & Xiang, B. (2019). Multi-passage BERT: A globally normalized BERT model for open-domain question answering. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP-IJCNLP)(pp.5878–5882). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1599
  16. Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., & Manning, C. D. (2018). HotpotQA: A dataset for diverse, explainable multi-hop question answering. In E. Riloff, D. Chiang, J. Hockenmaier, & J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in NaturalLanguageProcessing(pp.2369–2380).Associationfor Computational Linguistics. https://doi.org/10.18653/v1/D18-1259