نوع مقاله : مقاله پژوهشی
نویسندگان
1 گروه مدیریت فناوری اطلاعات، واحد بین المللی کیش، دانشگاه آزاد اسلامی، کیش، ایران
2 گروه مدیریت صنعتی، واحد تهران مرکزی، دانشگاه آزاد اسلامی، تهران، ایران نویسنده مسئول: moh.afsharkazemi@iauctb.ac.ir
3 گروه مدیریت صنعتی و فناوری اطلاعات، دانشکده مدیریت و حسابداری، دانشگاه شهید بهشتی، تهران، ایران
4 گروه مدیریت عملیات و فناوری اطلاعات، دانشکده مدیریت و حسابداری، دانشگاه علامه طباطبائی، تهران، ایران
چکیده
با توجه به اهمیت صنعت گردشگری در اقتصاد و فرهنگ کشور و کمبود منابع دادهای تخصصی برای آموزش مدلهای زبانی فارسی، این تحقیق درصدد پر کردن این خلأ برآمد. هدف از این پژوهش، ارتقای عملکرد سیستمهای پاسخگویی به سؤالات گردشگری فارسی و ارائه پاسخهای دقیق به پرسوجوهای مرتبط با جاذبهها و مقاصد گردشگری ایران بوده است. بدین منظور، مجموعهای از متون معتبر گردشگری گردآوری و پردازش شد و با بهرهگیری از رویکردهای پیشرفته پردازش زبان طبیعی، از جمله مهندسی پرامپت (Prompt Engineering)، حدود ۲۰ هزار رکورد پرسش و پاسخ تولید گردید. علاوهبراین، استفاده از راهبرد نوآورانه پنجره پاسخ (Answer Window) موجب ارتقای کیفیت دادههای تولیدشده شد. در ادامه، مدل ToKA-BERT با این مجموعهداده فاینتیون گردید و نتایج ارزیابیها نشان داد که عملکرد آن در حوزه پرسش و پاسخ تخصصی گردشگری نسبت به مدلهای عمومی زبان فارسی برتری قابلتوجهی دارد. مجموعهداده و مدل نهایی در پلتفرم Hugging Face منتشر شده و میتوانند بهعنوان منبعی ارزشمند برای توسعه سیستمهای هوشمند در صنعت گردشگری و بهبود تجربه کاربران مورد استفاده قرار گیرند
کلیدواژهها
موضوعات
عنوان مقاله [English]
Proposing an Artificial Intelligence Model for Extracting Persian Question–Answer in the Tourism Domain
نویسندگان [English]
- Shirin Amini 1
- Mohamad Ali Afshar Kazemi 2
- Akbar Alam Tabriz 3
- Seyed Mohammad Ali Khatami Firouzabadi 4
1 Department of Information Technology Management, KI.C., Islamic Azad University, Kish, Iran
2 IDepartment of Industrial Management, CT.C., Islamic Azad University, Tehran, Iran Corresponding Author: moh.afsharkazemi@iauctb.ac.ir
3 Department of Industrial Management and Information Technology, Faculty of Management and Accounting, Shahid Beheshti University, Tehran, Iran
4 Department of Operations Management and Information Technology, Faculty of Management and Accounting, Allameh Tabataba’i University, Tehran, Iran
چکیده [English]
Given the importance of the tourism industry in the country’s economy and culture, and the shortage of domain-specific datasets for training Persian language models, this research aims to fill the existing gap. The objective of this study is to improve the performance of Persian tourism-related question answering systems and to provide accurate responses to queries concerning Iranian attractions and destinations. To achieve this, a collection of reliable tourism texts was gathered and processed, and using advanced natural language processing approaches, including Prompt Engineering, approximately 20,000 question–answer records were generated. In addition, the application of the innovative Answer Window strategy enhanced the quality of the produced data. Subsequently, the ToKA-BERT model was fine-tuned with this dataset, and evaluation results revealed its significant superiority in specialized tourism question answering compared to general-purpose Persian QA models. Both the dataset and the final model have been published on the Hugging Face platform and can serve as valuable resources for developing intelligent systems in the tourism industry and enhancing user experience.
کلیدواژهها [English]
- ToKA-BERT
- IntelligentQuestion-Answering
- Iranian Tourism
- Natural Language Processing (NLP)
- SQuAD-
- Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., Rosenberg, M., Song, X., Stoica, A., Tiwary, S., & Wang, T. (2018). MS MARCO: A human generated machine reading comprehension dataset (No. arXiv:1611.09268). arXiv. https://doi.org/10.48550/arXiv.1611.09268
- Fajcik, M., Jon, J., & Smrz, P. (2021). Rethinking the objectives of extractive question answering (No. arXiv:2008.12804, Version 4). arXiv. https://doi.org/10.48550/arXiv.2008.12804
- Farahani, M., Gharachorloo, M., Farahani, M., & Manthouri, M. (2021). ParsBERT: Transformer-based model for Persian language understanding. Neural Processing Letters, 53(6), 3831–3847. https://doi.org/10.1007/s11063-021-10528-4
- Joshi, M., Choi, E., Weld, D., & Zettlemoyer, L. (2017). TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In R. Barzilay & M.-Y. Kan (Eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1601–1611).AssociationforComputationalLinguistics. https://doi.org/10.18653/v1/P17-1147
- Kazemi, A., Zojaji, Z., Malverdi, M., Mozafari, J., Ebrahimi, F., Abadani, N., Varasteh, M. R., & Nematbakhsh, M. A. (2023). FarsNewsQA: A deep learning-based question answering system for the Persian news articles. Information Retrieval Journal, 26(1), 3. https://doi.org/10.1007/s10791-023-09417-2
- Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova,
- Masumi, M., Majd, S. S., Shamsfard, M., & Beigy, H. (2024). FaBERT: Pre-training BERT on Persian blogs (No. arXiv:2402.06617).arXiv. https://doi.org/10.48550/arXiv.2402.06617
- Mozafari, J., Kazemi, A., Moradi, P., & Nematbakhsh, M. A. (2022). PerAnSel: A novel deep neural network-based system for Persian question answering. Computational Intelligence and Neuroscience,2022(1),3661286. https://doi.org/10.1155/2022/3661286
- Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ questions for machine comprehension of text(No.arXiv:1606.05250).arXiv. https://doi.org/10.48550/arXiv.1606.05250
- Rosenthal, S., Sil, A., Florian, R., & Roukos, S. (2024). CLAPNQ: Cohesive long-form answers from passages in natural questions for RAG systems (No. arXiv:2404.02103). arXiv. https://doi.org/10.48550/arXiv.2404.02103
- SadraeiJavaheri,M.,Moghaddaszadeh,A.,Molazadeh,M.,Naeiji, F.,Aghababaloo,F.,Rafiee,H.,Amirmahani,Z.,Abedini,T.,Sheikhi, F. Z., & Salehoof, A. (2024). TookaBERT: A step forward forPersianNLU(No.arXiv:2407.16382).arXiv. https://doi.org/10.48550/arXiv.2407.16382
- Soh, Y. J., & Zhao, J. (2024). A step towards mixture of grader: Statistical analysis of existing automatic evaluation metrics (No. arXiv:2410.10030,Version1).arXiv. https://doi.org/10.48550/arXiv.2410.10030
- Taghizadeh, N., Doostmohammadi, E., Seifossadat, E., Rabiee, H. R., & Tahaei, M. S. (2021). SINA-BERT: A pre-trained language model for analysis of medical texts in Persian (No. arXiv:2104.07613).arXiv. https://doi.org/10.48550/arXiv.2104.07613
- Tran, S. Q., Do, G.-H., Do, P. N.-T., Kretchmar, M., & Du, X. (2023). AGent: A novel pipeline for automatically creating unanswerablequestions(No.arXiv:2309.05103).arXiv. https://doi.org/10.48550/arXiv.2309.05103
- Wang, Z., Ng, P., Ma, X., Nallapati, R., & Xiang, B. (2019). Multi-passage BERT: A globally normalized BERT model for open-domain question answering. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP-IJCNLP)(pp.5878–5882). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1599
- Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., & Manning, C. D. (2018). HotpotQA: A dataset for diverse, explainable multi-hop question answering. In E. Riloff, D. Chiang, J. Hockenmaier, & J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in NaturalLanguageProcessing(pp.2369–2380).Associationfor Computational Linguistics. https://doi.org/10.18653/v1/D18-1259