مدل سازی در مهندسی

مدل سازی در مهندسی

استخراج موضوع از متون فارسی با استفاده از چهارچوب BERTopic، مدل‌های تعبیه زبانی و خوشه بندی متن

نوع مقاله : مقاله کامپیوتر

نویسندگان
دانشکده مهندسی برق و کامپیوتر، دانشگاه سمنان، سمنان، ایران
چکیده
با رشد اطلاعات، استخراج دانش از مجموعه‌های متنی اهمیت یافته است. استخراج موضوع، تکنیکی بدون نظارت در یادگیری ماشین است که مضامین پنهان اسناد را کشف می‌کند. در این مقاله، با الهام از BERTopic، روشی بدون نظارت برای استخراج موضوع از متون فارسی ارائه شده است. روش پیشنهادی از مدل تعبیه‌زبانی LaBSE برای تبدیل متون به بردار‌های تعبیه استفاده می‌کند. سپس، با استفاده ازUMAP، ابعاد بردارهای تعبیه را کاهش می‌دهد و پس از آن، با استفاده از الگوریتم خوشه‌بندی K-Means، متون مشابه را در خوشه‌های یکسان قرار می‌دهد. سپس با تشکیل ماتریس خوشه-توکن و تکنیک بازنمایی موضوعات، موضوعات مختلف را به ازای هر خوشه از متن استخراج می‌کند. ما LaBSE را با مدل‌های تعبیه‌زبانی XLM-R ،ParsBERT،Paraphrase-multilingual-MiniLM-L12-v2 ،Shiraz و HooshvareLab (RoBERTa) مقایسه کردیم. همچنین مقایسه‌ای بین الگوریتم‌های K-Means و HDBSCAN انجام دادیم. برای ارزیابی روش پیشنهادی، از مجموعه داده عصر ایران استفاده شد. معیار انسجام (NPMI) و معیار ارزیابی انسانی عملکرد روش‌ پیشنهادی را تأیید کردند. در الگوریتمHDBSCAN ، مدلHooshvare (RoBERTa) بر اساس معیار انسجام، و مدل‌ ParsBERT بر اساس ارزیابی انسانی بهترین نتایج را ارائه داد. در K-Means، مدل Paraphrase-multilingual- MiniLM-L12-v2 مطابق معیار انسجام و LaBSE مطابق ارزیابی انسانی، نتایج بهتری داشت. برتری K-Means نسبت به HDBSCAN نیز تأیید شد. همچنین با استفاده از دو مجموعه داده عصر ایران و تسنیم به صورت جداگانه، روش پیشنهادی با مدل‌های فاکتورسازی ماتریس غیرمنفی، تخصیص دیریکله پنهان و تحلیل معنایی پنهان مقایسه شد. نتایج مقایسه، عملکرد برجسته روش پیشنهادی را نشان می‌دهد.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Topic Extraction from Persian Texts Using BERTopic Framework, Language Embedding Models, and Text Clustering

نویسندگان English

Zahra Fallahi
Mohammad Rahmanimanesh
Faculty of Electrical and Computer Engineering, Semnan University, Semnan, Iran
چکیده English

With the growth of information, extracting knowledge from textual collections has become essential. Topic modeling is an unsupervised machine learning technique that uncovers the hidden themes in documents. In this paper, inspired by BERTopic, we present an unsupervised method for topic modeling on Persian texts. The proposed approach employs the LaBSE language embedding model to convert texts into embedding vectors, then reduces their dimensions using UMAP, and finally groups similar texts into clusters using the K-Means algorithm. Next, by forming a cluster-token matrix and applying a topic representation technique, various topics are extracted from each cluster. We compared LaBSE model with other language embedding models including XLM-R, ParsBERT, Paraphrase-multilingual-MiniLM-L12-v2, Shiraz, and HooshvareLab (RoBERTa). We also compared the K-Means and HDBSCAN clustering algorithms. For evaluation, the AsreIran dataset was used, and both the coherence evaluation metric (NPMI) and human evaluation confirmed the proposed method’s performance. In HDBSCAN, Hooshvare (RoBERTa) yielded the best coherence, while ParsBERT excelled in human evaluation. In K-Means, Paraphrase-multilingual-MiniLM-L12-v2 performed best in terms of coherence and LaBSE in human evaluation. The superiority of K-Means over HDBSCAN was also verified. Furthermore, using the AsreIran and Tasnim datasets separately, the proposed method was compared with non-negative matrix factorization, latent Dirichlet allocation, and latent semantic analysis, with results demonstrating its outstanding performance.

کلیدواژه‌ها English

Topic extraction
BERTopic
NMF
LDA
LSA
Persian text
[1] Moradi, Mona, Mohammad Rahmanimanesh, and Ali Shahzadi. "On evaluating the collaborative research areas: A case study." Journal of King Saud University-Computer and Information Sciences 34, no. 2 (2022): 408-420.
[2] Hofmann, Thomas. "Unsupervised learning by probabilistic latent semantic analysis." Machine Learning 42, no. 1 (2001): 177-196.
[3] Blei, David M., Andrew Y. Ng, and Michael I. Jordan. "Latent dirichlet allocation." Journal of Machine Learning Research 3, no. Jan (2003): 993-1022.
[4] Lee, Daniel D., and H. Sebastian Seung. "Learning the parts of objects by non-negative matrix factorization." Nature 401, no. 6755 (1999): 788-791.
[5] Daneshfar, Fatemeh, Amin Golzari Oskouei, Maryam Dorosti, and Mohammad Javad Aghajani. "An Improved Deep Text Clustering via Local Manifold of an Autoencoder Embedding." Journal of Modeling in Engineering 23, no. 80 (2025): 307-324.
[6] Moodi, Fatemeh, and Fatemeh Sadri. "Aspect-based sentiment analysis based on users' comments in an online marketplace." Journal of Modeling in Engineering (2025).doi:10.22075/jme.2025.33944.2657.
[7] Grootendorst, Maarten. "BERTopic: Neural topic modeling with a class-based TF-IDF procedure." arXiv preprint arXiv:2203.05794 (2022). [Online]. Available: http://arxiv.org/abs/2203.05794
[8] McInnes, L., J. Healy, N. Saul, and L. Großberger. "UMAP: Uniform manifold approximation and projection." Journal of Open Source Software 3, no. 29, (2018): 861.
[9] Wang, Zhongyi, Jing Chen, Jiangping Chen, and Haihua Chen. "Identifying interdisciplinary topics and their evolution based on BERTopic." Scientometrics 129, no. 11 (2024): 7359-7384.
[10] Egger, Roman, and Joanne Yu. "A topic modeling comparison between lda, nmf, top2vec, and bertopic to demystify twitter posts." Frontiers in sociology 7 (2022): 886498.
[11] Kermani, Hossein. "Framing the pandemic on persian twitter: gauging networked frames by topic modeling." American Behavioral Scientist 69, no. 10, (2023), doi: 10.1177/00027642231207078.
[12] Ghasemi, Saeideh, and Amir H. Jadidinejad. "Persian text classification via character-level convolutional neural networks." In 2018 8th Conference of AI & Robotics and 10th RoboCup Iranopen International Symposium (IRANOPEN), pp. 1-6. IEEE, 2018.
[13] Gholizade, Masoume, Hadi Soltanizadeh, and Mohammad Rahmanimanesh. "A Survey of Transfer Learning and Categories." Modeling and Simulation in Electrical and Electronics Engineering 1, no. 3 (2021): 17-25.
[14] Gholizade, Masoume, Hadi Soltanizadeh, Mohammad Rahmanimanesh, and Shib Sankar Sana. "A review of recent advances and strategies in transfer learning." International Journal of System Assurance Engineering and Management 16, (2025): 1123–1162.
[15] Moradi, Mona, Mohammad Rahmanimanesh, and Ali Shahzadi. "Transfer learning for concept drifting data streams in heterogeneous environments." Knowledge and Information Systems 66, no. 5 (2024): 2799-2857.
[16] Ghorbanali, Alireza, Mohammad Karim Sohrabi, and Farzin Yaghmaee. "Multiple transfer learning-based multimodal sentiment analysis using weighted convolutional neural network ensemble." Journal of Modeling in Engineering 21, no. 72 (2023): 83-97.
[17] Moradi, Mona, Mohammad Rahmanimanesh, Ali Shahzadi, and Reza Monsefi. "Smooth unsupervised domain adaptation considering uncertainties." Information Sciences 648 (2023): 119602.
Moradi, Mona, Mohammad Rahmanimanesh, and Ali Shahzadi. "Unsupervised domain adaptation by incremental learning for concept drifting data streams." International Journal of Machine Learning and Cybernetics 15, no. 9 (2024): 4055-4078.
[19] Moradi, Mona, Mohammad Rahmanimanesh, and Ali Shahzadi. "Learning from streaming data with unsupervised heterogeneous domain adaptation." International Journal of Data Science and Analytics 19, no. 1 (2025): 61-81.
[20] Rashidi, Dalile, Mohammad Rahmanimanesh, and Mohsen Shafiei Nikabadi. "Fuzzy Logic-Based Unsupervised Sentiment Analysis and Opinion Mining: Applications in Market Research." New Marketing Research Journal 14, no. 1 (2024): 127-146.
[21] Anoop, Vijayalekshi Sundaresan, S. Asharaf, and P. Deepak. "Unsupervised concept hierarchy learning: a topic modeling guided approach." Procedia computer science 89 (2016): 386-394.
[22] Ebrahimi, Fezzeh, Mohammad Dehghani, and Fatemah Makkizadeh. "Analysis of persian bioinformatics research with topic modeling." BioMed Research International 2023, no. 1 (2023): 3728131.
[23] Amirkhani, Hossein, Mohammad AzariJafari, Soroush Faridan-Jahromi, Zeinab Kouhkan, Zohreh Pourjafari, and Azadeh Amirak. "Farstail: A persian natural language inference dataset." Soft Computing (2023): 1-13.
[24] Farahani, Mehrdad, Mohammad Gharachorloo, Marzieh Farahani, and Mohammad Manthouri. "Parsbert: Transformer-based model for persian language understanding." Neural Processing Letters 53, no. 6 (2021): 3831-3847.
[25] Abuzayed, Abeer, and Hend Al-Khalifa. "BERT for Arabic topic modeling: An experimental study on BERTopic technique." Procedia computer science 189 (2021): 191-194.
[26] Xiao, Kejing, Zhaopeng Qian, and Biao Qin. "A graphical decomposition and similarity measurement approach for topic detection from online news." Information Sciences 570 (2021): 262-277.
[27] Garcia, Klaifer, and Lilian Berton. "Topic detection and sentiment analysis in Twitter content related to COVID-19 from Brazil and the USA." Applied soft computing 101 (2021): 107057.
[28] Dehghani, Mohammad, and Fezzeh Ebrahimi. "ParsBERT topic modeling of Persian scientific articles about COVID-19." Informatics in Medicine Unlocked 36 (2023): 101144.
[29] Asnawi, Mohammad Hamid, Anindya Apriliyanti Pravitasari, Tutut Herawan, and Triyani Hendrawati. "The combination of contextualized topic model and MPNet for user feedback topic modeling." IEEE Access 11 (2023): 130272-130286.
[30] Samsir, Samsir, Reagan Surbakti Saragih, Selamat Subagio, Rahmad Aditiya, and Ronal Watrianthos. "BERTopic modeling of natural language processing abstracts: Thematic structure and trajectory." Jurnal Media Informatika Budidarma 7, no. 3 (2023): 1514-1520.
[31] Ranjbar-Khadivi, Mehrdad, Shahin Akbarpour, Mohammad-Reza Feizi-Derakhshi, and Babak Anari. "Persian topic detection based on Human Word association and graph embedding." arXiv preprint arXiv:2302.09775 (2023).
[32] Wu, Xiaobao, Thong Nguyen, and Anh Tuan Luu. "A survey on neural topic models: methods, applications, and challenges." Artificial Intelligence Review 57, no. 2 (2024): 18.
[33] Khazeni, Mohsen, Mohammad Heydari, and Amir Albadvi. "Persian Slang Text Conversion to Formal and Deep Learning of Persian Short Texts on Social Media for Sentiment Classification." arXiv preprint arXiv:2403.06023 (2024).
[34] Ahmadi, Parvin, Mahmoud Tabandeh, and Iman Gholampour. "Persian text classification based on topic models." In 2016 24th Iranian Conference on Electrical Engineering (ICEE), pp. 86-91. IEEE, 2016.
[35] McInnes, L. "How HDBSCAN Works." [Online]. Available:https://hdbscan.readthedocs.io/en/latest/how_hdbscan_works.html. [Accessed: Aug. 8, 2024].
[36] Grootendorst, M. "Topic reduction." BERTopic documentation, [Online]. Available: https://maartengr.github.io/BERTopic/getting_started/topicreduction/topicreduction.html. [Accessed: Aug. 8, 2024].
[37] HooshvareLab. "roberta-fa-zwnj-base: A pretrained RoBERTa model for persian text." Hugging Face, 2024. [Online]. Available: https://huggingface.co/HooshvareLab/roberta-fa-zwnj-base. [Accessed: Aug. 8, 2024].
[38] Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. "Roberta: A robustly optimized bert pretraining approach." arXiv preprint arXiv:1907.11692 (2019).
[39] Sun, Zhiqing, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. "Mobilebert: a compact task-agnostic bert for resource-limited devices." arXiv preprint arXiv:2004.02984 (2020).
[40] R. SalehiChegeni, "Lifeweb AI team, Lifeweb language models," GitHub, (2024). [Online]. Available: https://github.com/lifeweb-ir/LM. [Accessed: June. 8, 2024].
[41] Feng, Fangxiaoyu, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. "Language-agnostic BERT sentence embedding." (2020), [Online]. Available: http://arxiv.org/abs/2007.01852
[42] Takahashi, S. "paraphrase-multilingual-MiniLM-L12-v2," GitHub, (2024). [Online]. Available: https://github.com/shinichiro-takahashi-sbr/paraphrase-multilingual-MiniLM-L12-v2/blob/main/README.md. [Accessed: Aug. 8, 2024].
[43] Reimers, Nils, and Iryna Gurevych. "Sentence-bert: Sentence embeddings using siamese bert-networks." arXiv preprint arXiv:1908.10084 (2019).
[44] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. "Unsupervised cross-lingual representation learning at scale." (2019), [Online]. Available: http://arxiv.org/abs/1911.02116
[45] Ghafouri, Arash, Mohammad Amin Abbasi, and Hassan Naderi. "AriaBERT: A pre-trained Persian BERT model for natural language understanding." (2023).
[46] Farahani, M. “m3hrdadfi/sentence-transformers.” Accessed: Aug. 08, 2024. [Online]. Available: https://github.com/m3hrdadfi/sentence-transformers
[47] MacQueen, J. "Multivariate observations." In Proceedings ofthe 5th Berkeley Symposium on Mathematical Statisticsand Probability, vol. 1, pp. 281-297. 1967.
[48] Gandhi, A. "Topic modeling with LSA, PLSA, LDA & lda2Vec." Jun. 2021, Accessed: Aug. 08, 2024. [Online]. Available: https://nanonets.com/blog/topic-modeling-with-lsa-plsa-lda-lda2vec/
[49] Oshriyeh, O. "Applied data science in tourism: Interdisciplinary approaches, methodologies, and applications." Information Technology & Tourism 25, no. 1, (2023):133–136, doi: 10.1007/s40558-023-00243-2.
[50] Pourmand, A. "TasnimNews Dataset (Farsi - Persian)," Kaggle, (2022). [Online]. Available: https://www.kaggle.com/datasets/amirpourmand/tasnimdataset. [Accessed: July 15, 2024].
[51] HooshvareLab. "Persian News," (2020). [Online]. Available: https://hooshvare.github.io/docs/datasets/tc#persian-news. [Accessed: July 30, 2024].
[52] Bouma, Gerlof. "Normalized (pointwise) mutual information in collocation extraction." Proceedings of GSCL 30 (2009): 31-40.
[53] Dieng, Adji B., Francisco JR Ruiz, and David M. Blei. "Topic modeling in embedding spaces." (2019), [Online]. Available: http://arxiv.org/abs/1907.04907
[54] Cai, Guoray, Feng Sun, and Yongzhong Sha. "Interactive Visualization for Topic Model Curation." In IUI Workshops. 2018.
[55] Chen, Yong, Hui Zhang, Rui Liu, Zhiwen Ye, and Jianying Lin. "Experimental explorations on short text topic mining between LDA and NMF based Schemes." Knowledge-Based Systems 163 (2019): 1-13.
 
دوره 23، شماره 83
زمستان 1404
صفحه 217-235

  • تاریخ دریافت 16 دی 1403
  • تاریخ بازنگری 08 فروردین 1404
  • تاریخ پذیرش 17 فروردین 1404