چکیده مقاله
Document clustering is critical to information retrieval IR as it enhances user navigation, semantic organization, and exploration of large text collections Current clustering techniques, though, are marred by poor accuracy and semantic inconsistency, with many misclassifying relevant documents as noise and using superficial textual representations This study aims to develop a clustering pipeline that produces semantically meaningful and structurally coherent groups of documents to support more effective IR We propose a method that combines SBERT embeddings for deep semantic representation, UMAP for structure preserving dimensionality reduction, and HDBSCAN for flexible, density based clustering without needing to predefine the number of clusters Experimental evaluations on the 20 Newsgroups dataset reveal that our optimal setting with the paraphrase mpnet base v2 model obtains a Silhouette Score of 0 6853, ARI of 0 7865, and NMI of 0 8186 These results illustrate the promise of embedding based clustering methods to greatly improve the interpretability and effectiveness of IR systems on real world text collections
کلیدواژهها
نویسندگان
شیوه ارجاع
Mohammadiha, Mahdi and Sadreddini, Mohammad Hassan and Mohammadi Zanjireh, Morteza,1404,Document Clustering Using Deep Pre-trained Language Model Embeddings for Information Retrieval,The Second National Conference on the Era of Technology Explosion: Artificial Intelligence, a Transformation in Industry, Trade, and Supply Chain,Tabriz
ارائهشده در
مجموعه مقالات دومین کنفرانس ملی عصر انفجار تکنولوژی؛ هوش مصنوعی، تحولی در صنعت، تجارت و زنجیره تامین17 مهر 1404 · تبریز