چکیده مقاله
Effective retrieval of biomedical information presents a significant challenge due to terminological complexity and semantic ambiguity Traditional keyword based methods like BM25 often fail to capture the user's semantic intent To address this, we propose and empirically evaluate a multi stage ranking architecture designed for high precision retrieval Our pipeline initiates with two parallel retrieval stages: a sparse lexical retriever BM25 and a dense semantic retriever using a Bi Encoder model multi qa MiniLM L6 cos v1 The resulting candidate lists are then fused using Reciprocal Rank Fusion RRF to leverage their complementary strengths In the final stage, a more powerful Cross Encoder model ms marco MiniLM L 6 v2 re ranks the top 100 candidates from the fused list to achieve fine grained relevance scoring Evaluated on the standard TREC COVID dataset, our complete pipeline demonstrates substantial performance gains at each stage, culminating in a final Precision@10 of 0 808 and an nDCG@10 of 0 754 This represents a significant relative improvement of 68% and 69%, respectively, over the BM25 baseline These results validate the efficacy of a cascaded retrieve fuse rerank architecture Our work underscores the synergistic value of combining sparse, dense, and cross attention models, providing a robust framework for developing high performance information retrieval systems in specialized domains
کلیدواژهها
نویسندگان
شیوه ارجاع
Shabanian, Asa and Asl Nemati, Alireza and Mohammadi Zanjireh, Morteza,1404,A Multi-Stage Ranking Pipeline for High-Precision Medical Information Retrieval,The Second National Conference on the Era of Technology Explosion: Artificial Intelligence, a Transformation in Industry, Trade, and Supply Chain,Tabriz
ارائهشده در
مجموعه مقالات دومین کنفرانس ملی عصر انفجار تکنولوژی؛ هوش مصنوعی، تحولی در صنعت، تجارت و زنجیره تامین17 مهر 1404 · تبریز