چکیده مقاله
The rapid advancement of generative artificial intelligence has enabled the creation of highly realistic synthetic media, commonly known as deepfakes These manipulated multimedia contents, particularly videos and images, pose significant threats to information integrity, personal privacy, public trust, and national security Detecting deepfakes has therefore emerged as one of the most pressing challenges in modern multimedia systems research This paper proposes a novel hybrid framework for deepfake detection that combines the feature extraction capabilities of Convolutional Neural Networks CNNs with the long range dependency modeling power of Transformer architectures The proposed model operates on facial regions extracted from video frames and leverages multi scale spatial features alongside self attention mechanisms to capture both local texture artifacts and global contextual inconsistencies introduced during the synthesis process We evaluate the proposed approach on three widely used benchmark datasets: FaceForensics , Celeb DF, and DFDC Experimental results demonstrate that the hybrid CNN Transformer model achieves detection accuracy of 97 4% on FaceForensics , outperforming several state of the art methods while maintaining competitive performance across cross dataset evaluation scenarios The paper also provides an analysis of failure cases, discusses the challenges posed by highly compressed and low resolution deepfakes, and outlines future research directions for developing more robust and generalizable detection systems
کلیدواژهها
نویسندگان
شیوه ارجاع
Behboodi, Amir Masoud,1405,AI-Based Deepfake Detection in Multimedia Content Using CNN and Transformer Models,The 9th international conference on artificial intelligence and its future prospects in electrical, computer, mechanical and telecommunication engineering sciences,Mashhad
ارائهشده در
مجموعه مقالات نهمین کنفرانس بین المللی هوش مصنوعی و چشم انداز آینده آن در علوم مهندسی برق، کامپیوتر، مکانیک و مخابرات16 خرداد 1405 · مشهد