چکیده مقاله
Speech recognition from visual data is in important step towards communication when audio is not available This paper considers several hand crafted features including HOG, MBH, DCT, LBP, MTC, and their combinations for recognizing speech from a sequence of images Several classifiers including SVM, decision trees, K nearest neighbor algorithm and the sub space K nearest algorithm were tested feature evaluation Further, the application of PCA for dimensionality reduction was considered in this study Two sets of tests were carried out in this study: lip pose recognition and recognition of isolated words For evaluation, the MIRACL VC1 data set was considered Self dependent tests reached an accuracy of over 95% while in the self independent tests, the maximum accuracy of recognition was about 52%
کلیدواژهها
نویسندگان
شیوه ارجاع
Jafari Sheshpoli, Ali and Nadian-Ghomsheh, Ali,1396,Temporal and Spatial Features for Visual Speech Recognition,Fifth International Conference on Electrical and Computer Engineering with Emphasis on Indigenous Knowledge,Tehran
ارائهشده در
مجموعه مقالات پنجمین کنفرانس بین المللی مهندسی برق و کامپیوتر با تاکید بر دانش بومی19 بهمن 1396 · تهران