چکیده مقاله
TTS, or text to speech, is a complicated process thatcan be accomplished through appropriate modeling using deeplearning methods In order to implement deep learning models, asuitable dataset is required Since there is a scarce amount ofwork done in this field for the Persian language, this paper willintroduce the single speaker dataset: ArmanTTS We comparedthe characteristics of this dataset with those of various prevalentdatasets to prove that ArmanTTS meets the necessary standardsfor teaching a Persian text to speech conversion model We alsocombined the Tacotron 2 and HiFi GAN to design a model thatcan receive phonemes as input, with the output being thecorresponding speech 4 0 value of MOS was obtained from realspeech, 3 87 value was obtained by the vocoder prediction and2 98 value was reached with the synthetic speech generated bythe TTS model
کلیدواژهها
نویسندگان
شیوه ارجاع
Shamgholi, Mohammd Hasan and Saeedi, Vahid and Peymanfard, Javad and Alhabib, Leila and Zeinali, Hossein,1401,ArmanTTS single-speaker Persian dataset,1st International Conference and 6th National Conference on Computers, information technology and applications of artificial intelligence
ارائهشده در
مجموعه مقالات اولین کنفرانس بین المللی و ششمین کنفرانس ملی کامپیوتر، فناوری اطلاعات و کاربردهای هوش مصنوعی3 اسفند 1401 · اهواز