An Approach for Audio-Visual Content Understanding of Video using Multimodal Deep Learning Methodology

Boztepe, Emre; Karakaya, Bedirhan; Karasulu, BAHADIR; Ünlü, İsmet

doi:10.35377/saucis...1139765

An Approach for Audio-Visual Content Understanding of Video using Multimodal Deep Learning Methodology

Atıf İçin Kopyala

Boztepe E. B., Karakaya B., Karasulu B., Ünlü İ.

Sakarya University Journal of Computer and Information Sciences, cilt.5, sa.2, ss.181-207, 2022 (Hakemli Dergi)

Yayın Türü: Makale / Tam Makale
Cilt numarası: 5 Sayı: 2
Basım Tarihi: 2022
Doi Numarası: 10.35377/saucis...1139765
Dergi Adı: Sakarya University Journal of Computer and Information Sciences
Derginin Tarandığı İndeksler: TR DİZİN (ULAKBİM), Index Copernicus
Sayfa Sayıları: ss.181-207
Çanakkale Onsekiz Mart Üniversitesi Adresli: Evet

Özet

This study contains an approach for recognizing the sound environment class from a video to understand the spoken content with its sentimental context via some sort of analysis that is achieved by the processing of audio-visual content using multimodal deep learning methodology. This approach begins with cutting the parts of a given video which the most action happened by using deep learning and this cutted parts get concanarated as a new video clip. With the help of a deep learning network model which was trained before for sound recognition, a sound prediction process takes place. The model was trained by using different sound clips of ten different categories to predict sound classes. These categories have been selected by where the action could have happened the most. Then, to strengthen the result of sound recognition if there is a speech in the new video, this speech has been taken. By using Natural Language Processing (NLP) and Named Entity Recognition (NER) this speech has been categorized according to if the word of a speech has connotation of any of the ten categories. Sentiment analysis and Apriori Algorithm from Association Rule Mining (ARM) processes are preceded by identifying the frequent categories in the concanarated video and helps us to define the relationship between the categories owned. According to the highest performance evaluation values from our experiments, the accuracy for sound environment recognition for a given video's processed scene is 70%, average Bilingual Evaluation Understudy (BLEU) score for speech to text with VOSK speech recognition toolkit's English language model is 90% on average and for Turkish language model is 81% on average. Discussion and conclusion based on scientific findings are included in our study.