Speech/music audio classification for publicity insertion and DRM

dc.audience.educationlevelEstudis de primer/segon cicle
dc.audience.mediatorEscola d'Enginyeria de Telecomunicació i Aeroespacial de Castelldefels
dc.contributorTarrés Ruiz, Francisco
dc.contributor.authorGil Moreno, Fausto
dc.date.accessioned2018-02-26T16:40:11Z
dc.date.available2018-02-26T16:40:11Z
dc.date.issued2018-02-15
dc.date.updated2018-02-17T05:26:39Z
dc.description.abstractThe goal of this project is to develop, implement and optimize an existing method called Continuous Frequency Activation (CFA). The aim is to try to solve the problem that exists when advertising in randomly introduced in TV programmes/films/audio podcasts/etc. that can generate discomfort in the viewer. The basic idea is to avoid introducing adverts in the middle of a conversation. The final criteria will be selected taking into account metadata of video (change of plane, scene, fade-out, etc.) and audio (voice, music). To do that, we have developed and algorithm capable of discriminate between music and voice. This algorithm has been developed exclusively for this purpose and does not require base or date training to be trained. Previously to the creation of the algorithm, different existent methods of discrimination between music and voice have been studied and their pros and cons have been analysed. After performing the study, the method that has been selected is The Continuous Frequency Activation (CFA). CFA is one of the methods with better statistic results and it is not necessary to obtain large data bases for its training. The implementation of this algorithm has been performed using MATLAB®. Data base have been used in the realization of the trials, using five different musical style: classic music, Blues, electronical music, Jazz and Speech. The audio files from each different music style have been edited using the software called Audacity®. After performing all the tests, it can be said that the developed algorithm works correctly and it is able to discern music from voice in a very high percentage of cases (97.55%). With the results obtained after the trials, it can be said that this method could be used by companies that are involved in the fields of media, television (Antena 3, Telecinco, etc.) and/or audio podcast. The goal is to automatically introduce publicity in audio podcast format at the most appropriate moment.
dc.identifier.urihttps://hdl.handle.net/2117/114521
dc.language.isoeng
dc.publisherUniversitat Politècnica de Catalunya
dc.rights.accessOpen Access
dc.rights.urihttp://creativecommons.org/licenses/by-nc-sa/3.0/es/
dc.subjectÀrees temàtiques de la UPC::Enginyeria de la telecomunicació
dc.subjectÀrees temàtiques de la UPC::Informàtica::Aplicacions de la informàtica
dc.subject.lcshClassification--Music
dc.subject.lcshSound--Recording and reproducing
dc.subject.lemacSo -- Enregistrament i reproducció
dc.subject.lemacClassificació -- Música
dc.subject.otherDigital Audio
dc.subject.otherSignal Processing
dc.subject.otherAudio Descriptors
dc.subject.otherMachine Learning
dc.titleSpeech/music audio classification for publicity insertion and DRM
dc.typeMaster thesis
dspace.entity.typePublication

Fitxers

Paquet original

Mostrant 1 - 1 de 1
Carregant...
Miniatura
Nom:
memoria.pdf
Mida:
1.65 MB
Format:
Adobe Portable Document Format