Efficient Motion Detection: A Machine Learning-Based Compressed Domain Approach for Real-Time Processing
| dc.audience.degree | MÀSTER UNIVERSITARI EN INNOVACIÓ I RECERCA EN INFORMÀTICA (Pla 2012) |
| dc.audience.educationlevel | Màster |
| dc.audience.mediator | Facultat d'Informàtica de Barcelona |
| dc.contributor | Herrero Zaragoza, José Ramón |
| dc.contributor | Eid, Majd |
| dc.contributor.author | Petrilli, Stefano |
| dc.contributor.other | Universitat Politècnica de Catalunya. Departament d'Arquitectura de Computadors |
| dc.date.accessioned | 2026-01-07T14:01:56Z |
| dc.date.issued | 2025-07-03 |
| dc.date.updated | 2025-07-23T09:35:24Z |
| dc.description.abstract | In Video Management Systems (VMSs), rising camera resolutions and complex vision algorithms make computational resources scarce. Video compression codecs like H.264 are essential, as uncompressed streams are impractical for transfer and storage. However, video analysis requires decoding, adding significant overhead and limiting the scalability of current state-of-the-art solutions. This work presents a method to predict motion directly from the compressed video stream without relying on the pixel values of the video. The technique presented makes use of the metadata included in the video compressed video stream (more specifically motion vectors) and uses machine learning to classify if the metadata in the frame correspond to motion. The focus of this work is on the efficient use of computational resources and on improving scalability. The experimental results reported in this work demonstrate that the proposed compressed-domain motion detection framework offers a viable alternative to the more resource intensive alternatives currently in use. The capability of the presented method to identify motion is remarkable, the recall obtained on the testing dataset is 100% which means that the machine learning model can always identify motion just using the metadata extracted from the stream. Additionally, the experimental results also demonstrate that this method allows for a remarkable increase in speed as the modified decoder proposed in this work is twice as fast as the industry standard FFmpeg decoder and six times as fast as the H264Bsd baseline decoder. In practice, this means that the decoding of the compressed video can be replaced by the extraction of the motion vectors which is proven to be faster than a full decoding. The same metadata present in H.264 and used in the proposed solution are also present in other video compression algorithms which makes this technique possibly applicable also to H.265 and H.266 compressed video streams. This thesis also includes an environmental impact analysis and ethical implications of the proposed solution which highlights the potential positive impact on the CO2 emissions and on the potential positive contribution to security. |
| dc.identifier.slug | 197164 |
| dc.identifier.uri | https://hdl.handle.net/2117/449733 |
| dc.language.iso | eng |
| dc.publisher | Universitat Politècnica de Catalunya |
| dc.rights.access | Open Access |
| dc.subject | Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic |
| dc.subject.lcsh | Video compression |
| dc.subject.lcsh | Computer vision |
| dc.subject.lcsh | Machine learning |
| dc.subject.lemac | Compressió de vídeo |
| dc.subject.lemac | Visió per ordinador |
| dc.subject.lemac | Aprenentatge automàtic |
| dc.title | Efficient Motion Detection: A Machine Learning-Based Compressed Domain Approach for Real-Time Processing |
| dc.type | Master thesis |
| dspace.entity.type | Publication |
Fitxers
Paquet original
1 - 1 de 1



