Laporkan Masalah

Perbandingan Human Activity Recognition YOWOv3 Berbasis You Only Look Once dengan Pergantian Modul Dataloader untuk Efisiensi Performa Training

Raden Mas Benediktus Suryo Wicaksono, Dr. I Wayan Mustika, S.T., M.Eng.; Prof. Ir. Selo, S.T., M.T., M.Sc., Ph.D., IPU, ASEAN Eng.

2026 | Tesis | S2 Teknologi Informasi

Human Activity Recognition (HAR) adalah bidang penting dalam kecerdasan buatan yang bertujuan mendeteksi dan mengklasifikasikan aktivitas manusia secara otomatis. HAR memiliki banyak aplikasi, seperti keamanan, kesehatan, interaksi manusia-komputer, dan lingkungan cerdas. Pendekatan HAR bisa menggunakan kamera melalui visi komputer (CV). Kini, banyak penelitian HAR memanfaatkan Deep Learning (DL), terutama Convolutional Neural Network (CNN) berbasis 2D dan 3D, yang efektif menangkap data spasial (gambar) dan temporal (urutan gerakan) yang digabung menjadi spasioetmporal. Salah satu metode deteksi objek real-time yang banyak digunakan adalah YOLO (You Only Look Once), dengan versi Ultralytics di YOLOv8 dan YOLOv11. Untuk CV HAR spasiotemporal, arsitektur YOWOv3 menggabungkan deteksi spasial dari YOLO dengan analisis temporal dari CNN 3D. YOWOv3 adalah berbasis dataloader Torchvision dan PIL yang mana didesain untuk visualisasi data, tetapi kombinasi kedua modul memperlambat waktu training. Selain itu, YOLOv11 masih belum dibuktikan secara terperinci untuk dimasukkan ke YOWOv3. Dengan itu, proses pemuatan data diganti memakai modul Albumentations dan YOLOv8 diganti ke YOLOv11 sebagai perbandingan. Perbandingan ini berkontribusi terhadap pengembangan YOWOv3, salah satunya menuju efisiensi program. Evaluasi dilakukan dengan dataset MSCOCO untuk YOLO, dan AVAv2.2 untuk YOWOv3. Hasil modifikasi menunjukkan waktu pelatihan YOWOv3 turun drastis dari 1:10:00 menjadi 0:48:35. Namun, pergantian dataloader menghasilkan peningkatan loss, terutama saat menggantikan YOLOv8 ke YOLOv11. Dalam 9 epoch, YOWOv3 dengan YOLOv8 menghasilkan loss 7.933, sedangkan YOLOv11 8.0369. Ini menunjukkan tidak selamanya pergantian dataloader memenuhi efisiensi keseluruhan di mana terdapat tradeoff dari dataloader karena pendekatan optimisasi dataloader terhadap arsitektur dan hardware yang berbeda.

Human Activity Recognition (HAR) is a vital field in artificial intelligence, supporting applications in security, healthcare, human-computer interaction, and smart environments. HAR aims to automatically identify and classify human activities using either sensor-based or computer vision-based methods. Recent progress highlights deep learning, particularly Convolutional Neural Networks (CNNs), 2D and 3D, which are the key for comprehensive spatiotemporal modeling of human actions. The Ultralytics YOLO series, especially YOLOv8 and YOLOv11, is widely used. Integrating YOLO’s spatial detection with 3D CNN-based temporal analysis, the YOWOv3 architecture offers advanced capabilities for spatiotemporal HAR. However, performance-wise, YOWOv3 is still slow especially for the training time. The use of Torchvision +PIL dataloader causes performance bottlenecks that affect training time. On the other hand, YOLOv11 has not been thoroughly measured in YOWOv3. With that, this work proposes performance comparison to YOWOv3 by comparing YOLOv8 and YOLOv11 to YOWOv3 and changing dataloader between Torchvision+PIL dan Albumentations. This contributes to YOWOv3 development for performance efficiency. The evaluation uses the MSCOCO dataset for YOLO and the AVAv2.2 dataset for YOWOv3, ensuring thorough benchmarking across varied activity domains. Results show that the comparison has significant reduction in YOWOv3 training time with, from 1:10:00 to 0:48:35, using Albumentations and changing from YOLOv8 to YOLOv11. However, there is a tradeoff where in 9 epochs YOWOv3 using YOLOv11 has higher loss at 8.0369, compared with YOLOv8 at 7.933. This gives insight not every efficient dataloader can improve model performance, as each dataloader has different integration process between model architecture and hardware.

Kata Kunci : computer vision, object detection, human activity recognition, HAR, spatiotemporal, Ultralytics

  1. S2-2026-509455-abstract.pdf  
  2. S2-2026-509455-bibliography.pdf  
  3. S2-2026-509455-tableofcontent.pdf  
  4. S2-2026-509455-title.pdf