Laporkan Masalah

Implementasi Analisis Data Honeypot Berbasis YARA dan Large Language Model Untuk Identifikasi Ancaman Siber

Emanuel Dananjaya Kusumadewa Biansa Dua, Ir. Nur Rohman Rosyid, S.T., M.T., D.Eng., IPM.

2026 | Tugas Akhir | D4 TEKNOLOGI JARINGAN

Meningkatnya ancaman siber mendorong penggunaan Honeypot sebagai sistem pendeteksi serangan. Namun, data yang dihasilkan pada log serangan T-Pot Honeypot belum dapat dianalisis secara kontekstual untuk mengungkap karakteristik malware dan pola serangan secara menyeluruh. Penelitian ini mengembangkan sistem pipeline terintegrasi yang mencakup pengumpulan dan sinkronisasi data Honeypot menggunakan mekanisme dual-output, analisis artefak berbahaya menggunakan YARA dan Large Language Model (LLM) untuk klasifikasi serangan, serta visualisasi hasil analisis dalam bentuk dasbor pemantauan menggunakan ELK Stack. Hasil pengujian sinkronisasi menunjukkan integritas data terjaga sempurna dengan selisih berkas antar server hanya 0,17?ri total 57 GB, deteksi YARA mencapai tingkat deteksi 46,2?ri sampel berkas yang berisi data Honeypot dengan 26?rkas berhasil diidentifikasi sebagai ancaman yang tidak terdeteksi VirusTotal, serta penggunaan sumber daya server tercatat sebesar 5,93% CPU dan 5,13 GB RAM pada skenario sinkronisasi Rsync, dan 1,27% CPU dan 4,9 GB RAM pada skenario pengiriman Logstash. Hasil penelitian ini membuktikan bahwa proses pengumpulan dan analisis data Honeypot menggunakan YARA dan LLM mampu memberikan informasi tambahan dalam bentuk visualisasi dasbor yang berguna bagi analis keamanan dalam mengidentifikasi ancaman.

The increasing prevalence of cyber threats has driven the adoption of Honeypots as an attack detection system. However, the data generated from T-Pot Honeypot attack logs has not been able to be analyzed contextually to reveal malware characteristics and attack patterns comprehensively. This research develops an integrated pipeline system that encompasses Honeypot data collection and synchronization using a dual-output mechanism, malicious artifact analysis using YARA and Large Language Model (LLM) for attack classification, as well as visualization of analysis results in the form of a monitoring dashboard using the ELK Stack. The synchronization test results show that data integrity is perfectly maintained with a file difference between servers of only 0.17% out of a total of 57 GB, YARA detection achieves a detection rate of 46.2% from sample files containing Honeypot data with 26% of files successfully identified as threats undetected by VirusTotal, and server resource usage recorded at 5.93% CPU and 5.13 GB RAM in the Rsync synchronization scenario, and 1.27% CPU and 4.9 GB RAM in the Logstash transmission scenario. The results of this research prove that the process of collecting and analyzing Honeypot data using YARA and LLM is able to provide additional information in the form of dashboard visualizations that are useful for security analysts in identifying threats.

Kata Kunci : Honeypot, YARA, Large Language Model, MITRE ATT&CK, Threat Intelligence

  1. D4-2026-506099-abstract.pdf  
  2. D4-2026-506099-bibliography.pdf  
  3. D4-2026-506099-tableofcontent.pdf  
  4. D4-2026-506099-title.pdf