LightGBM-Based Intrusion Detection on the UNSW-NB15 Dataset: Feature Selection and Evaluation Across Validation Protocols
- Authors
-
-
Sameeruddin Shaik
Author
-
- Keywords:
- Network Intrusion Detection, LightGBM, UNSW-NB15, Feature Selection, Validation Protocols, Machine Learning
- Abstract
-
Network intrusion detection systems face a fundamental challenge: studies using the same benchmark dataset are difficult to compare because each adopts different validation protocols. This paper investigates how evaluation methodology affects reported intrusion detection performance. Using UNSW-NB15, we train a LightGBM classifier with feature engineering, preprocessing, and RFECV-based feature selection, then apply the same model across every evaluation protocol identified in the literature. By controlling for model architecture while varying only the validation approach, we quantify the performance variability introduced by methodological choices. Our evaluation preserves the original UNSW-NB15 train/test split (170,000 training, 90,000 test instances). The independent test set evaluation achieves 92.72% F1 (91.73% accuracy, 94.72% recall, 90.80% precision) with a 12.00% false positive rate and a 9.20% false alarm rate. Results show that validation protocols fixed splits, cross-validation folds, and pooled datasets produce systematically different metrics even with identical models. Our findings demonstrate why standardized evaluation practices are essential for fair cross-study comparison, and support LightGBM as an effective model for network anomaly detection.
- References
- Downloads
- Published
- 2026-08-22
- Issue
- Vol. 1 No. 4 (2026)
- Section
- Articles
- License
-
Copyright (c) 2026 International Journal of Intelligent Systems and Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.
