A production-grade, hybrid Intrusion Detection System that combines a high-accuracy supervised machine learning classifier (XGBoost) with an unsupervised deep learning Autoencoder (PyTorch) to detect both known attacks and zero-day anomalies in network traffic — with full SHAP-based explainability for SOC analysts.
- Project Overview
- System Architecture
- Key Results
- Dataset
- Project Structure
- Setup & Installation
- Running the Pipeline
- Phase-by-Phase Breakdown
- Explainability (SHAP)
- Generalization Testing
- Limitations & Future Work
- Technologies Used
Traditional IDS solutions rely entirely on known attack signatures, leaving networks vulnerable to novel (zero-day) exploits. This project presents a Hybrid IDS that solves this gap using a two-stage decision engine:
-
Stage 1 — Known Attack Detection (XGBoost): A gradient-boosted tree classifier trained on 2.2 million labelled network flows from the CICIDS2017 dataset. Achieves 99.89% accuracy on known attack families including DoS, DDoS, PortScan, and Botnet traffic.
-
Stage 2 — Zero-Day Anomaly Detection (Autoencoder): A PyTorch Encoder-Decoder neural network trained exclusively on benign traffic. Any packet that XGBoost passes as "safe" is then reconstructed by the Autoencoder. A high reconstruction error (MSE > 0.0116) flags the packet as an unknown anomaly.
The system is fully explainable via SHAP (SHapley Additive exPlanations), enabling SOC analysts to understand exactly why any packet was flagged as malicious.
Incoming Network Traffic (Flow Features)
│
▼
┌─────────────────────┐
│ Preprocessing │ ← Clean, normalise, encode labels
└─────────┬───────────┘
│
▼
┌─────────────────────┐
│ XGBoost Classifier │ ← Stage 1: Known Attack Detection
└─────────┬───────────┘
│
┌─────────┴───────────┐
│ │
Attack ✓ Benign? (uncertain)
│ │
│ ┌──────────▼──────────┐
│ │ PyTorch Autoencoder│ ← Stage 2: Zero-Day Detection
│ │ (Reconstruction │
│ │ Error Analysis) │
│ └──────────┬──────────┘
│ │
│ ┌──────────┴──────────┐
│ │ │
│ MSE > 0.0116 MSE ≤ 0.0116
│ (Anomaly) ✓ (Safe) ✓
│ │
└──────────▼
ALERT / BLOCK
| Model | Accuracy | Precision | Recall | F1-Score | ROC-AUC |
|---|---|---|---|---|---|
| Random Forest (Baseline) | 99.85% | 99.75% | 99.62% | 99.69% | 99.99% |
| XGBoost (Optimised) | 99.90% | 99.75% | 99.83% | 99.79% | 99.99% |
| Autoencoder (Standalone) | 83.93% | 92.60% | 37.22% | 53.09% | 75.99% |
| Hybrid IDS (Final System) | 99.17% | 96.87% | 99.83% | 98.33% | 99.99% |
Note: The Autoencoder's lower standalone recall is expected — it was designed to detect unknown anomalies, not classify known labelled attack families. Its real strength emerges inside the Hybrid system.
| Attack Type | Samples |
|---|---|
| DoS Hulk | 230,124 |
| PortScan | 158,804 |
| DDoS | 128,027 |
| DoS GoldenEye | 10,293 |
| DoS Slowloris | 5,796 |
| DoS Slowhttptest | 5,499 |
| Bot | 1,956 |
| Infiltration | 36 |
| Heartbleed | 11 |
Our Phase 4 feature importance analysis reduced the feature space from 78 → 39 features while retaining 99% of the model's predictive power, halving inference time with zero accuracy loss.
| Dataset | Purpose | Rows | Features |
|---|---|---|---|
CICIDS2017 (combine.csv) |
Training & evaluation | 2,214,469 | 78 |
| UNSW-NB15 | Generalization testing | 246,996 | 45 |
- CICIDS2017: University of New Brunswick — Place as
data/raw/combine.csv - UNSW-NB15: Kaggle — mrwellsdavid/unsw-nb15 — Place CSV files in
data/raw/
Intelligent-Intrusion-Detection-System-for-Encrypted-Network-Traffic/
│
├── data/
│ ├── raw/
│ │ ├── combine.csv ← CICIDS2017 raw dataset
│ │ └── UNSW_NB15_testing-set.csv ← UNSW-NB15 generalization dataset
│ └── processed/
│ ├── cleaned_data.csv ← After Phase 1 preprocessing
│ └── optimized_data.csv ← After Phase 4 feature selection
│
├── models/
│ ├── random_forest_baseline.pkl ← Trained Random Forest
│ ├── xgboost_baseline.pkl ← Trained XGBoost (primary classifier)
│ ├── autoencoder.pth ← PyTorch Autoencoder weights
│ └── ae_scaler.pkl ← MinMaxScaler for Autoencoder input
│
├── results/
│ ├── baseline_metrics.json ← Phase 3 evaluation metrics
│ ├── ae_threshold.json ← Autoencoder anomaly threshold
│ ├── evaluation_metrics.json ← Phase 7 full comparison metrics
│ ├── generalization_metrics.json ← Phase 8 UNSW-NB15 results
│ ├── class_distribution.png ← Phase 2 EDA
│ ├── correlation_heatmap.png ← Phase 2 EDA
│ ├── top_features_histograms.png ← Phase 2 EDA
│ ├── feature_importance.png ← Phase 4 XGBoost importances
│ ├── cumulative_importance.png ← Phase 4 feature selection curve
│ ├── reconstruction_error.png ← Phase 5 Autoencoder MSE dist.
│ ├── hybrid_confusion_matrix.png ← Phase 6 Hybrid IDS matrix
│ ├── model_comparison.png ← Phase 7 ROC-AUC + metrics bar
│ ├── generalization_confusion_matrix.png← Phase 8 cross-dataset matrix
│ ├── shap_summary_bar.png ← Phase 9 global importance
│ ├── shap_beeswarm.png ← Phase 9 feature direction
│ └── shap_waterfall.png ← Phase 9 single prediction
│
├── scripts/
│ ├── preprocess.py ← Phase 1: Data cleaning & encoding
│ ├── eda.py ← Phase 2: Exploratory analysis
│ ├── train_baseline.py ← Phase 3: RF + XGBoost training
│ ├── feature_optimization.py ← Phase 4: Feature selection
│ ├── train_autoencoder.py ← Phase 5: PyTorch Autoencoder
│ ├── hybrid_ids.py ← Phase 6: Hybrid decision engine
│ ├── evaluate.py ← Phase 7: Full model comparison
│ ├── generalization_test.py ← Phase 8: Cross-dataset testing
│ └── shap_explain.py ← Phase 9: SHAP explainability
│
├── requirements.txt
└── README.md
- Python 3.10 or higher
- pip
git clone https://github.com/yourusername/Intelligent-Intrusion-Detection-System-for-Encrypted-Network-Traffic.git
cd Intelligent-Intrusion-Detection-System-for-Encrypted-Network-Trafficpip install -r requirements.txt
pip install torchNote: TensorFlow 2.16.1 (listed in requirements.txt) is not compatible with Python 3.14+. This project uses PyTorch as an equivalent alternative for the Autoencoder.
- CICIDS2017 → Place as
data/raw/combine.csv - UNSW-NB15 (optional, for Phase 8) → Place CSV files in
data/raw/
Run each phase sequentially from the project root directory:
# Phase 1 — Preprocess the raw dataset
python scripts/preprocess.py
# Phase 2 — Exploratory Data Analysis
python scripts/eda.py
# Phase 3 — Train baseline ML models (Random Forest + XGBoost)
python scripts/train_baseline.py
# Phase 4 — Feature optimisation (select top 39 features)
python scripts/feature_optimization.py
# Phase 5 — Train the Autoencoder on normal traffic only
python scripts/train_autoencoder.py
# Phase 6 — Run the Hybrid IDS decision engine
python scripts/hybrid_ids.py
# Phase 7 — Formal evaluation: compare all models with ROC-AUC curves
python scripts/evaluate.py
# Phase 8 — Generalization test on UNSW-NB15 (requires dataset in data/raw/)
python scripts/generalization_test.py
# Phase 9 — SHAP explainability plots
python scripts/shap_explain.pyAll output plots are saved to results/ and models are saved to models/.
| Phase | Script | Description | Output |
|---|---|---|---|
| 0 | — | Environment setup & project structure | data/, models/, results/, scripts/ |
| 1 | preprocess.py |
Clean data, drop NaN/Inf, binary-encode labels | cleaned_data.csv |
| 2 | eda.py |
Class distribution, correlation heatmap, histograms | 3 PNG plots |
| 3 | train_baseline.py |
Train Random Forest (50 trees) + XGBoost (100 estimators) | 2 .pkl model files |
| 4 | feature_optimization.py |
Extract importances, reduce 78 → 39 features (99% power retained) | optimized_data.csv, 2 plots |
| 5 | train_autoencoder.py |
Train PyTorch Encoder-Decoder on benign traffic only; calculate MSE threshold | autoencoder.pth, ae_threshold.json |
| 6 | hybrid_ids.py |
Two-stage decision: XGBoost → Autoencoder for benign packets | Confusion matrix plot |
| 7 | evaluate.py |
Head-to-head comparison + ROC-AUC curves for all 4 models | evaluation_metrics.json, comparison plot |
| 8 | generalization_test.py |
Cross-dataset test on UNSW-NB15 with semantic feature mapping | generalization_metrics.json, matrix plot |
| 9 | shap_explain.py |
SHAP TreeExplainer: Summary Bar, Beeswarm, Waterfall plots | 3 SHAP PNG plots |
Model explainability is critical for real-world SOC deployment. Using SHAP (SHapley Additive exPlanations) with TreeExplainer on the XGBoost model, the system can:
- Globally rank which network features most commonly drive attack predictions.
- Locally explain why a specific packet was flagged — step-by-step, feature by feature.
Top contributing features identified:
Average Packet SizeBwd Packet Length StdBwd Header LengthMax Packet LengthFwd IAT MaxDestination PortInit_Win_bytes_forwardFlow IAT Mean
Testing the Hybrid IDS on the UNSW-NB15 dataset revealed significant performance degradation, which is an important and honest finding:
| Dataset | Accuracy | F1-Score |
|---|---|---|
| CICIDS2017 (trained on) | 99.17% | 98.33% |
| UNSW-NB15 (unseen) | 42.95% | 0.06% |
Root Cause — Dataset Shift:
- Only ~24 of 78 required features could be semantically mapped from UNSW-NB15, with remaining features zeroed out.
- The attack families (Fuzzer, Exploit, Reconnaissance, Backdoor) are statistically distinct from CICIDS2017 attack families.
- Both datasets were generated in different lab environments.
Implications for Future Work: This result motivates several important research directions including transfer learning, domain adaptation, and online/continual learning for real-world IDS deployments.
| Limitation | Proposed Future Work |
|---|---|
| Model trained on static lab data | Implement online learning for continuous adaptation |
| Low cross-dataset generalization | Explore transfer learning and domain adaptation |
| Autoencoder threshold is fixed | Use reinforcement learning for dynamic threshold adjustment |
| Feature-based (requires CICFlowMeter) | Integrate with live packet capture for real-time inference |
| No multi-class classification | Extend to multi-class IDS to identify specific attack families |
| Library | Version | Purpose |
|---|---|---|
pandas |
2.2.1 | Data manipulation |
numpy |
1.26.4 | Numerical computing |
scikit-learn |
1.4.1 | ML models, preprocessing, metrics |
xgboost |
2.0.3 | Gradient-boosted classifier |
torch (PyTorch) |
2.12.1 | Autoencoder deep learning |
shap |
0.51.0 | Model explainability |
matplotlib |
3.8.3 | Visualizations |
seaborn |
0.13.2 | Statistical plots |
joblib |
1.3.2 | Model serialization |
imbalanced-learn |
0.12.0 | Class imbalance handling |
- CICIDS2017 Dataset: Sharafaldin, I., Lashkari, A. H., & Ghorbani, A. A. (2018). Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization. ICISSP. Link
- UNSW-NB15 Dataset: Moustafa, N., & Slay, J. (2015). UNSW-NB15: a comprehensive data set for network intrusion detection systems. MilCIS. Link
- XGBoost: Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. KDD. Link
- SHAP: Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS. Link
- PyTorch: Paszke, A., et al. (2019). PyTorch: An Imperative Style, High-Performance Deep Learning Library. NeurIPS. Link
