Intrusion detection in IoT and IIoT networks must operate under tight resource constraints, yet most published machine learning-based IDS solutions report accuracy on held-out data without addressing whether the trained model can actually run on the target hardware. We address this gap with an end-to-end study spanning dataset preprocessing, model training, INT8 quantisation, and physical execution on two real microcontrollers. Five supervised classifiers—Logistic Regression, Decision Tree (depth 5), Random Forest, XGBoost, and LightGBM—plus an MLP deep learning baseline are evaluated on binary and ten-class intrusion detection tasks using the TON_IoT network dataset. A 5-fold stratified cross-validation confirms stable performance across splits, with LightGBM reaching (Formula presented.). Models are then exported through three quantisation pipelines: m2cgen C code generation for the two lightest classifiers, TensorFlow Lite Micro full-integer INT8 for the MLP (9.34× size reduction to 13.03 KB), and a custom post-training INT8 binary format for XGBoost and LightGBM (18.91× compression for LightGBM to 73.85 KB). All five quantised models are deployed to an Arduino Mega 2560 (ATmega2560, 16 MHz, 8 KB SRAM) and an ESP32-C3 SuperMini (RISC-V, 160 MHz, 400 KB SRAM) and benchmarked on physical hardware across 500 timed inferences per model (250 per input class), with firmware predictions confirmed to match the Python 3.11 float model on both test vectors. The Decision Tree achieves 5.6 µs inference on the ESP32-C3; LightGBM INT8 ((Formula presented.)) provides the best accuracy–size trade-off among ensemble models. Cross-platform comparison reveals that the RISC-V device is 5.8–7.8× faster than the 8-bit AVR for identical model code. A cross-domain evaluation using CIC-IoT-Dataset2023 identifies large normalised distribution shifts (up to (Formula presented.) in packet asymmetry), quantifying the generalisation gap that remains an open challenge.

Lightweight Machine Learning Intrusion Detection for IoT/IIoT Networks: Quantisation Strategies and Physical Deployment on Resource-Constrained Microcontrollers

Kuznetsov O.
;
Arnesano M.;
2026-01-01

Abstract

Intrusion detection in IoT and IIoT networks must operate under tight resource constraints, yet most published machine learning-based IDS solutions report accuracy on held-out data without addressing whether the trained model can actually run on the target hardware. We address this gap with an end-to-end study spanning dataset preprocessing, model training, INT8 quantisation, and physical execution on two real microcontrollers. Five supervised classifiers—Logistic Regression, Decision Tree (depth 5), Random Forest, XGBoost, and LightGBM—plus an MLP deep learning baseline are evaluated on binary and ten-class intrusion detection tasks using the TON_IoT network dataset. A 5-fold stratified cross-validation confirms stable performance across splits, with LightGBM reaching (Formula presented.). Models are then exported through three quantisation pipelines: m2cgen C code generation for the two lightest classifiers, TensorFlow Lite Micro full-integer INT8 for the MLP (9.34× size reduction to 13.03 KB), and a custom post-training INT8 binary format for XGBoost and LightGBM (18.91× compression for LightGBM to 73.85 KB). All five quantised models are deployed to an Arduino Mega 2560 (ATmega2560, 16 MHz, 8 KB SRAM) and an ESP32-C3 SuperMini (RISC-V, 160 MHz, 400 KB SRAM) and benchmarked on physical hardware across 500 timed inferences per model (250 per input class), with firmware predictions confirmed to match the Python 3.11 float model on both test vectors. The Decision Tree achieves 5.6 µs inference on the ESP32-C3; LightGBM INT8 ((Formula presented.)) provides the best accuracy–size trade-off among ensemble models. Cross-platform comparison reveals that the RISC-V device is 5.8–7.8× faster than the 8-bit AVR for identical model code. A cross-domain evaluation using CIC-IoT-Dataset2023 identifies large normalised distribution shifts (up to (Formula presented.) in packet asymmetry), quantifying the generalisation gap that remains an open challenge.
2026
Inglese
15
13
Arduino; edge AI; embedded deployment; ESP32; IIoT; intrusion detection; IoT security; LightGBM; machine learning; model quantisation; SHAP; TensorFlow Lite Micro; XGBoost
5
info:eu-repo/semantics/article
262
De Bernardis, E. P.; Kuznetsov, O.; Arnesano, M.; Zhansaya, P.; Sydykova, M.
1 Contributo su Rivista::1.1 Articolo in rivista
none
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11389/93091
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact