This paper introduces a novel hybrid architecture that integrates Kolmogorov-Arnold Networks (KANs) with traditional convolutional neural networks for visual recognition tasks in edge computing environments. KANs leverage the Kolmogorov-Arnold representation theorem to model multivariate continuous functions through compositions of univariate functions, offering potential advantages in parameter efficiency and representational capacity. Our approach combines CNN-based feature extraction with KAN-based classification to exploit the complementary strengths of both paradigms. Through extensive experiments on the Visual Wake Words dataset, we demonstrate that our hybrid architecture achieves 82.3% accuracy while maintaining moderate parameter usage (78.5K parameters) and reasonable inference latency. Unlike conventional approaches that focus on extremely low-resolution inputs, our model processes 128×128-pixel images, preserving more visual details without compromising computational efficiency. Comparative analysis reveals that our approach outperforms several specialized lightweight architectures by 4.7-5.5 percentage points in accuracy while requiring fewer computational resources than larger models with similar performance. Additionally, we provide insights into optimizing inference through batch processing, achieving a 26× speedup when using batch size 32. This work expands the design space for efficient neural architectures beyond traditional CNNs and demonstrates that KAN-based models represent a promising direction for resource-aware visual computing at the edge.

Advancing Visual Recognition with Kolmogorov-Arnold Networks: A Novel Hybrid Architecture for Edge Computing Applications

Kuznetsov O.
;
Randieri C.
2025-01-01

Abstract

This paper introduces a novel hybrid architecture that integrates Kolmogorov-Arnold Networks (KANs) with traditional convolutional neural networks for visual recognition tasks in edge computing environments. KANs leverage the Kolmogorov-Arnold representation theorem to model multivariate continuous functions through compositions of univariate functions, offering potential advantages in parameter efficiency and representational capacity. Our approach combines CNN-based feature extraction with KAN-based classification to exploit the complementary strengths of both paradigms. Through extensive experiments on the Visual Wake Words dataset, we demonstrate that our hybrid architecture achieves 82.3% accuracy while maintaining moderate parameter usage (78.5K parameters) and reasonable inference latency. Unlike conventional approaches that focus on extremely low-resolution inputs, our model processes 128×128-pixel images, preserving more visual details without compromising computational efficiency. Comparative analysis reveals that our approach outperforms several specialized lightweight architectures by 4.7-5.5 percentage points in accuracy while requiring fewer computational resources than larger models with similar performance. Additionally, we provide insights into optimizing inference through batch processing, achieving a 26× speedup when using batch size 32. This work expands the design space for efficient neural architectures beyond traditional CNNs and demonstrates that KAN-based models represent a promising direction for resource-aware visual computing at the edge.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11389/93123
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact