EDBT 2026 Demo / reviewers in the wild / expert
Manolis Sifalakis
dblp:75/3932
· DBLP profile ↗
24ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-0949-2094ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 3 first-authorArtificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Partner Project: A Deep Learning Platform Targeting Embedded Hardware for Edge-AI Applications (NEUROKIT2E)abstractThe goal of the NEUROKIT2E project is to create an open-source Deep Learning framework for edge and embedded AI built around an established European value chain. This framework, called AIDGE, supports a wide range of application areas that operate independently and serve a global user community. It provides easy and fast full-stack solutions from Neural Network design and optimization to AI application development all the way down to hardware implementations while enabling code generation for application-specific targets. This platform provides flexibility for academic users in the AI domain to explore and innovate while allowing them the possibility to prototype systems, ensuring their work aligns well with industrial needs. This paper presents the results and achievements of the first part of this three-year project, along with its roadmap and expected outcomes. Rajendra Bishnoi, Mohammad Amin Yaldagard, Said Hamdioui, Kanishkan Vadivel, Manolis Sifalakis, Nicolás Rodríguez 0002, Pedro Julián, Lothar Ratschbacher, Maen Mallah, Yogesh Ramesh Patil, Fabian Chersi |
DATE | 5 |
| 2025 | SENMap: Multi-objective dataflow mapping & synthesis for hybrid scalable neuromorphic systemsabstractThis paper introduces SENMap, a mapping and synthesis tool for a scalable energy efficient neuromorphic computing architecture frameworks. SENECA a flexible architectural design optimized for executing edge AI SNN/ANN inference applications efficiently. To speed up the silicon tapeout and chip design for SENECA, an accurate emulator SENSIM was designed. While SENSIM supports direct mapping of SNNs on neuromorphic architectures, as the SNN/ANN grow in size, achieving optimal mapping for objectives like energy, throughput, area, and accuracy becomes challenging. This paper introduces SENMap, flexible mapping software for efficiently mapping large SNN/ANN applications onto adaptable architectures. SENMap considers architectural, pretrained SNN/ANN realistic examples, and event rate-based parameters and is open-sourced along with SENSIM to aid flexible neuromorphic chip design before fabrication. Experimental results show SENMap enables 40 percent energy improvements for a baseline SENSIM operating on timestep asynchronous mode of operation. SENMap is designed in such a way that it facilitates mapping large spiking neural networks for future modifications as well.1 Prithvish Nembhani, Oliver Rhodes, Guangzhi Tang, Alexandra F. Dobrita, Yingfu Xu, Kanishkan Vadivel, Kevin Shidqi, Paul Detterer, Mario Konijnenburg, Gert-Jan van Schaik, Manolis Sifalakis, Zaid Al-Ars, Amirreza Yousefzadeh |
IJCNN | 11 |
| 2025 | Event-based optical flow on neuromorphic processor: ANN vs. SNN comparison based on activation sparsification
Yingfu Xu, Guangzhi Tang, Amirreza Yousefzadeh, Guido de Croon, Manolis Sifalakis |
Neural Networks | 5 |
| 2024 | Invited: Neuromorphic Vision Modalities in the NimbleAI 3D ChipabstractThis paper provides an overview of the ongoing work to enable novel modalities of passive monocular neuromorphic vision in the NimbleAI sensing-processing architecture; namely, foveated and light-field event-driven vision with selective visual attention. The latter vision modality encodes 3D visual surroundings as sparse visual events in a 4D spatiotemporal domain, adding depth to current representation of visual information delivered by Dynamic Vision Sensors (DVS). The NimbleAI architecture implements hardware support for efficient execution of mainstream computer vision algorithms and AI models using these visual inputs. The architecture is designed to harness the latest advancements in 3D silicon integration, making it possible to squeeze sensing and spiking circuitry, memory, and processing engines into a miniature silicon volume. Xabier Iturbe, Bernabé Linares-Barranco, Sio-Hoi Ieng, Arne Erdmann, Luca Peres, Oliver Rhodes, Rafael Tornero, Manolis Sifalakis, Marcel D. van de Burgwal, Amirreza Yousefzadeh, Maha Kooli, Riccardo Alidori, Pavel Zaykov |
DAC | 8 |
| 2024 | SENSIM: An Event-driven Parallel Simulator for Multi-core Neuromorphic SystemsabstractIn this paper, we present SENSIM, which is an open-source simulator designed specifically for the SENECA neuromorphic processor. This simulator is unique in that it combines features from both hardware-specific and hardware-agnostic spiking neural network simulators, resulting in a hybrid event-driven and time-step-driven simulation approach. This allows for flexibility between accuracy and speed during different stages of simulation. Our work highlights the open-source SENSIM platform, which enables the mapping of large-scale SNN/DNN models to the SENECA cores, as well as the benchmarking of crucial KPIs such as power and latency estimations1. Prithvish Nembhani, Kanishkan Vadivel, Guangzhi Tang, Mohammad Tahghighi, Gert-Jan van Schaik, Manolis Sifalakis, Zaid Al-Ars, Amirreza Yousefzadeh |
IJCNN | 6 |
| 2024 | Co-optimized training of models with synaptic delays for digital neuromorphic acceleratorsabstractConfigurable delays are a basic feature in many neuromorphic neural network hardware accelerators. However, they have been rarely used in model implementations, despite their promising impact on performance and efficiency in tasks that exhibit complex dynamics, as it has been unclear how to optimize them. In this work, we propose a framework to train and deploy in digital neuromorphic hardware highly performing spiking neural networks (SNNs) where apart from the synaptic weights, the delays are also co-optimized. We consider synaptic (i.e. per-synapse) delays and evaluate them in two neuromorphic digital hardware platforms: Intel's Loihi and Imec's Seneca. Leveraging spike-based back-propagation-through-time, the training process accounts for both platform constraints, such as synaptic weight precision and the total number of parameters per core, as a function of the network size. In addition, a delay pruning technique is used to reduce memory footprint with a low cost in performance. The evaluated benchmark involves several models for solving the SHD (Spiking Heidelberg Digits) classification task, where minimal accuracy degradation during the transition from software to hardware is demonstrated. To our knowledge, this is the first work show-casing how to train and deploy hardware-aware models parameterized with synaptic delays, on multicore neuromorphic hardware accelerators. Alberto Patiño-Saucedo, Roy Meijer, Paul Detterer, Amirreza Yousefzadeh, Laura Garrido-Regife, Bernabé Linares-Barranco, Manolis Sifalakis |
ISCAS | 7 |
| 2023 | Empirical study on the efficiency of Spiking Neural Networks with axonal delays, and algorithm-hardware benchmarkingabstractThe role of axonal synaptic delays in the efficacy and performance of artificial neural networks has been largely unexplored. In step-based analog-valued neural network models (ANNs), the concept is almost absent. In their spiking neuroscience-inspired counterparts, there is hardly a systematic account of their effects on model performance in terms of accuracy and number of synaptic operations. This paper proposes a methodology for accounting for axonal delays in the training loop of deep Spiking Neural Networks (SNNs), intending to efficiently solve machine learning tasks on data with rich temporal dependencies. We then conduct an empirical study of the effects of axonal delays on model performance during inference for the Adding task [1]–[3], a benchmark for sequential regression, and for the Spiking Heidelberg Digits dataset (SHD) [4], commonly used for evaluating event-driven models. Quantitative results on the SHD show that SNNs incorporating axonal delays instead of explicit recurrent synapses achieve state-of-the-art, over 90% test accuracy while needing less than half trainable synapses. Additionally, we estimate the required memory in terms of total parameters and energy consumption of accomodating such delay-trained models on a modern neuromorphic accelerator [5], [6]. These estimations are based on the number of synaptic operations and the reference GF-22nm FDX CMOS technology. As a result, we demonstrate that a reduced parameterization, which incorporates axonal delays, leads to approximately 90% energy and memory reduction in digital hardware implementations for a similar performance in the aforementioned task. Alberto Patiño-Saucedo, Amirreza Yousefzadeh, Guangzhi Tang, Federico Corradi, Bernabé Linares-Barranco, Manolis Sifalakis |
ISCAS | 6 |
| 2023 | Open the box of digital neuromorphic processor: Towards effective algorithm-hardware co-designabstractSparse and event-driven spiking neural network (SNN) algorithms are the ideal candidate solution for energy-efficient edge computing. Yet, with the growing complexity of SNN algorithms, it isn't easy to properly benchmark and optimize their computational cost without hardware in the loop. Although digital neuromorphic processors have been widely adopted to benchmark SNN algorithms, their black-box nature is problematic for algorithm-hardware co-optimization. In this work, we open the black box of the digital neuromorphic processor for algorithm designers by presenting the neuron processing instruction set and detailed energy consumption of the SENeCA neuromorphic architecture. For convenient benchmarking and optimization, we provide the energy cost of the essential neuromorphic components in SENeCA, including neuron models and learning rules. Moreover, we exploit the SENeCA's hierarchical memory and exhibit an advantage over existing neuromorphic processors. We show the energy efficiency of SNN algorithms for video processing and online learning, and demonstrate the potential of our work for optimizing algorithm designs. Overall, we present a practical approach to enable algorithm designers to accurately benchmark SNN algorithms and pave the way towards effective algorithm-hardware co-design. Guangzhi Tang, Ali Safa, Kevin Shidqi, Paul Detterer, Stefano Traferro, Mario Konijnenburg, Manolis Sifalakis, Gert-Jan van Schaik, Amirreza Yousefzadeh |
ISCAS | 7 |
| 2022 | Delta Activation Layer exploits temporal sparsity for efficient embedded video processingabstractActivation sparsity can improve compute efficiency and resource utilization in sparsity-aware neural network accelerators. While spatial sparsification of activations is a popular topic in DNN literature, introducing and exploiting spatio-temporal sparsity is a topic much less explored in DNN literature. However, it is in perfect resonance with the trend in DNNs, to shift from static-data signal processing (e.g., image processing) to stream-data real-time signal processing (e.g., video and audio) in embedded edge devices. Towards the goal of exploiting temporal sparsity, in this paper, we introduce a new DNN layer (called Delta Activation Layer), whose sole purpose is to promote both spatial as well as temporal sparsity of activations during training with time-distributed data. One may use the Delta Activation Layer either during vanilla training or during a refinement phase. We have implemented the Delta Activation Layer as an extension of the standard TensorFlow-Keras (2.0) library and applied it on training deep neural networks with several datasets including the Human Action Recognition (UCF101) dataset. We report an almost 3x improvement of activation sparsity, with recoverable loss of model accuracy after prolonged training. All source code for this project is available upon request for academic purposes. Amirreza Yousefzadeh, Manolis Sifalakis |
IJCNN | 2 |
| 2020 | Tera-scale coordinate descent on GPUs
Thomas P. Parnell, Celestine Dünner, Kubilay Atasu, Manolis Sifalakis, Haralampos Pozidis |
Future Gener. Comput. Syst. | 4 |
| 2017 | Linear-complexity relaxed word Mover's distance with GPU accelerationabstractThe amount of unstructured text-based data is growing every day. Querying, clustering, and classifying this big data requires similarity computations across large sets of documents. Whereas low-complexity similarity metrics are available, attention has been shifting towards more complex methods that achieve a higher accuracy. In particular, the Word Mover's Distance (WMD) method proposed by Kusner et al. is a promising new approach, but its time complexity grows cubically with the number of unique words in the documents. The Relaxed Word Mover's Distance (RWMD) method, again proposed by Kusner et al., reduces the time complexity from qubic to quadratic and results in a limited loss in accuracy compared with WMD. Our work contributes a low-complexity implementation of the RWMD that reduces the average time complexity to linear when operating on large sets of documents. Our linear-complexity RWMD implementation, henceforth referred to as LC-RWMD, maps well onto GPUs and can be efficiently distributed across a cluster of GPUs. Our experiments on real-life datasets demonstrate 1) a performance improvement of two orders of magnitude with respect to our GPU-based distributed implementation of the quadratic RWMD, and 2) a performance improvement of three to four orders of magnitude with respect to our distributed WMD implementation that uses GPU-based RWMD for pruning. Kubilay Atasu, Thomas P. Parnell, Celestine Dünner, Manolis Sifalakis, Haralampos Pozidis, Vasileios Vasileiadis, Michail Vlachos, Cesar Berrospi, Abdel Labbi |
IEEE BigData | 4 |
| 2017 | Understanding and optimizing the performance of distributed machine learning applications on apache sparkabstractIn this paper we explore the performance limits of Apache Spark for machine learning applications. We begin by analyzing the characteristics of a state-of-the-art distributed machine learning algorithm implemented in Spark and compare it to an equivalent reference implementation using the high performance computing framework MPI. We identify critical bottlenecks of the Spark framework and carefully study their implications on the performance of the algorithm. In order to improve Spark performance we then propose a number of practical techniques to alleviate some of its overheads. However, optimizing computational efficiency and framework related overheads is not the only key to performance — we demonstrate that in order to get the best performance out of any implementation it is necessary to carefully tune the algorithm to the respective trade-off between computation time and communication latency. The optimal trade-off depends on both the properties of the distributed algorithm as well as infrastructure and framework-related characteristics. Finally, we apply these technical and algorithmic optimizations to three different distributed linear machine learning algorithms that have been implemented in Spark. We present results using five large datasets and demonstrate that by using the proposed optimizations, we can achieve a reduction in the performance difference between Spark and MPI from 20x to 2x. Celestine Dünner, Thomas P. Parnell, Kubilay Atasu, Manolis Sifalakis, Haralampos Pozidis |
IEEE BigData | 4 |
| 2017 | On Hardware Programmable Network Dynamics With a Chemistry-Inspired AbstractionabstractChemical algorithms are statistical control algorithms described and represented as chemical reaction networks. They are analytically tractable, they reinforce a deterministic state-to-dynamics relation, they have configurable stability properties, and they are directly implemented in state space using a high-level visual representation. These properties make them attractive solutions for traffic shaping and generally the control of dynamics in computer networks. In this paper, we present a framework for deploying chemical algorithms on field programmable gate arrays. Besides substantial computational acceleration, we introduce a low-overhead approach for hardware-level programmability and re-configurability of these algorithms at runtime, and without service interruption. We believe that this is a promising approach for expanding the control-plane programmability of software defined networks (SDN), to enable programmable network dynamics. To this end, the simple high-level abstractions of chemical algorithms offer an ideal northbound interface to the hardware, aligned with other programming primitives of SDN (e.g., flow rules). Massimo Monti, Manolis Sifalakis, Christian F. Tschudin, Marco Luise |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | A packet rewriting core for Information Centric NetworkingabstractAlthough very similar, Information Centric Networking implementations like CCNx or NDN have not converged (yet). Moreover, high-level extensions like named functions and low-level considerations for IoT or logical link control will lead to even more variety in formats and experiments, also including slight (and not well-documented) changes of the ICN protocols' behavior. Based on our experience with the multi-protocol implementation CCN-lite, we have developed a domains-specific intermediate language ICCL (Information Centric Core Language). Initially, the purpose of ICCL was to document the common as well as diverging aspects of the various approaches and to provide a run-time which can directly execute an ICCL program. But as it turned out, ICCL became a tool beyond modeling the behavior of ICN forwarders. For example, by including appropriate primitives to access local data sources, one can now seamlessly include in an ICCL program the access to data repositories or serve on-demand data from sensors. This enables easy programmability of network functionality and contributes to “do-it-yourself” networking. We discuss the benefit of this fusion for the development of ICN functionality and report on performance measurements for our prototype system. Christopher Scherb, Manolis Sifalakis, Christian F. Tschudin |
CCNC | 2 |
| 2013 | An Empirical Study of Receiver-Based AIMD Flow-Control Strategies for CCNabstractThe Content Centric Network (CCNx) protocol introduces a new routing and forwarding paradigm for the waist of the future Internet architecture. Below CCN, a volatile set of transport and flow-control strategies is envisioned, which can match better different service requirements, than current TCP/IP technology does. Although a broad range of possibilities has not been explored yet, there are already several proponents of strategies that come close to TCP's well-known flow control mechanism (timeout-driven, window-based, and AIMD- operated). In this paper, we carry out an empirical exploration of some proposed strategies in order to assert their feasibility and efficiency. Our contributions are twofold: First, we establish if receiver-based, timeout-driven, AIMD operated flow-control on Interest transmissions is sufficiently effective for CCN in a future Internet deployment, where it may co-exist with TCP. In this process we compare the performance of three different variants of this strategy, in presence of multi-homed content in the network (one of them proposed by the authors). Second, we provide indicators for the general efficiency of timeout-based flow-control at the CCN receiver, in presence of in-network caching, and exhibit some of the challenges faced by such strategies. Stefan Braun 0002, Massimo Monti, Manolis Sifalakis, Christian F. Tschudin |
ICCCN | 3 |
| 2013 | CCN & TCP co-existence in the future Internet: Should CCN be compatible to TCP?
Stefan Braun 0002, Massimo Monti, Manolis Sifalakis, Christian F. Tschudin |
IM | 3 |
| 2013 | A chemical-inspired approach to design distributed rate controllers for packet networks
Massimo Monti, Manolis Sifalakis, Thomas Meyer 0004, Christian F. Tschudin, Marco Luise |
IM | 2 |
| 2011 | A Generic Service Interface for Cloud Networks
Manolis Sifalakis, Christian F. Tschudin, Sylvain Martin, Lucio Studer Ferreira, Luís M. Correia 0001 |
CLOSER | 1 |
| 2011 | Greedy Routing on Hierarchical ClustersabstractGreedy geometric routing is one of the most appealing options for sensor and ad-hoc networks due to its simplicity, scalability and mainly stateless operation. However, its performance may vary when a low-dimensional embedding results in a graph with void regions and local minima locations. In this paper, we advocate the benefits of clustering for improving the performance of greedy routing over an existing embedding. Our contributions include: (a) an examination of the effects of clustering on influencing routing near void regions and overcoming problems with local minima, (b) a generic framework for enabling a cluster view for greedy routes, that can work with any low-dimensional embedding, and (c) a quantitative evaluation of the performance improvement we achieve, using some reference clustering approaches, over a number of virtual coordinate systems for geometric routing. Ghazi Bouabene, Manolis Sifalakis, Christian F. Tschudin |
EUC | 2 |
| 2011 | Functional composition in future networks
Manolis Sifalakis, Andreas Louca, Ghazi Bouabene, Michael Fry 0001, Andreas Mauthe, David Hutchison 0001 |
Comput. Networks | 1 |
| 2010 | Event detection and correlation for network environmentsabstractAutonomic communication has the aim of supporting fast-evolving network technologies and services, and of reducing the burden in managing complex and dynamic network environments. Networks with desirable self- properties should be more adaptable to changing conditions and would enable greater flexibility and functional scalability. A necessary condition for realising these benefits is a heightened level of network awareness; this requires not merely the capacity to monitor the system and network state, but also the ability to characterise the operational environment and its dynamic shifts. As argued in the literature, patterns of change can be detected through cross-correlation of monitored or sensory inputs expressed in events. The profiling of the temporal ordering and other relationships encoded in these patterns can provide contextual information suitable for reasoning and adaptation tasks. In this paper, we present the design framework and initial evaluation of an Information Sensing system that aims to enable awareness through an integrated event detection-correlation mechanism. In the context of autonomic networks, it offers a more lightweight solution than traditional active database-oriented event systems, and it has better performance than log-post analysis processing; its design also provides for a distributed detection facility. Manolis Sifalakis, Michael Fry 0001, David Hutchison 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2007 | On the Long-Range Dependent Behaviour of Unidirectional Packet Delay of Wireless TrafficabstractIn contrast to aggregate inter-packet metrics that quantify the arrival processes of aggregate traffic at a single point in the network, intraflow end-to-end per-packet performance metrics assess the level of service quality experienced by certain traffic types while routed over a network path. Consequently, while the former is influenced by the superimposition of a large number of concurrent ON/OFF sources, the latter mainly depends on the specific transport mechanisms employed by individual flows in conjunction with the temporal resource contention along the end-to-end path. In this paper, we have used a ubiquitous measurement technique to assess the unidirectional end-to-end delay characteristics experienced by diverse sets of IPv6 flows routed over heterogeneous wireless network configurations. We analysed numerous traces to find that, when viewed as time series data, often exhibit long-range dependence manifested by Hurst parameter estimates greater than 0.5. Our results suggest that the end-to-end packet delay can be bursty across multiple time scales even at the microflow level, implying high performance variability during sufficiently long-lived application sessions. We anticipate that the quantification of such intraflow phenomena can enable applications to optimise and adjust their operation in the face of potential performance degradation. Dimitrios P. Pezaros, Manolis Sifalakis, David Hutchison 0001 |
GLOBECOM | 2 |
| 2006 | A highly flexible service composition framework for real-life networks
Stefan Schmid 0002, Tim Chart, Manolis Sifalakis, Andrew Scott 0001 |
Comput. Networks | 3 |
| 2005 | Network address hopping: a mechanism to enhance data protection for packet communicationsabstractThis paper proposes a novel mechanism to enhance data protection in communications across untrusted networks. The approach is based on the principle of network address hopping, whereby a data stream is spread across multiple end-to-end connections. The aim is to obscure the data exchange between two peers by shuffling the communication pattern. The hopping pattern - a shared secret between the communication participants - defines the sequence for the address hopping and determines the data spreading. Besides a description of the basic operation of the network address hopping mechanism, the paper evaluates the level of protection it can offer for end-to-end communications. This theoretical analysis is accompanied by a quantitative evaluation of the processing overhead of our prototype implementation on commodity end hosts. Manolis Sifalakis, Stefan Schmid 0002, David Hutchison 0001 |
ICC | 1 |