EDBT 2026 Demo / reviewers in the wild / expert
Orlando Moreira
dblp:33/2059
· DBLP profile ↗
25ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-2362-4169ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksabstractDeep Neural Networks (DNNs) are accurate but compute-intensive, leading to substantial energy consumption during inference. Exploiting temporal redundancy through ∆-Σ convolution [26] in video processing has proven to greatly enhance computation efficiency. However, temporal ∆-Σ DNNs typically require substantial memory for storing neuron states to compute inter-frame differences, hindering their on-chip deployment. To mitigate this memory cost, directly compressing the states can disrupt the linearity of temporal ∆-Σ convolution, causing accumulated errors in long-term ∆-Σ processing. Thus, we propose MEET, an optimization framework for MEmory-Efficient Temporal ∆-Σ DNNs. MEET transfers the state compression challenge to a well-established weight compression problem by trading fewer activations for more weights and introduces a co-design of network architecture and suppression method to optimize for mixed spatial-temporal execution. Evaluations on three vision applications demonstrate a reduction of 5.1∼13.3 × in total memory compared to the most computation-efficient temporal DNNs, while preserving the computation efficiency and model accuracy in long-term ∆-Σ processing. MEET facilitates the deployment of temporal ∆-Σ DNNs within on-chip memory of embedded event-driven platforms, empowering low-power edge processing. Zeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira |
CVPR | 6 |
| 2025 | Neuromorphic Edge Computing: Challenges, Opportunities, and Current SolutionsabstractNeuromorphic computing is emerging as a paradigm for high-performance, energy-efficient edge intelligence. Yet the transition from laboratory prototypes to deployable edge platforms is slowed by four intertwined obstacles: (1) complex near-sensor integration, where spiking inference must co-locate with analogue sensing to minimise latency and data-movement energy; (2) novel event-based optimisation, requiring weight compression and sparsity techniques tailored to event-driven workloads; (3) heterogeneous integration of emerging devices, such as RRAM and other non-volatile memories, into reliable, manufacturable stacks; and (4) novel security risks, including spike-pattern side channels and model-specific attacks that demand to develop neuromorphic security primitives. This paper surveys the state of the art across these four fronts, drawing on recent advances in spiking microcontrollers, mixed-precision compute-in-memory fabrics, sparsity-aware compilation, and hardware-anchored security primitives (physical unclonable functions, true random number generators, computing-in-memory-based cryptography). By distilling lessons from academic research and industrial prototyping, the paper outlines current solutions and future research directions aimed at accelerating the adoption of neuromorphic platforms in real-world edge AI systems. Federico Corradi, Amir Zjajo, Letícia Maria Veiras Bolzani, Milos Krstic, Orlando Moreira, Zeqi Zhu, Farhad Merchant |
ISLPED | 5 |
| 2024 | ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration
Zeqi Zhu, Alberto García Ortiz, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira |
ECCV (11) | 6 |
| 2024 | CATS: Combined Activation and Temporal Suppression for Efficient Network InferenceabstractBrain-inspired event-driven processors execute deep neural networks (DNNs) in a sparsity-aware manner, leading to superior performance compared to conventional platforms. In the pursuit of higher event sparsity, prior studies suppress non-zero events by either eliminating the intra-frame activations (spatially) or leveraging the redundancy in the inter-frame differences for a video (temporally). However, we have empirically observed that simultaneously enhancing activation and temporal sparsity can lead to a synergistic suppression outcome. To this end, we propose an end-to-end event suppression training approach CATS −− Combined Activation and Temporal Suppression for efficient network inference. It utilizes a gradient-based method to search for the optimal temporal thresholds per layer while penalizing the presence of events in both spatial and temporal domains. Our experimental results show that CATS achieves 2 ∼ 6× higher event suppression compared to the inherent ReLU suppression across a wide range of vision applications, consistently outperforming the state-of-the-art (SOTA) methods by a significant margin at all accuracy levels. Furthermore, a case study on the commercial event-driven processor GrAI-VIP highlights that the induced event sparsity in SSD on the EgoHands dataset can be efficiently translated into a performance enhancement of 2.5× in FPS, 2.1× in latency, and 3.8× in energy consumption, while maintaining the model accuracy. Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Ibrahim Batuhan Akkaya, Egor Bondarev, Orlando Moreira |
WACV | 6 |
| 2023 | NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated ChipsabstractThe NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2. Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov |
DATE | 32 |
| 2023 | Synapse Compression for Event-Based Convolutional-Neural-Network AcceleratorsabstractManufacturing-viable neuromorphic chips require novel compute architectures to achieve the massively parallel and efficient information processing the brain supports so effortlessly. The most promising architectures for that are spiking/event-based, which enables massive parallelism at low complexity. However, the large memory requirements for synaptic connectivity are a showstopper for the execution of modern convolutional neural networks (CNNs) on massively parallel, event-based architectures. The present work overcomes this roadblock by contributing a lightweight hardware scheme to compress the synaptic memory requirements by several thousand times—enabling the execution of complex CNNs on a single chip of small form factor. A silicon implementation in a 12-nm technology shows that the technique achieves a total memory-footprint reduction of up to 374× compared to the best previously published technique at a negligible area overhead. Lennart Bamberg, Arash Pourtaherian, Luc Waeijen, Anupam Chahar, Orlando Moreira |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | ARTS: An adaptive regularization training schedule for activation sparsity explorationabstractBrain-inspired event-based processors have attracted considerable attention for edge deployment because of their ability to efficiently process Convolutional Neural Networks (CNNs) by exploiting sparsity. On such processors, one critical feature is that the speed and energy consumption of CNN inference are approximately proportional to the number of non-zero values in the activation maps. Thus, to achieve top performance, an efficient training algorithm is required to largely suppress the activations in CNNs. We propose a novel training method, called Adaptive-Regularization Training Schedule (ARTS), which dramatically decreases the non-zero activations in a model by adaptively altering the regularization coefficient through training. We evaluate our method across an extensive range of computer vision applications, including image classification, object recognition, depth estimation, and semantic segmentation. The results show that our technique can achieve 1.41 × to 6.00 × more activation suppression on top of ReLU activation across various networks and applications, and outperforms the state-of-the-art methods in terms of training time, activation suppression gains, and accuracy. A case study for a commercially-available event-based processor, Neuronflow, shows that the activation suppression achieved by ARTS effectively reduces CNN inference latency by up to 8.4 × and energy consumption by up to 14.1 ×. Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Lennart Bamberg, Egor Bondarev, Orlando Moreira |
DSD | 6 |
| 2020 | NeuronFlow: a neuromorphic processor architecture for Live AI applicationsabstractNeuronflow is a neuromorphic, many core, data flow architecture that exploits brain-inspired concepts to deliver a scalable event-based processing engine for neuron networks in Live AI applications. Its design is inspired by brain biology, but not necessarily biologically plausible. The main design goal is the exploitation of sparsity to dramatically reduce latency and power consumption as required by sensor processing at the Edge. Orlando Moreira, Amirreza Yousefzadeh, Fabian Chersi, Gokturk Cinserin, Rik-Jan Zwartenkot, Ajay Kapoor, Peng Qiao, Peter Kievits, Mina A. Khoei, Louis Rouillard, Aimee Ferouge, Jonathan Tapson, Ashoka Visweswara |
DATE | 1 |
| 2018 | Evolutionary multi-level acyclic graph partitioningabstractDirected graphs are widely used to model data flow and execution dependencies in streaming applications. This enables the utilization of graph partitioning algorithms for the problem of parallelizing execution on multiprocessor architectures under hardware resource constraints. However due to program memory restrictions in embedded multiprocessor systems, applications need to be divided into parts without cyclic dependencies. This can be done by a subsequent second graph partitioning step with an additional acyclicity constraint. Orlando Moreira, Merten Popp, Christian Schulz 0003 |
GECCO | 1 |
| 2017 | Automatic Control Flow Generation for OpenVX GraphsabstractHeterogeneous platforms with large numbers of processing elements (PEs) have been proposed to satisfy the computational requirements of computer vision applications. Limiting the incurred communication cost here is key to meet the power constraints of embedded devices.We present a new heuristic to reduce communication among PEs and to external memory by aggregating inter-process communication and pipelining image processing functions. The application is specified as an OpenVX graph, an industry standard for vision applications, though our method is applicable to dataflow in general. We use dataflow and graph-based analysis techniques to map the application at configuration time to a hardware platform with strong program memory constraints. We show that our approach can yield a reduction of up to 53% in communication compared to other OpenVX implementations. Merten Popp, Stef van Son, Orlando Moreira |
DSD | 3 |
| 2017 | Graph Partitioning with Acyclicity ConstraintsabstractGraphs are widely used to model execution dependencies in applications. In particular, the NP-complete problem of partitioning a graph under constraints receives enormous attention by researchers because of its applicability in multiprocessor scheduling. We identified the additional constraint of acyclic dependencies between blocks when mapping streaming applications to a heterogeneous embedded multiprocessor. Existing algorithms and heuristics do not address this requirement and deliver results that are not applicable for our use-case. In this work, we show that this more constrained version of the graph partitioning problem is NP-complete and present heuristics that achieve a close approximation of the optimal solution found by an exhaustive search for small problem instances and much better scalability for larger instances. In addition, we can show a positive impact on the schedule of a real imaging application that improves communication volume and execution time. Orlando Moreira, Merten Popp, Christian Schulz 0003 |
SEA | 1 |
| 2017 | Special Section: Integrating Dataflow, Embedded Computing and ArchitectureabstractNo abstract available. Twan Basten, Orlando Moreira, Robert de Groote |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Automatic HAL generation for embedded multiprocessor systemsabstractAutomated hardware design flows considerably speed up the development of embedded systems and are a useful asset during architecture exploration phase. However, any existing software has to be adapted for every new system. In this work we will demonstrate, how a Hardware Abstraction Layer (HAL) for device addresses and properties can be automatically generated from a formal system description while providing sufficient abstraction from hardware details. Comparison to earlier projects show that this saves between 40--50 person-weeks of work per IP. Necessary device specific information is stored using a novel approach that allows the compiler to remove unused data. We will show with a real imaging application that this can reduce the amount of memory used by the HAL from ≈ 9.5% to ≈ 5.1% of the available scratchpad memory. It will also be shown that the overhead of this HAL only depends on the level of abstraction that is used by the application and that performance and memory usage will equal a hard-coded solution in the case that an application uses compile-time constant device identifiers. Merten Popp, Orlando Moreira, Wim Yedema, Menno Lindwer |
EMSOFT | 2 |
| 2016 | Response modeling runtime schedulers for timing analysis of self-timed dataflow graphs
Alok Lele, Orlando Moreira, Pieter J. L. Cuijpers, Kees van Berkel 0001 |
J. Syst. Archit. | 2 |
| 2016 | Buffer allocation for real-time streaming applications running on heterogeneous multi-processors without back-pressure
Hrishikesh Salunkhe, Alok Lele, Orlando Moreira, Kees van Berkel 0001 |
J. Syst. Archit. | 3 |
| 2015 | FP-scheduling for mode-controlled dataflow: a case study
Alok Lele, Orlando Moreira, Kees van Berkel 0001 |
DATE | 2 |
| 2015 | Buffer Allocation for Dynamic Real-Time Streaming Applications Running on a Multi-processor without Back-PressureabstractBuffer allocation for real-time streaming applications, modeled as dataflow graphs, minimizes the total memory consumption while reserving sufficient space for each data production without overwriting any live data and guaranteeing the satisfaction of real-time constraints. We focus on the problem of buffer allocation for systems without back-pressure. Since systems without back-pressure lack blocking behavior at the side of the producer, buffer allocation requires both best- and worst-case timing analysis. Moreover, the dynamic (data-dependent) behavior in these applications makes buffer allocation challenging from the best- and worst-case- timing analysis perspective. We argue that static dataflow cannot conveniently express the dynamic behavior of these applications, leading to overallocation of memory resources. Mode-controlled Dataflow (MCDF) is a restricted form of dynamic dataflow that allows mode switching at runtime and static analysis of real-time constraints. In this paper, we address the problem of buffer allocation for MCDF graphs scheduled on systems without back-pressure. We consider practically relevant applications that can be modeled in MCDF using recurrent-choice mode sequence that consists of the mode sequences of equal length; it provides tractable analysis. Our contribution is a buffer allocation algorithm that achieves up to 36% reduction in total memory consumption compared to the current state-of-the-art for an LTE and an LTE Advanced receiver use cases. Hrishikesh Salunkhe, Alok Lele, Orlando Moreira, Kees van Berkel 0001 |
DSD | 3 |
| 2014 | Mode-Controlled Dataflow based modeling & analysis of a 4G-LTE receiverabstractToday's smartphones and tablets contain multiple cellular modems to support 2G/3G/4G standards, including Long Term Evolution (LTE). They run on complex multi-processor hardware platforms and have to meet hard real-time constraints. Dataflow modeling can be used to design an LTE receiver. Static dataflow allows a rich set of analysis techniques, but is too restrictive to model the dynamic behavior in many realistic applications, including LTE receivers. Dynamic dataflow allows modeling of many realistic applications, but does not support rigorous temporal analysis. Mode-Controlled Dataflow (MCDF) is a restricted form of dynamic dataflow, and allows the same analysis techniques as static dataflow, in principle. We prove that MCDF is sufficiently expressive to handle the dynamic behavior of a realistic LTE receiver, by systematically and stepwise developing a complete MCDF model for an LTE receiver. Hrishikesh Salunkhe, Orlando Moreira, Kees van Berkel 0001 |
DATE | 2 |
| 2012 | A new data flow analysis model for TDMabstractThis paper proposes a new data ow model for analyzing the worst-case temporal behavior of resource arbitration through Time Division Multiplexing (TDM). Alok Lele, Orlando Moreira, Pieter J. L. Cuijpers |
EMSOFT | 2 |
| 2012 | Hard-Real-Time Scheduling on a Weakly Programmable Multi-core Processor with Application to Multi-standard Channel DecodingabstractIn Software Defined Radio (SDR), some or all of the physical layer functions are implemented by software. In this paper, we focus on the channel decoding part of SDR. We use Synchronous Data Flow (SDF) and Cyclo-Static Data Flow (CSDF) graphs to model channel decoding functions. We want to tackle the problem of scheduling a dynamic mix of multiple radios with throughput constraints on a multi-standard multi-channel channel decoder. The decoder consists of a Micro-Controller Unit (MCU) and several weakly programmable Hardware Units (HU) with internal states and very limited buffer sizes. Each HU has a Round Robin (RR) scheduler hosted on the MCU. To reduce scheduling overhead, RR schedules applications at coarse granularity. Due to limited buffer sizes, some tasks of an application are tightly coupled. We propose a so-called coupled scheduling policy, which is a relaxation of strict gang scheduling, to concurrently schedule these tasks. We propose a technique to model coupled scheduling in (C)SDF graphs. Under our scheduling policies, we also design an admission controller to guarantee the throughput requirements of running applications. To verify the approach, we have implemented a simulation system to run DVB-SH and DVB-T concurrently and independently. Orlando Moreira, Rick J. M. Nas, Kees van Berkel 0001 |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2011 | Power Minimisation for Real-Time Dataflow ApplicationsabstractEnergy efficient execution of applications is important for many reasons, e.g. time between battery charges, device temperature. Voltage and Frequency Scaling (VFS) enables applications to be run at lower frequencies on hardware resources thereby consuming less power. Real-time applications have deadlines that must be met otherwise their output is devalued. Dataflow modelling of real-time applications enables off-line verification of the application's temporal requirements. In this paper we describe a method to reduce the combined static and dynamic energy consumption using a Dynamic VFS (DVFS) technique for dataflow modelled real-time applications that may be mapped onto multiple hardware resources. We achieve this by using an application's static slack in order to perform DVFS while still satisfying the application's temporal requirements. We show that by formulating a dataflow modelled application and its mapping as a convex optimisation problem, with energy consumption as the objective function, the problem can be solved with a generic convex optimisation solver, producing an energy optimal constant frequency per application task. Our method allows task frequencies to be constrained such that, e.g. one frequency per application or per processor may be achieved. Andrew Nelson 0001, Orlando Moreira, Anca Mariana Molnos, Sander Stuijk, Ba Thang Nguyen, Kees Goossens |
DSD | 2 |
| 2010 | Buffer Sizing for Rate-Optimal Single-Rate Data-Flow Scheduling RevisitedabstractSingle-Rate Data-Flow (SRDF) graphs, also known as Homogeneous Synchronous Data-Flow (HSDF) graphs or Marked Graphs, are often used to model the implementation and do temporal analysis of concurrent DSP and multimedia applications. An important problem in implementing applications expressed as SRDF graphs is the computation of the minimal amount of buffering needed to implement a static periodic schedule (SPS) that is optimal in terms of execution rate, or throughput. Ning and Gao [1] propose a linear-programming-based polynomial algorithm to compute this minimal storage amount, claiming optimality. We show via a counterexample that the proposed algorithm is not optimal. We prove that the problem is, in fact, NP-complete. We give an exact solution, and experimentally evaluate the degree of inaccuracy of the algorithm of Ning and Gao. Orlando Moreira, Twan Basten, Marc Geilen, Sander Stuijk |
IEEE Trans. Computers | 1 |
| 2007 | Scheduling multiple independent hard-real-time jobs on a heterogeneous multiprocessorabstractThis paper proposes a scheduling strategy and an automatic scheduling flow that enable the simultaneous execution of multiple hard-real-time dataflow jobs. Each job has its own execution rate and starts and stops independently from other jobs, at instants unknown at compile-time, on a multiprocessor system-on-chip. We show how a combination of Time-Division Multiplex (TDM) and static-order scheduling can be modeled as additional nodes and edges on top of the dataflow representation of the job using Single-Rate Dataflow semantics to enable tight worst-case temporal analysis. We also propose algorithms to find combined TDM/static order schedules for jobs that guarantee a requested minimum throughput and maximum latency, while minimizing the usage of processing resources. We illustrate the usage of these techniques for a combination of Wireless LAN and TD-SCDMA radio jobs running on a prototype Software-Defined Radio platform. Orlando Moreira, Frederico Valente, Marco Bekooij |
EMSOFT | 1 |
| 2005 | Multiprocessor Resource Allocation for Hard-Real-Time Streaming with a Dynamic Job-MixabstractAn embedded multiprocessor that can run multiple hard-real-time (HRT) jobs simultaneously has to guarantee that enough resources are available to meet the timing constraints. It is essential that both application model and hardware be tailored to this goal. Moreover, suitable resource allocation and scheduling are needed. This paper proposes a resource allocator that gives guarantees for HRT streaming applications. Because new jobs arrive during operation, resource allocation is performed at run-time. This provides admission control. Resource budget enforcement is handled by local schedulers. We formalize our resource allocation problem and show that it is NP-complete. We developed heuristics to tackle the problem during runtime and evaluated them. A modified First-fit Vector Bin-Packing algorithm provides a good solution; it can allocate 95% of the resources, while handling a large number of job arrivals and departures on a heavily loaded system. Orlando Moreira, Jan David Mol, Marco Bekooij, Jef L. van Meerbergen |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2004 | Predictable Embedded Multiprocessor System Design
Marco Bekooij, Orlando Moreira, Peter Poplavko, Bart Mesman, Milan Pastrnak, Jef L. van Meerbergen |
SCOPES | 2 |