Manil Dev Gomony

dblp:20/11157 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-5889-0785ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 9 first-author · 16 since 2021Software engineering, systems software and programming languages · 8 · 5 first-author · 3 since 2021Computer networks · 4 · 3 since 2021
YearPublicationVenuePosition
2026 LOKI: a 0.266 pJ/SOP Digital SNN Accelerator with Multi-Cycle Clock-Gated SRAM in 22 nm
abstract
Bio-inspired sensors like Dynamic Vision Sensors (DVS) and silicon cochleas are often combined with Spiking Neural Networks (SNNs), enabling efficient, event-driven processing similar to biological sensory systems. To realize the low-power constraints of the edge, the SNN should run on a hardware architecture that can exploit the sparse nature of the spikes. In this paper, we introduce LOKI, a digital architecture for Fully-Connected (FC) SNNs. By using Multi-Cycle Clock-Gated (MCCG) SRAMs, LOKI can operate at 0.59 V, while running at a clock frequency of 667 MHz. At full throughput, LOKI only consumes $0.266 \mathrm{pJ} /$ SOP. We evaluate LOKI on both the Neuromorphic MNIST (N-MNIST) and the Keyword Spotting (KWS) tasks, achieving 98.0 % accuracy at 119.8 nJ /inference and $93.0 \%$ accuracy at 546.5 nJ /inference respectively.
Rick Luiken, Lorenzo Pes, Manil Dev Gomony, Sander Stuijk
ASP-DAC3
2026 Voltage Aware Approximate CGRA Synthesis for Energy Efficient DNN Inference
Georgios Alexandris, Panagiotis Chaidos, Alexis Maras, Barry de Bruin, Manil Dev Gomony, Henk Corporaal, Dimitrios Soudris, Sotirios Xydis
DATE5
2026 CRAFT: A Co-Adaptive Freeze and Train Strategy for Dynamic Neural Networks
Priscilla Sharon Allwin, Manil Dev Gomony, Marc Geilen
WCNC2
2026 LOREN: Low-Rank-Based Code-Rate Adaptation in Neural Receivers
Bram Van Bolderik, Vlado Menkovski, Sonia M. Heemstra de Groot, Manil Dev Gomony
WCNC4
2026 MEAN: Mixture-of-Experts Neural Receiver - Architecture and Performance Analysis
abstract
Neural network-based wireless receivers, also known as neural receivers , have demonstrated superior performance over traditional receivers, but come with greater computational complexity. The need to use these networks on energy-conscious edge devices is increasing, necessitating energy-efficient and adaptive neural receivers to function well under varying channel conditions. Transitioning static neural receivers to dynamic models using the concepts of Dynamic Neural Networks (DyNN) could reduce runtime computational complexity and energy consumption. This could be achieved by adapting the network architecture during run-time and exploit the varying channel conditions in a mobile communication system to reduce the computational complexity. This work introduces MEAN, a novel hard-gated Mixture-of-Experts (MOE) based neural receiver architecture. The main idea behind MEAN is to use several smaller Signal-to-Noise-Ratio (SNR) expert networks are used during run-time to create a network that selects the correct expert for the current data input to reduce complexity. The paper consists of the following key contributions. (1) MEAN architecture based on [ 25 ]: A hard-gated MoE model that dynamically selects the most suitable expert for each input dynamically based on current channel conditions. (2) Comprehensive system-level performance analysis of MEAN across different code rates and modulation schemes, including validation of expert selection during inference. (3) Loss function optimization for the gating network to promote the activation of specific experts in designated noise regions. This enhancement reduces gating network complexity while improving expert selection accuracy over previous MEAN implementations. (4) Hardware analysis of a gate-level implementation of MEAN synthesized in 22nm FD-SOI technology. The design is generated using a custom High-Level Synthesis (HLS) framework, providing detailed power and area evaluations. The proposed MEAN architecture for a Single Input Multiple Output (SIMO) wireless system achieves a 37.11% reduction in total power consumption with only an 11.04% increase in area, while maintaining accuracy comparable to that of a static neural network.
Bram Van Bolderik, Vlado Menkovski, Sonia M. Heemstra de Groot, Manil Dev Gomony
ACM Trans. Embed. Comput. Syst.4
2025 Multi-Partner Project: Securing Future Edge-AI Processors in Practice (CONVOLVE)
abstract
Artificial Intelligence (AI) has had a profound impact on our contemporary society, and it is indisputable that it will continue to play a significant role in the future. To further enhance AI experience and performance, a transition from large-scale server applications towards AI-powered edge devices is inevitable. In fact, current projections indicate that the market for Smart Edge Processors (SEPs) will grow beyond 70 Billion USD by 2026 [1]. Such a shift comes with major challenges, as these devices have limited computing and energy resources yet need to be highly performant. Additionally, security mechanisms need to be implemented to protect against diverse attack vectors as attackers now have physical access to the device. Besides cryptographic keys, Intellectual Property (IP), including neural network weights, may also be potential targets. The CONVOLVE [2] project (currently in its intermediate stage) follows a holistic approach to address these challenges and establish the EU in a leading position in embedded, ultra-low-power and secure processors for edge computing. It encompasses novel hardware technologies, end-to-end integrated workflows, and a security-by-design approach. This paper highlights the security aspects of future edge-AI processors by illustrating challenges encountered in CONVOLVE, the solutions we pursue including some early results, and directions for future research.
Sven Argo, Henk Corporaal, Alejandro Garza, Marc Geilen, Manil Dev Gomony, Tim Güneysu, Adrian Marotzke, Fouwad Jamil Mir, Jan Richter-Brockmann, Jeffrey Smith 0001, Mottaqiallah Taouil, Said Hamdioui
DATE5
2025 Dependable Neuromorphic Computing-in-Memory Architectures
Farhad Merchant, Ankit Bende, Markus Fritscher, Shahar Kvatinsky, Simranjeet Singh, Vikas Rana, Regina Dittmann, Keerthi Dorai Swamy Reddy, Christian Wenger, Fouwad Jamil Mir, Mottaqiallah Taouil, Manil Dev Gomony, Said Hamdioui, Henk Corporaal
ETS12
2024 Run-time Non-uniform Quantization for Dynamic Neural Networks in Wireless Communication
abstract
Dynamic Neural Networks (DyNN) offer the ability to adapt their structure, parameters, or precision dynamically, making them suitable for systems with rapidly changing environmental conditions, such as wireless communication. Traditional uniform quantization, if applied in DyNNs, will result in unnecessary switching power as the precision requirements are different at different environment conditions. To address this issue, we present two main contributions. 1) An offline non-uniform quantization algorithm enabling run-time quantization adaptation while preserving system performance. 2) A low-overhead dynamic data-gating architecture facilitating run-time non-uniform quantization. The proposed algorithm facilitates dynamic data-gating of up to 8-bits for QPSK demodulation parameters with no performance loss in a Digital Video Broadcast (DVB-S.2) receiver simulation. The DyNN architecture with data-gating, synthesized using GF 22-nm FDSOI CMOS technology achieves a 43% total power reduction with a minimal 3% area overhead compared to the architecture without data-gating.
Priscilla Sharon Allwin, Manil Dev Gomony, Marc Geilen
ASPDAC2
2024 Invited: Achieving PetaOps/W Edge-AI Processing
abstract
Artificial Intelligence (AI) supported by Deep Artificial Neural Networks (ANNs) is booming and already used in many applications, with impressive results, and we are still its infancy. For many sensing applications it would be advantageous if we could move AI from cloud to Edge. However this requires huge improvements in energy-efficiency. The CONVOLVE project (convolve.eu) aims at enabling smart edge devices through a concerted effort at all layers of the design stack. This ranges from using much more efficient models and mappings, like exploiting Spiking Neural Networks (SNNs), to new processing architectures, like compute-in-memory (CIM), use of approximation, and using new device technology, like memristors. However these latter changes make HW more susceptible to noise and other disturbances. Online continuous learning (i.e. adapting weights) may alleviate these problems. This paper shows several CONVOLVE developments in the crucial areas of CIM architectures, SNN accelerators and online learning.
Manil Dev Gomony, Bas Ahn, Rick Luiken, Yashvardhan Biyani, Anteneh Gebregiorgis, Axel Laborieux, Friedemann Zenke, Said Hamdioui, Henk Corporaal
DAC1
2024 Agile Design-Space Exploration of Dynamic Layer-Skipping in Neural Receivers
abstract
Dynamic Neural Networks (DyNN) adapt their structure during runtime for improved performance and lower power consumption. DyNNs benefit neural network-based wireless receivers, or in short neural receivers, that adapt their performance under varying channel conditions. However, DyNN architectures are not yet explored sufficiently in the literature as this would require a framework for simultaneously evaluating system-level performance and accurate hardware Power-Performance-Area (PPA) under different dynamic scenarios. This paper presents two main contributions: (1) An automated framework that bridges a system-level performance evaluation model of a wireless system and a High-Level Synthesis (HLS) tool for agile Design-Space Exploration (DSE) of DyNNs. (2) A novel DyNN architecture for neural receivers in wireless systems that skips a variable number of layers according to varying channel conditions. Our proposed neural receiver architecture with layer-skipping for a Single Input Multi Output (SIMO) wireless system when implemented in 22 nm FD-SOI technology shows power consumption savings of up to 59.2% at the cost of 200% increase in area compared to a static network.
Bram Van Bolderik, Souradip Sarkar, Vlado Menkovski, Sonia M. Heemstra de Groot, Manil Dev Gomony
DSD5
2024 LLRSymNet: A Low-Complex Neural Network for LLR Estimation Through Symmetry Exploitation
abstract
In wireless communication receivers, soft demodulation translates the noisy received symbols into soft decision information, typically in the form of Log Likelihood Ratios (LLR). Of late, Neural Network (NN)-based demodulators show promise, offering improved performance. However, there is a growing concern about the rising complexity in NN based LLR estimation, particularly with denser modulation schemes. This increased complexity leads to higher area and power consumption, posing challenges for efficiency in edge device applications. To address this issue, this paper proposes a novel NN architecture named LLRSymNet, which works by exploiting the symmetries found in the LLR functions of the bits in the received symbols, which stems from the symmetries present in the modulation constellations used for transmission. By doing so, it reduces the computational complexity compared to existing NN-based soft demodulators. Experimental evaluation of LLRSymNet within the DVB-S.2 receiver system demonstrates performance enhancements and achieves complexity reductions of up to 75% for M-PSK and M-APSK modulation schemes, and up to 81% for denser M-QAM modulations, compared to conventional NN architectures.
Priscilla Sharon Allwin, Manil Dev Gomony, Marc Geilen
GLOBECOM2
2024 AFSRAM-CIM: Adder Free SRAM-Based Digital Computation-in-Memory for BNN
abstract
Binary Neural Networks (BNNs) have demonstrated significant advantages in reducing computation and memory costs, all while maintaining acceptable accuracy on various image detection tasks. Thus, BNNs have the potential to support practical cognitive tasks on resource-constrained platforms, such as edge computing devices. To realize this, SRAM-based digital Computation-in-Memory (CIM) has gained growing attention as it overcomes the analog CIM architecture bottlenecks such as limited computing accuracy due to process variation, non-linearity, power and area-hungry Analog-to-Digital Converters (ADCs), etc. However, digital CIM architectures are highly dominated by power-hungry adder-trees, which can nullify the benefits of SRAM-based digital CIM. To address this issue, this paper proposes an adder free SRAM-based digital CIM, AFSRAM-CIM, for BNN acceleration. The proposed CIM architecture utilizes a multi-functional 10-T SRAM cell-based crossbar array and a new energy-efficient approach to perform the popcount operation. Simulation results using the MNIST dataset show that the proposed architecture maintains the state-of-the-art inference accuracy of 99.21% with only 11.86 fJ energy per operation. Moreover, AFSRAM-CIM achieves over$3\times$energy and$\approx 17\times$area savings when compared to the conventional digital CIM approaches.
Asmae El Arrassi, Mohammad Amin Yaldagard, Xingjian Tao, Taha Shahroodi, Fouwad Jamil Mir, Yashvardhan Biyani, Manil Dev Gomony, Anteneh Gebregiorgis, Rajiv V. Joshi, Said Hamdioui
VLSI-SoC7
2024 MEAN: Mixture-of-Experts Based Neural Receiver
abstract
Dynamic Neural Networks (DyNN) adapt its network architecture at run-time compared to static neural networks. DyNNs benefit in wireless receivers based on neural networks (or neural receivers), that need to adapt their performance under varying channel conditions. Mixture-of-Experts (MoE) is one efficient way to realize a DyNN in neural receivers in which several smaller expert networks are dynamically combined in the run-time to create a dedicated network according to the channel requirements. This paper presents a novel hard-gated (also known as sparsely-gated) MoE-based neural receiver architecture called MEAN for Single Input Multiple Output (SIMO)-based wireless communication systems. Our proposed MEAN architecture for a SIMO wireless system shows a reduction of up to 50% in the number of active layers during runtime.
Bram Van Bolderik, Vlado Menkovski, Sonia M. Heemstra de Groot, Manil Dev Gomony
VLSI-SoC4
2024 A Scalable Hardware Architecture for Efficient Learning of Recurrent Neural Networks at the Edge
abstract
Edge devices can execute pre-trained Artificial Intelligence (AI) models optimized on large Graphical Processing Units (GPU) but often need fine-tuning for real-world data. This process, known as edge learning, is crucial for personalized learning for tasks such as speech and gesture recognition and often requires recurrent neural networks (RNNs). However, training RNNs on edge devices faces challenges due to limited resources. We propose a system for RNN training through sequence partitioning using the Forward Propagation Through Time (FPTT) training method, facilitating edge learning. Our optimized HW/SW co-design for FPTT is the first of its kind. In our work, we have implemented the complete computational process for training Long Short-Term Memory (LSTM) networks using FPTT, and we have optimized and explored the hardware architecture leveraging the Chipyard framework. Our findings indicate considerable memory savings, with only a slight increase in latency, when training small-batch size sequential MNIST (S-MNIST) data.
Yicheng Zhang 0006, Manil Dev Gomony, Henk Corporaal, Federico Corradi
VLSI-SoC2
2024 Reconfigurable Signal Processing and DSP Hardware Generator for 5G and Beyond Transmitters
abstract
The digital front-end of the communication transceivers envisioned for fifth-generation (5G) and beyond requires highly configurable high-performance digital signal processing (DSP) hardware operating at very high sampling rates to accommodate increasing signal bandwidths and support a range of modulation schemes and transmitter architectures. In this article, we present an efficient implementation of a highly configurable DSP hardware generator that can generate high-performance DSP hardware for multiple transmitter architectures including Cartesian, polar, outphasing, and multilevel outphasing modulators. The generated hardware unit, which consists of multistage multirate filters and other required DSP operations, runs at sample rates up to 4 GHz. The hardware supports an adjacent channel leakage ratio (ACLR) down to −48 dB and an error vector magnitude (EVM) of 0.78% with a 7-bit phase signal at a sampling rate of 4 GHz for multilevel outphasing modulation. Digital synthesis of the circuit in a 5-nm complimentary metal-oxide semiconductor (CMOS) process yields a core area consumption of 0.01 mm2 and an estimated power consumption of 37.2 mW for a 200-MHz bandwidth 5G new radio (NR) baseband (BB) signal.
Agnimesh Ghosh, Andrei Spelman, Tze Hin Cheung, Dhanashree Boopathy, Kari Stadius, Manil Dev Gomony, Mikko Valkama, Jussi Ryynänen, Marko Kosunen, Vishnu Unnikrishnan 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2023 PetaOps/W edge-AI $\mu$ Processors: Myth or reality?
abstract
With the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality.
Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal
DATE1
2023 Dependability of Future Edge-AI Processors: Pandora's Box
abstract
This paper addresses one of the directions of the HORIZON EU CONVOLVE project being dependability of smart edge processors based on computation-in-memory and emerging memristor devices such as RRAM. It discusses how how this alternative computing paradigm will change the way we used to do manufacturing test. In addition, it describes how these emerging devices inherently suffering from many non-idealities are calling for new solutions in order to ensure accurate and reliable edge computing. Moreover, the paper also covers the security aspects for future edge processors and shows the challenges and the future directions.
Manil Dev Gomony, Anteneh Gebregiorgis, Moritz Fieback, Marc Geilen, Sander Stuijk, Jan Richter-Brockmann, Rajendra Bishnoi, Sven Argo, Lara Arche Andradas, Tim Güneysu, Mottaqiallah Taouil, Henk Corporaal, Said Hamdioui
ETS1
2023 BrainTTA: A 28.6 TOPS/W Compiler Programmable Transport-Triggered NN SoC
abstract
Accelerators designed for deep neural network (DNN) inference with extremely low operand widths, down to 1-bit, have become popular due to their ability to significantly reduce energy consumption during inference. This paper introduces a compiler-programmable flexible System-on-Chip (SoC) with mixed-precision support. This SoC is based on a Transport-Triggered Architecture (TTA) that facilitates efficient implementation of DNN workloads. By shifting the complexity of data movement from the hardware scheduler to the exposed-datapath compiler, DNN workloads can be implemented in an energy efficient yet flexible way. The architecture is fully supported by a compiler and can be programmed using C/C++/OpenCL. The SoC is implemented using 22nm FDX technology and achieves a peak energy efficiency of 28.6/14.9/2.47 TOPS/W for binary, ternary, and 8-bit precision, respectively, while delivering a throughput of 614/307/77 GOPS. Compared to state-of-the-art (SotA), this work achieves up to 3.3x better energy efficiency compared to other programmable solutions.
Maarten Molendijk, Floran de Putter, Manil Dev Gomony, Pekka Jääskeläinen, Henk Corporaal
ICCD3
2022 Dilate-Invariant Temporal Convolutional Network for Real-Time Edge Applications
abstract
Temporal Convolutional Networks (TCNs) involving mono channels as input, have shown superior performance compared to state-of-the-art sequence detection recursive networks in a variety of applications. TCNs leverage the concept of dilated causal convolution for a wider receptive field coverage of input (mono) channels, which requires scaling the delay between input samples in Multiply-Accumulate (MAC) units in different layers. We demonstrate a possible data-flow transformation to convert a dilated convolution to a non-dilated convolution to remove such need for delay scaling while maintaining the same receptive field. The new data-flow transformation allows for hardware units to be shared across all layers with single-delay units between the MAC units. We demonstrate how such data-flow transformation can be easily achieved using generic Finite Impulse Response (FIR) filter modules, simplifying the deployment of TCNs. We validate the predicted savings using Cadence Stratus High-Level Synthesis (HLS). A gesture recognition case study using ultrasound is synthesized achieving 25% savings in both energy and area if the data-flow transformation is applied.
Emad A. Ibrahim, Bart van den Dool, Sayandip De, Manil Dev Gomony, Jos Huisken, Marc Geilen
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 μ-Genie: A Framework for Memory-Aware Spatial Processor Architecture Co-Design Exploration
abstract
Spatial processor architectures are essential to meet the increasing demand in performance and energy efficiency of both embedded and high performance computing systems. Due to the growing performance gap between memories and processors, the memory system of ten determines the overall performance and power consumption in silicon. The interdependency between memory system and spatial processor architectures suggests that they should be co-designed. For the same reason, state-of-the-art design methodologies for processor architectures are ineffective for spatial processor architectures because they do not include the memory system. In this paper, we present μ -Genie: an automated framework for co-design-space exploration of spatial processor architecture and the memory system, starting from an application description in a high-level programming language. In addition, we propose a spatial processor architecture template that can be configured at design-time for optimal hardware implementation. To demonstrate the effectiveness of our approach, we show a case study of co-designing a spatial processor using different memory technologies.
Giulio Stramondo, Manil Dev Gomony, Bartek Kozicki, Cees T. A. M. de Laat, Ana Lucia Varbanescu
DSD2
2019 A Reconfigurable Architecture for Posit Arithmetic
abstract
Physical layer of modern communication systems involves several floating point computations that require higher numerical fidelity and dynamic range for achieving the maximum data rate. Posit number format is a promising alternative for floating point computation as it provides a better dynamic range and numerical fidelity for the same number of bits used for representation. However, the Posit number format has not yet been thoroughly studied in the context of signal processing algorithms in the physical layer. In addition, a configurable Posit arithmetic hardware is essential for adapting to the ever changing communication standards and to configure the dynamic range according to algorithmic demands. The main contributions in this paper are: 1) Performance analysis of common signal processing algorithms using the Posit number format in comparison with the IEEE Standard for Single Precision Floating Point arithmetic (IEEE 754). 2) A novel reconfigurable hardware accelerator for Posit arithmetic operations and comparison with the state-of-the-art Posit and IEEE Floating Point arithmetic architectures. Although our proposed architecture consumes over 3× energy and area compared to IEEE Single Precision arithmetic, our results show that using the Posit number format for physical layer algorithms results in significant performance gain. We achieved over 15 dB and 25 dB gain for FFT and matrix multiplication algorithms. In addition, compared to state-of-the-art Posit arithmetic architecture, our proposed architecture resulted in over 2× speedup in operating frequency and 35% savings in energy consumption.
Souradip Sarkar, Purushotham Murugappa, Manil Dev Gomony
DSD3
2018 Quater-imaginary base for complex number arithmetic circuits
abstract
Arithmetic operations involving complex numbers are widely used in the signal processing functions in the physical layer of modern wireless and wireline communication systems, electronic instrumentation and control systems. With the ever increasing throughput requirements of such systems, the power consumption of the hardware realization is increasing beyond the allowed budget. Arithmetic circuits based on binary numeral system that have been optimized rigorously over the past few decades are currently being used for the computation involving complex numbers. In this paper, we present the potential of arithmetic circuits for complex number computations based on the Quater-imaginary (QI) base numeral system to reduce power consumption. We show that for a simple multiplier implementation in the QI base, the savings in power and area consumption could be up to 40% when synthesized in 28nm TSMC standard cell technology node.
Souradip Sarkar, Manil Dev Gomony
DATE2
2017 A Globally Arbitrated Memory Tree for Mixed-Time-Criticality Systems
abstract
Embedded systems are increasingly based on multi-core platforms to accommodate a growing number of applications, some of which have real-time requirements. Resources, such as off-chip DRAM, are typically shared between the applications using memory interconnects with different arbitration polices to cater to diverse bandwidth and latency requirements. However, traditional centralized interconnects are not scalable as the number of clients increase. Similarly, current distributed interconnects either cannot satisfy the diverse requirements or have decoupled arbitration stages, resulting in larger area, power and worst-case latency. The four main contributions of this article are: 1) a Globally Arbitrated Memory Tree (GAMT) with a distributed architecture that scales well with the number of cores, 2) an RTL-level implementation that can be configured with five arbitration policies (three distinct and two as special cases), 3) the concept of mixed arbitration policies that allows the policy to be selected individually per core, and 4) a worst-case analysis for a mixed arbitration policy that combines TDM and FBSP arbitration.We compare the performance of GAMT with centralized implementations and show that it can run up to four times faster and have over 51 and 37 percent reduction in area and power consumption, respectively, for a given bandwidth.
Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens
IEEE Trans. Computers1
2015 A generic, scalable and globally arbitrated memory tree for shared DRAM access in real-time systems
Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens
DATE1
2015 A Real-Time Multichannel Memory Controller and Optimal Mapping of Memory Clients to Memory Channels
abstract
Ever-increasing demands for main memory bandwidth and memory speed/power tradeoff led to the introduction of memories with multiple memory channels, such as Wide IO DRAM. Efficient utilization of a multichannel memory as a shared resource in multiprocessor real-time systems depends on mapping of the memory clients to the memory channels according to their requirements on latency, bandwidth, communication, and memory capacity. However, there is currently no real-time memory controller for multichannel memories, and there is no methodology to optimally configure multichannel memories in real-time systems. As a first work toward this direction, we present two main contributions in this article: (1) a configurable real-time multichannel memory controller architecture with a novel method for logical-to-physical address translation and (2) two design-time methods to map memory clients to the memory channels, one an optimal algorithm based on an integer programming formulation of the mapping problem, and the other a fast heuristic algorithm. We demonstrate the real-time guarantees on bandwidth and latency provided by our multichannel memory controller architecture by experimental evaluation. Furthermore, we compare the performance of the mapping problem formulation in a solver and the heuristic algorithm against two existing mapping algorithms in terms of computation time and mapping success ratio. We show that an optimal solution can be found in 2 hours using the solver and in less than 1 second with less than 7% mapping failure using the heuristic for realistically sized problems. Finally, we demonstrate configuring a Wide IO DRAM in a high-definition (HD) video and graphics processing system to emphasize the practical applicability and effectiveness of this work.
Manil Dev Gomony, Benny Akesson, Kees Goossens
ACM Trans. Embed. Comput. Syst.1
2014 Coupling TDM NoC and DRAM controller for cost and performance optimization of real-time systems
abstract
Existing memory subsystems and TDM NoCs for real-time systems are optimized independently in terms of cost and performance by configuring their arbiters according to the bandwidth and/or latency requirements of their clients. However, when they are used in conjunction, and run in different clock domains, i.e. they are decoupled, there exists no structured methodology to select the NoC interface width and operating frequency for minimizing area and/or power consumption. Moreover, the multiple arbitration points, one in the NoC and the other in the memory subsystem, introduce additional overhead in the worst-case guaranteed latency. These makes it hard to design cost-efficient real-time systems. The three main contributions in this paper are: (1) We present a novel methodology to couple any existing TDM NoC with a realtime memory controller and compute the different NoC interface width and operating frequency combinations for minimal area and/or power consumption. (2) For two different TDM NoC types, one a packet-switched and the other circuit-switched, we show the trade-off between area and power consumption with the different NoC configurations, for different DRAM generations. (3) We compare the coupled and decoupled architectures with the two NoCs, in terms of guaranteed worst-case latency, area and power consumption by synthesizing the designs in 40 nm technology. Our experiments show that using a coupled architecture in a system consisting of 16 clients results in savings of over 44% in guaranteed latency, 18% and 17% in area, 19% and 11% in power consumption for a packet-switched and a circuit-switched TDM NoC, respectively, with different DRAM types.
Manil Dev Gomony, Benny Akesson, Kees Goossens
DATE1
2013 Architecture and optimal configuration of a real-time multi-channel memory controller
abstract
Optimal utilization of a multi-channel memory, such as Wide IO DRAM, as shared memory in multi-processor platforms depends on the mapping of memory clients to the memory channels, the granularity at which the memory requests are interleaved in each channel, and the bandwidth and memory capacity allocated to each memory client in each channel. Firm real-time applications in such platforms impose strict requirements on shared memory bandwidth and latency, which must be guaranteed at design-time to reduce verification effort. However, there is currently no real-time memory controller for multichannel memories, and there is no methodology to optimally configure multi-channel memories in real-time systems. This paper has four key contributions: (1) A real-time multi-channel memory controller architecture with a new programmable Multi-Channel Interleaver unit. (2) A novel method for logical-to-physical address translation that enables inter-leaving memory requests across multiple memory channels at different granularities. (3) An optimal algorithm based on an Integer Linear Program (ILP) formulation to map memory clients to memory channels considering their communication dependencies, and to configure the memory controller for minimum bandwidth utilization. (4) We experimentally evaluate the run-time of the algorithm and show that an optimal solution can be found within 15 minutes for realistically sized problems. We also demonstrate configuring a multi-channel Wide IO DRAM in a High-Definition (HD) video and graphics processing system to emphasize the effectiveness of our approach.
Manil Dev Gomony, Benny Akesson, Kees Goossens
DATE1
2012 DRAM selection and configuration for real-time mobile systems
abstract
The performance and power consumption of mobile DRAMs (LPDDRs) depend on the configuration of system-level parameters, such as operating frequency, interface width, request size, and memory map. In mobile systems running both real-time and non-real-time applications, the memory configuration must satisfy bandwidth requirements of real-time applications, meet the power consumption budget, and offer the best average-case execution time to the non-real-time applications. There is currently no well-defined methodology for selecting a suitable memory configuration for real-time mobile systems. The worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across generations have furthermore not been investigated. This paper has two main contributions. 1) We analyze the worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across three generations: LPDDR, LPDDR2 and Wide-IO-based 3D-stacked DRAM. 2) Based on our analysis, we propose a methodology for selecting memory configurations in real-time mobile systems.We show that LPDDR (32-bit IO), LPDDR2 (32-bit IO) and 3D-DRAM (128-bit IO) provide worst-case bandwidth up to 0.75 GB/s, 1.6 GB/s and 3.1 GB/s, respectively. We furthermore show for an H.263 decoder that LPDDR2 and 3D-DRAM reduce power consumption with up to 25% and 67%, respectively, compared to LPDDR, and reduce the execution time with up to 18% and 25%.
Manil Dev Gomony, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens
DATE1
2012 Leveraging 802.11n frame aggregation to enhance QoS and power consumption in Wi-Fi networks
Daniel Camps-Mur, Manil Dev Gomony, Xavier Pérez Costa, Sebastià Sallent
Comput. Networks2