EDBT 2026 Demo / reviewers in the wild / expert
Nguyen Anh Vu Doan
dblp:184/5651 · also Ng. Anh Vu Doan
· DBLP profile ↗
13ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-8156-9025ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ApproxiMorph: Energy-Efficient Neuromorphic System With Layer-Wise Approximation of Spiking Neural Networks and 3-D-Stacked SRAMabstractThis paper proposes ApproxiMorph, a comprehensive framework for both software and hardware co-design, targeting energy-efficient AI applications using 3D-IC-based neuromorphic systems. By leveraging parallel interconnections and high-bandwidth communications inherent to 3D-ICs, and the noise-resilience characteristics of spiking neural networks (SNNs), ApproxiMorph achieves significant power savings by exploiting (1) approximate implementation of neuron cells, (2) layer-wise approximation of SNNs through the heuristic exploration algorithm, (3) reduced-voltage operation in the 3D-stacked SRAM, and (4) incorporating a weight-tuning method. As a result, to search for the energy-optimal layer-wise approximation, ApproxiMorph explores only 0.44—0.67% of all possible combinations, achieving a 28.06% power saving for additions with a 0.60% accuracy loss in comparison to the baseline SNN for MNIST. In the VGG16 for CIFAR-10, ApproxiMorph searches around 103 combinations from over 1017 possible solutions, resulting in a 29.16% power saving with slight accuracy gain. Furthermore, integrating all methods enhances the accuracy of approximate implementations and demonstrates higher error resilience than accurate implementations. Ryoji Kobayashi, Ngo-Doanh Nguyen, Ben A. Abderazek, Nguyen Anh Vu Doan, Khanh N. Dang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Generating and Predicting Output Perturbations in Image SegmentersabstractImage segmentation applications are a core component of safety-critical autonomous software pipelines. Sensor data input noise can lead to segmentation output corruption that threatens safety in both DNN- and transformer-based segmenters. Previous work has proposed methods for generating malicious noise to cause DNN- and transformer-based object detection and classification output corruption. We perform the same task for image segmentation applications using genetic algorithms for optimization. We then propose a novel method to predict whether an input image will yield a corrupted segmentation output due to noise. We evaluate the optimal noise generation and corruption prediction on state-of-the-art image segmenters YOLOv8 and DETR. We observe that we can (a) cause segmentation output corruption with noise that is undetectable to the human eye and unrelated to the corrupted region of the image; and (b) predict output corruption due to image noise with over 96% accuracy. Matthew Bozoukov, Nguyen Anh Vu Doan, Bryan Donyanavard |
DATE | 2 |
| 2023 | Butterfly Effect Attack: Tiny and Seemingly Unrelated Perturbations for Object DetectionabstractThis work aims to explore and identify tiny and seemingly unrelated perturbations of images in object detection that will lead to performance degradation. While tininess can naturally be defined using$L_{p}$norms, we characterize the degree of “unrelatedness” of an object by the pixel distance between the occurred perturbation and the object. Triggering errors in prediction while satisfying two objectives can be formulated as a multi-objective optimization problem where we utilize genetic algorithms to guide the search. The result successfully demonstrates that (invisible) perturbations on the right part of the image can drastically change the outcome of object detection on the left. An extensive evaluation reaffirms our conjecture that transformer-based object detection networks are more susceptible to butterfly effects in comparison to single-stage object detection networks such as YOLOv5. Nguyen Anh Vu Doan, Arda Yüksel, Chih-Hong Cheng |
DATE | 1 |
| 2022 | AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic AlgorithmsabstractWe present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele |
DATE | 10 |
| 2022 | Region of interest based non-dominated sorting genetic algorithm-II: an invite and conquer approachabstractEvolutionary multi-objective optimization plays a vital role in solving many complex real-world optimization problems. Numerous approaches have been proposed over the years, and popular methods such as NSGA and its variants incorporate non-dominated sorting selection into evolutionary genetic algorithms to extract competing Pareto-optimal solutions from all over the objective space. However, in applications where the decision-maker is interested in a region of interest, a global optimization wastes effort to find irrelevant solutions outside of the preferred region. In this work, we propose an approach named ROI-NSGA-II to limit the optimization effort to a region of interest defined by the boundaries provided by the decision-maker. The ROI-NSGA-II invites the classical NSGA-II algorithm into the desired region using a modified dominance relation and conquers solutions within this region using a modified crowding distance based selection. The effectiveness of our approach is demonstrated on a set of benchmark problems with up to ten objectives and a real-world application, and the results are compared to a state-of-the-art R-NSGA-II. Manu Manuel, Benjamin Hien, Simon Conrady, Arne Kreddig, Nguyen Anh Vu Doan, Walter Stechele |
GECCO | 5 |
| 2021 | Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core ProcessorsabstractWe propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art. Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
DATE | 6 |
| 2021 | HW-FlowQ: A Multi-Abstraction Level HW-CNN Co-design Quantization MethodologyabstractModel compression through quantization is commonly applied to convolutional neural networks (CNNs) deployed on compute and memory-constrained embedded platforms. Different layers of the CNN can have varying degrees of numerical precision for both weights and activations, resulting in a large search space. Together with the hardware (HW) design space, the challenge of finding the globally optimal HW-CNN combination for a given application becomes daunting. To this end, we propose HW-FlowQ, a systematic approach that enables the co-design of the target hardware platform and the compressed CNN model through quantization. The search space is viewed at three levels of abstraction, allowing for an iterative approach for narrowing down the solution space before reaching a high-fidelity CNN hardware modeling tool, capable of capturing the effects of mixed-precision quantization strategies on different hardware architectures (processing unit counts, memory levels, cost models, dataflows) and two types of computation engines (bit-parallel vectorized, bit-serial). To combine both worlds, a multi-objective non-dominated sorting genetic algorithm (NSGA-II) is leveraged to establish a Pareto-optimal set of quantization strategies for the target HW-metrics at each abstraction level. HW-FlowQ detects optima in a discrete search space and maximizes the task-related accuracy of the underlying CNN while minimizing hardware-related costs. The Pareto-front approach keeps the design space open to a range of non-dominated solutions before refining the design to a more detailed level of abstraction. With equivalent prediction accuracy, we improve the energy and latency by 20% and 45% respectively for ResNet56 compared to existing mixed-precision search methods. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Nguyen Anh Vu Doan, Christian Unger, Naveen Shankar Nagaraja, Maurizio Martina, Walter Stechele |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2020 | A Lightweight Nonlinear Methodology to Accurately Model Multicore Processor PowerabstractMany power management algorithms demand accurate and fine-grained runtime estimations of dynamic core power. In the absence of fine-grained power sensors, model-based estimations are needed. Such power models commonly approximate the switching activity of logic gates using performance counters while assuming a linear performance counter/power relation at a fixed frequency and voltage. It has been shown that this relation cannot be captured accurately enough with purely linear models and that well-established nonlinear modeling techniques, e.g., polynomial modeling, easily overfit the underlying performance/power relations. Although neural-network-based modeling has shown to accurately capture nonlinear relations, it has a large training and inference overhead which is too high for fine-grained models on core-level and estimation rates in the range of 1-10 kHz. We propose a methodology for nonlinear transformation of specific performance counters to increase power modeling accuracy at constant frequency and voltage with a relatively low overhead for both model generation and run-time application over a linear model. Furthermore, we use least-angle regression (LARS) to determine a ranking of the performance counter inputs for use in linear and nonlinear modeling and show that the transformed performance counters are better suited for power modeling. The generated dynamic power model consisting of a nonlinear transformation block and a linear regression block reduces relative estimation error on average by 4% and in worst-case scenarios by 7% compared to state-of-the-art fine-grained linear power models. Compared to a state-of-the-art polynomial regression model our proposed approach reduces the relative estimation error by 10% in worst-case scenarios. Mark Sagi, Nguyen Anh Vu Doan, Martin Rapp, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | GenMap: A Genetic Algorithmic Approach for Optimizing Spatial Mapping of Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) are expected to be used for embedded systems, Internet of Things (IoT) devices, and edge computing thanks to their high-energy efficiency and programmability. In essence, a CGRA is an array of numerous processing elements. To exploit this abundant computation resource, a compiler for CGRAs has to fulfill more tasks compared that for general-purpose processors. Therefore, many studies have proposed optimization methods, especially for application mapping, because the performance and energy efficiency strongly depend on optimization at compile time. However, many works focus only on performance improvement or resource minimization, although such optimization objectives are not always appropriate when considering various use cases. In this work, we propose GenMap, an application mapping framework using multiobjective optimization based on a genetic algorithm so that users can set optimization criteria as needed. Besides, it provides aggressive power optimization using our dynamic power model and leakage minimization technique. The proposed method is applied to three fabricated CGRA chips for evaluation. Experimental results show that GenMap achieves 15.7% reduction of wire length while keeping processing element utilization when compared with conventional methods. In addition, according to real chip experiments, 12.1%-46.8% of energy consumption is reduced, and up to $2\times $ speedup is archived for several architectures when compared with other two approaches. Takuya Kojima, Nguyen Anh Vu Doan, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Channel mapping strategies for effective protection switching in fail-operational hard real-time NoCsabstractWith Multi Processor System-on-Chips (MPSoC) scaling up to thousands of processing elements, bus-based solutions have been dropped in favor of Network-on-Chips (NoC) as proposed in [2]. However, MPSoCs are yet hesitantly adopted in safety-critical fields, mainly due to the difficulty of ensuring strict isolation between different applications running on a single MPSoC as well as providing communication with Guaranteed Service (GS) to critical applications. This is particularly difficult in the NoC as it constitutes a network of shared resources. Moreover, safety-critical applications require some degree of Fault-Tolerance (FT) to guarantee safe operation at all times. Max Koenen, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 2 |
| 2019 | APEC: improved acknowledgement prioritization through erasure coding in bufferless NoCsabstractBufferless NoCs have been proposed as they come with a decreased silicon area footprint and a reduced power consumption, when compared to buffered NoCs. However, while known for their inherent simplicity, they suffer from early saturation and depend on additional measures to ensure reliable packet delivery, such as control protocols based on ACKs or NACKs. In this paper, we propose APEC, a novel concept for bufferless NoCs that allows to prioritize ACKs and NACKs over single payload flits of colliding packets by discarding the latter. Lightweight heuristic erasure codes are used to compensate for discarded payload flits. By trading off the erasure code overhead for packet retransmissions, a more efficient network operation is achieved. For ACK-based networks, APEC saturates at 2.1x and 2.875x higher generation rates than a conventional ACK-based bufferless NoC for packets between 5 and 17 flits. For NACK-based networks, APEC does not require concepts such as deflection routing or circuit-switched overlay NACK-networks, as prior work does. Therefore, it can simplify the network implementation compared to prior work while achieving similar performance. Michael Vonbun, Adrian Schiechel, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 3 |
| 2017 | Body bias optimization for variable pipelined CGRAabstractVariable Pipeline Cool Mega Array (VPCMA) is an low power Coarse Grained Reconfigurable Architecture (CGRA) based on the concept of CMA (Cool Mega Array). It implements a pipeline structure that can be configured depending on performance requirements, and the silicon on thin buried oxide (SOTB) technology that allows to control its body bias voltage to balance performance and leakage power. In this paper, we propose a methodology to optimize exactly with an Integer Linear Program the VPCMA body bias while considering simultaneously its variable pipeline structure. For the studied applications, we evaluate that it is possible to achieve an average reduction of energy consumption of 19.3% and 11.8% when compared to respectively the zero bias (without body bias control) and the uniform (control of the whole PE array) cases, while respecting performance constraints. Besides, with appropriate body bias control, it is possible to extend the possible performance, hence enabling broader trade-off analyzes between consumption and performance. These promising results show that applying an adequate optimization technique for the body bias control while simultaneously considering pipeline structures can not only enable further power reduction than previous methods, but also allow more trade-off analysis possibilities. Takuya Kojima, Naoki Ando, Hayate Okuhara, Nguyen Anh Vu Doan, Hideharu Amano |
FPL | 4 |
| 2017 | XYZ-Randomization using TSVs for Low-Latency Energy Efficient 3D-NoCsabstractIn this paper, we propose a method to design low latency and low energy networks for 3D Network-on-Chip (3D-NoC). Recent many-core processors require low-latency interconnection networks since the increasing number of cores limits the network performance. To achieve high performance in such many-core chips, small-world or random networks have been applied in the NoC field. However, the actual diameters and average shortest path lengths (ASPL) of these networks are far from the theoretical lower bound. In this work, we propose an approach based on the graph theory to design ultra low-latency topologies. We introduce a method to design a network that has low values of diameter and ASPL, with configurable upper bound of wire length, called opt ASPL. We also show that irregular topology, such as the topology used in opt ASPL, has a higher average energy consumption than general regular topology like 3D torus. In NoCs, energy budget and link length are limited, and thus such parameters must be carefully considered. Therefore, we introduce a multi-objective optimization for the ASPL and energy consumption called opt A/e which can obtain the Pareto optimal set useful for NoC designers. In a router with 64 nodes per chips and 4 chips stacked with a 3D-NoC, our proposed network optimized for energy consumption has a lower ASPL by 26.8% and a lower energy consumption by 10.9% compared to a 3D torus. Hiroshi Nakahara, Nguyen Anh Vu Doan, Ryota Yasudo, Hideharu Amano |
NOCS | 2 |