VLDB 2026 Research / reviewers in the wild / expert
Joren Dumoulin
dblp:380/5734
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0005-0692-1227ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU IntegrationabstractEmerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microscaling (MX) standard enables narrow bit-widths for inference and large dynamic ranges for training. However, existing MX multiply-accumulate (MAC) designs face a critical trade-off: integer accumulation requires expensive conversions from narrow floating-point products, while FP32 accumulation suffers from quantization losses and costly normalization. To address these limitations, we propose a hybrid precision-scalable reduction tree for MX MACs that combines the benefits of both approaches, enabling efficient mixed-precision accumulation with controlled accuracy relaxation. Moreover, we integrate an 8x8 array of these MACs into the state-of-the-art (SotA) NPU integration platform, SNAX, to provide efficient control and data transfer to our optimized precision-scalable MX datapath. We evaluate our design both on MAC and system level and compare it to the SotA. Our integrated system achieves an energy efficiency of 657, 1438-1675, and 4065 GOPS/W, respectively, for MXINT8, MXFP8/6, and MXFP4, with a throughput of 64, 256, and 512 GOPS. Stef Cuyckens, Xiaoling Yi, Robin Geens, Joren Dumoulin, Martin Wiesner, Chao Fang 0005, Marian Verhelst |
ASP-DAC | 4 |
| 2026 | The Configuration Wall: Characterization and Elimination of Accelerator Configuration OverheadabstractContemporary compute platforms increasingly offload compute kernels from CPU to integrated hardware accelerators to reach maximum performance per Watt. Unfortunately, the time the CPU spends on setup control and synchronization has increased with growing accelerator complexity. For systems with complex accelerators, this means that performance can be configuration-bound. Faster accelerators are more severely impacted by this overlooked performance drop, which we call the configuration wall. Prior work evidences this wall and proposes ad-hoc solutions to reduce configuration overhead. However, these solutions are not universally applicable, nor do they offer comprehensive insights into the underlying causes of performance degradation. In this work, we first introduce a widely-applicable variant of the well-known roofline model to quantify when system performance is configuration-bound. To move systems out of the performance-bound region, we subsequently propose a domain-specific compiler abstraction and associated optimization passes. We implement the abstraction and passes in the MLIR compiler framework to run optimized binaries on open-source architectures to prove its effectiveness and generality. Experiments demonstrate a geomean performance boost of 2x on the open-source OpenGeMM system, by eliminating redundant configuration cycles and by automatically hiding the remaining configuration cycles. Our work provides key insights in how accelerator performance is affected by setup mechanisms, thereby facilitating automatic code generation for circumventing the configuration wall. Josse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols, Xiaoling Yi, Ryan Antonio, Jackson Woodruff, Tobias Grosser, Marian Verhelst |
ASPLOS (1) | 3 |
| 2026 | A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
Xiaoling Yi, Ryan Antonio, Yunhao Deng, Fanchen Kong, Joren Dumoulin, Jun Yin 0001, Marian Verhelst |
ISCAS | 5 |
| 2026 | Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM InferenceabstractTernary weight quantization (e.g., BitNet b1.58) offers a promising path to mitigate the memory bandwidth bottleneck in Large Language Model (LLM) inference. However, conventional compute platforms lack native support for ternary-weight arithmetic, often relying on inefficient dequantization. Lookup table (LUT)-based hardware architectures provide an effective alternative by replacing multiplications with conditional additions, but their design space remains largely unexplored. Existing designs rely on heuristic parameter selection, lacking a systematic understanding of the architectural trade-offs. This work addresses this gap by formalizing the design space of ternary LUT-based accelerators and presenting an open-source hardware generator coupled with an analytical cost model, validated against synthesis in TSMC 16 nm technology. By spanning the full architectural space, this framework not only enables rapid design space exploration but also establishes a common footing for fair cross-design evaluation, which was previously hindered by inconsistent instantiations across published accelerators. Using this framework, we challenge several assumptions and design choices in recent literature. We demonstrate that the optimal architecture is fundamentally governed by the activation data type: while LUT-based reuse offers significant gains for high-cost arithmetic (e.g., FP16), it yields diminishing returns for small integer types. Furthermore, we show that maximizing core size consistently improves area density compared to highly tiled approaches. Our optimized designs achieve a $2.2 \times$ area reduction compared to multiplier-based baselines. Moreover, by benchmarking state-of-the-art implementations against our model, we reveal that correcting suboptimal parameters yields up to a $1.2 \times$ area improvement. These findings underscore that achieving state-of-the-art efficiency requires moving beyond heuristics to a model-driven design methodology. This transition is enabled by our open-source hardware generator available at https://github.com/KULeuven-MICAS/ternary-lut-dse. Robin Geens, Joran Heldens, Joren Dumoulin, Marian Verhelst |
ISPASS | 3 |
| 2026 | FlexiGen: An Automated AI Accelerator Generation Framework With Decoupled-Access-Execute and Dynamic DataflowsabstractModern tensor applications, especially artificial intelligence (AI) applications, are evolving rapidly, posing a significant demand for agile hardware design. While numerous hardware generators have been developed, they suffer from three significant limitations: 1) they are either limited to a single dataflow/data type generation, failing to cater to the computational requirements of diverse workloads; 2) or focus only on the array level optimization, omitting system-level effects, such as the influence of on-chip memory bandwidth and contention; and 3) customized workload mapping/configuration is needed, resulting in increased programming complexity. To address these challenges, we proposeFlexiGen, a flexible and extensible hardware generation framework, which targets diverse deep neural networks (DNN) tensor applications and can generate a complete synthesizable acceleration system at the RTL level with arbitrary dataflow and its combinations. Our key contributions are threefold: 1) we incorporate decoupled-access-execute architecture insideFlexiGen, enabling full system generation while maintaining flexibility and efficiency; 2) we propose a versatile spatial core generator that supports dynamic spatial dataflows and multiple data precisions in the same array and a compatible data streaming engine generator that can support arbitrary temporal dataflows and$N$-dimensional data access patterns; and 3) we leverage a uniform programming interface and provide a customized kernel library, enabling agile configuration programming. We conduct an intensive evaluation to demonstrate the versatility ofFlexiGenin dataflow accelerator generation and show the trade-offs of performance, area, and power across a wide range of dataflows and workloads at both the array level and system level. Our case study experiment showsFlexiGen’s usefulness as a hardware generator to rapidly generate desired dataflow acceleration systems. Compared with the state-of-the-art (SotA) hardware generation framework LEGO,FlexiGenachieves 36.79% and 57.16% less area and power when generating the same dual spatial dataflow design.FlexiGenis open-source and available athttps://github.com/KULeuven-MICAS/snax_cluster Xiaoling Yi, Man Shi, Joren Dumoulin, Jiacong Sun, Yunhao Deng, Ryan Antonio, Fanchen Kong, Marian Verhelst |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | OpenGeMM: A Highly-Efficient GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory CouplingabstractDeep neural networks (DNNs) face significant challenges when deployed on resource-constrained extreme edge devices due to their computational and data-intensive nature. While standalone accelerators tailored for specific application scenarios suffer from inflexible control and limited programmability, generic hardware acceleration platforms coupled with RISC-V CPUs can enable high reusability and flexibility, yet typically at the expense of system-level efficiency and low utilization. Xiaoling Yi, Ryan Antonio, Joren Dumoulin, Jiacong Sun, Josse Van Delm, Guilherme Paim, Marian Verhelst |
ASP-DAC | 3 |
| 2025 | An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator SystemsabstractHeterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movement efficiency and encounter compatibility issues in hardware and software. This prevents a unified approach that balances performance and ease of use. To this end, we present SNAX, an open-source integrated HW-SW framework enabling efficient multi-accelerator platforms through a novel hybrid-coupling scheme, consisting of loosely coupled asynchronous control and tightly coupled data access. SNAX brings reusable hardware modules designed to enhance compute accelerator utilization, and its customizable MLIR-based compiler to automate key system management tasks, jointly enabling rapid development and deployment of customized multi-accelerator compute clusters. Through extensive experimentation, we demonstrate SNAX’s efficiency and flexibility in a low-power heterogeneous SoC. Accelerators can be easily integrated and programmed to achieve >10× improvement in neural network performance compared to other accelerator systems while maintaining accelerator utilization of >90% in full system operation. Ryan Antonio, Joren Dumoulin, Xiaoling Yi, Josse Van Delm, Yunhao Deng, Guilherme Paim, Marian Verhelst |
ISLPED | 2 |
| 2024 | Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the EdgeabstractHybrid vision transformers combine the elements of conventional neural networks (NN) and vision transformers (ViT) to enable lightweight and accurate detection. However, several challenges remain for their efficient deployment on resource-constrained edge devices. The hybrid models suffer from a widely diverse set of NN layer types and large intermediate data tensors, hampering efficient hardware acceleration. To enable their execution at the edge, this paper proposes innovations across the hardware-scheduling stack: a.) At the lowest level, a configurable PE array supports all hybrid ViT layer types; b.) temporal loop re-ordering within one layer, enabling hardware support for normalization and softmax layers, minimizing on-chip data transfers; c.) further scheduling optimization employs layer fusion across inverted bottleneck layers to drastically reduce off-chip memory transfers. The resulting accelerator is implemented in 28nm CMOS, achieving a peak energy efficiency of 1.39 TOPS/W at 25.6 GMACs/s. Joren Dumoulin, Pouya Houshmand, Vikram Jain, Marian Verhelst |
ISCAS | 1 |