VLDB 2026 Research / reviewers in the wild / expert
Timon Evenblij
dblp:258/5860
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-5337-0617ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing Machine Learning Force Fields as Emerging Molecular Dynamics Workloads on Graphics Processing UnitsabstractMolecular dynamics (MD) simulates the time evolution of atomic systems governed by interatomic forces, and the fidelity of these simulations depends critically on the underlying force model. Classical force fields (CFFs) rely on fixed functional forms fitted to experimental or theoretical data, offering computational efficiency and broad applicability but limited accuracy in chemically diverse or reactive environments. In contrast, machine learning force fields (MLFFs) deliver near-quantum-chemical accuracy at molecular-mechanics cost by learning interatomic interactions directly from high-level electronic-structure data.While MLFFs offer improved accuracy at a fraction of the cost of quantum methods, they introduce significant computational overhead, particularly in descriptor evaluation and neural network inference. These operations pose challenges for parallel hardware due to irregular memory access, minimum data reuse and inefficient kernel execution.This work investigates the hardware performance of such models using poly-alanine chains, a novel benchmark molecule system(s) with controllable input size, which used as performance evaluation test cases highlighting the computational bottlenecks of the graphical processor units when scaling out MLFF simulations. The analysis identifies key bottlenecks in descriptor and force computation, memory handling, highlighting the opportunities for improvements in the emerging area of MLFF based MD in drug discovery, that has received limited attention from a computer architecture perspective. Udari De Alwis, Benjamin E. Mayer, Tom J. Ashby, Maria Barrera, Timon Evenblij, Joyjit Kundu |
ISPASS | 5 |
| 2025 | Energon: A Sustainability-Driven Modeling Framework for AI Data CentersabstractRecent exponential investments in data centers, purposefully built for artificial intelligence (AI) workloads, have raised significant concerns around the sustainability of AI. The need for holistic sustainability analysis during the initial design phase of data centers and their components has become increasingly stringent. Such analysis must account for future technological advances in hardware and software, uphold a high level of accuracy, and operate with high speed to explore the large design space of a data center system. Sustainability in this context is a multifaceted concept encompassing various metrics such as performance, power consumption, energy efficiency, physical footprint, cost, and both embodied and operational emissions. This work aims to continue the discussion around such early design phase sustainability analysis by introducing Energon: a uniquely positioned, fast, end-to-end framework for early hardware-software codesign of carbon-efficient AI data centers. Wenzhe Guo, Joyjit Kundu, Uras Tos, Giuliano Sisto, Cedric Rolin, L.-Å. Ragnarsson, Timon Evenblij |
ISPASS | 7 |
| 2025 | PARL: Page Allocation in hybrid main memory using Reinforcement Learning
Emil Karimov, Timon Evenblij, Saeideh Alinezhad Chamazcoti, Francky Catthoor |
J. Syst. Archit. | 2 |
| 2024 | Bank on Compute-Near-Memory: Design Space Exploration of Processing-Near-Bank ArchitecturesabstractNear-DRAM computing strategies advocate for providing computational capabilities close to where data is stored. Although this paradigm can effectively address the memory-to-processor communication bottleneck, it also presents new challenges: The strict resource constraints in the memory periphery demand careful tailoring of architectural elements. We herein propose a novel framework and methodology to explore compute-near-memory designs that interface to DRAM memory banks, demonstrating the area, energy, and performance tradeoffs subject to the architectural configuration. We exemplify this methodology by conducting two studies on compute-near-bank designs: 1) analyzing the interaction between control and data resources, and 2) exploring the integration of processing units with different DRAM standards. According to our study, the optimal size ratios between instruction and data capacity vary from$2\times $to$4\times $across benchmarks from representative application domains. The retrieved Pareto-optimal solutions from our framework improve state-of-the-art designs, e.g., achieving a 50% performance increase on matrix operations with 15% energy overhead relative to the FIMDRAM design. In addition, the exploration of DRAM shows the interplay between available internal bandwidth, performance, and area overhead. For example, a threefold increase in bandwidth rises performance by 47% across workloads at a 34% extra area cost. Rafael Medina 0001, Giovanni Ansaloni, Marina Zapater, Alexandre Levisse, Saeideh Alinezhad Chamazcoti, Timon Evenblij, Dwaipayan Biswas, Francky Catthoor, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | AMPeD: An Analytical Model for Performance in Distributed Training of TransformersabstractTransformers are a class of machine learning models that have piqued high interest recently due to a multitude of reasons. They can process multiple modalities efficiently and have excellent scalability. Despite these obvious advantages, training these large models is very time-consuming. Hence, there have been efforts to speed up the training process using efficient distributed implementations. Many different types of parallelism have been identified that can be employed standalone or in combination. However, naively combining different parallelization schemes can incur significant communication overheads, thereby potentially defeating the purpose of distributed training. Thus, it becomes vital to predict the right mapping of different parallelisms to the underlying system architecture. In this work, we propose AMPeD, an analytical model for performance in distributed training of transformers. It exposes all the transformer model parameters, potential parallelism choices (along with their mapping onto the system), the accelerator as well as system architecture specffications as tunable knobs, thereby enabling hardware-software co-design. With the help of 3 case studies, we show that the combinations of parallelisms predicted to be efficient by AMPeD conform with the results from the state-of-the-art literature. Using AMPeD, we also show that future distributed systems consisting of optical communication substrates can train large models up to 4× faster as compared to the current state-of the-art systems without modifying the peak computational power of the accelerators. Finally, we validate AMPeD with in-house experiments on real systems and via published literature. The max. observed error is limited to 12%. The model is available here: https://github.com/CSA-infra/AMPeD Diksha Moolchandani, Joyjit Kundu, Frederik Ruelens, Peter Vrancx, Timon Evenblij, Manu Perumkunnil Komalan |
ISPASS | 5 |
| 2023 | Exploring Pareto-Optimal Hybrid Main Memory Configurations Using Different Emerging MemoriesabstractMain memory system design and corresponding technology requirements have become increasingly challenging for data-dominated high-performance applications. To address the leakage and scalability issues of the conventional DRAM-based memory, new memory technologies with ultra-low leakage and potential for high scalability have been explored extensively over the last decade. However, none of them are mature enough to serve as a drop-in replacement for DRAM. In this paper, we propose a hybrid main memory system solution for utilizing new memory technologies with specific features, based on the target application characteristics and system configurations. To this end, we examine two new memories, 1S-1VCMA and IGZO-based DRAM, along with conventional DRAM in the context of hybrid main memory solutions for high-capacity and low-power Pareto-optimizations, respectively. To better evaluate the power and performance, we consider the page-fault modeling in our evaluations. The results of the simulation show that different combinations of memory technologies in the hybrid memory system, different memory capacities, and different storage systems could provide a promising solution in the system regarding the characteristics of running applications and the requirements of the system. Saeideh Alinezhad Chamazcoti, Mohit Gupta 0004, Hyungrock Oh, Timon Evenblij, Francky Catthoor, Manu Perumkunnil Komalan, Gouri Sankar Kar, Arnaud Furnémont |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Time-Dependent Electromigration Modeling for Workload-Aware Design-Space Exploration in STT-MRAMabstractElectromigration (EM) has been known as a reliability threatening factor for back-end-of-the-line interconnects. Spin-transfer torque magnetic RAM (STT-MRAM) is an emerging nonvolatile memory that has gained a lot of attention in recent years. However, relatively large operational current magnitude is a challenge for this technology, and hence, EM can be a potential reliability concern, even for the signal lines of this memory. A workload-aware EM modeling needs to capture time-dependent current density in the memory signal lines and to be able to predict the effect of the EM phenomenon on the interconnect for its entire lifetime. In this work, we present methods to effectively model the workload-dependent EM-induced meantime to failure (MTTF) in typical STT-MRAM arrays under a variety of realistic workloads. This allows performing the design-space exploration to co-optimize reliability and other design metrics. Mahta Mayahinia, Mehdi Baradaran Tahoori, Manu Perumkunnil Komalan, Houman Zahedmanesh, Kris Croes, Tommaso Marinelli, José Ignacio Gómez, Timon Evenblij, Gouri Sankar Kar, Francky Catthoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2021 | Memory Hierarchy Calibration Based on Real Hardware In-order Cores for Accurate SimulationabstractComputer system simulators are major tools used by architecture researchers. Two key elements play a role in the credibility of simulator results: (1) the simulator's accuracy, and (2) the quality of the baseline architecture. Some simulators, such as gem5, already provide highly accurate parameterized models. However, finding the right values for all these parameters to faithfully model a real architecture is still a problem. In this paper, we calibrate the memory hierarchy of an in-order core gem5 simulation to accurately model a real mobile Arm SoC. We execute small programs, which we design to stress specific parts of the memory system, to deduce key parameter values for the model. We compare the execution of SPEC CPU2006 benchmarks on the real hardware with the gem5 simulation. Our results show that our calibration reduces the average and worst-case IPC error by 36 % and 50%, respectively, when compared with a gem5 simulation configured with the default parameters. Quentin Huppert, Timon Evenblij, Manu Perumkunnil Komalan, Francky Catthoor, Lionel Torres, David Novo |
DATE | 2 |
| 2019 | A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCsabstractSpin Transfer Torque Magnetic RAM (STT-MRAM) is being extensively considered as a promising replacement for Last Level Caches (LLC), due to its high density, low leakage and non-volatility. However, writes to STT-MRAM are energy intensive and have a high latency. While the high dynamic energy consumption during writes can be compensated by the low static energy consumption, the high latency results in performance degradation. This work shows that in contrast to SRAM-based LLCs, the performance degradation for STT-MRAM is primarily due to bank contention, when trying to satisfy a read request while the bank is being written. We holistically explore the effects of cache banking and cache contention on energy and performance in the LLC of mobile multicore systems, with in-order cores or with out-of-order cores. The detail of the analysis is enabled by highly accurate cache models, based on a 28nm SRAM industry compiler, and an in-house developed STT-MRAM compiler, which generates full STT-MRAM macro designs with silicon-validated MTJ stack and complete parasitic extraction at the 28nm node. Our results show that there is a clear difference in the energy-performance optimal banking configuration between STT-MRAM caches and SRAM caches. These low contention STT-MRAM cache designs with the optimal number of banks save at least 60% cache energy while losing at most single digit percentages in system performance compared to SRAM cache designs. This show an increased potential of using STT-MRAM as a replacement for SRAM in an LLC. Timon Evenblij, Christian Tenllado, Manu Perumkunnil Komalan, Francky Catthoor, Sushil Sakhare, Peter Debacker, Gouri Sankar Kar, Arnaud Furnémont, Nicolas Bueno, José Ignacio Gómez |
ICCD | 1 |