Xiaoling Yi

dblp:333/3736 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 6 first-author · 18 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
abstract
Emerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microscaling (MX) standard enables narrow bit-widths for inference and large dynamic ranges for training. However, existing MX multiply-accumulate (MAC) designs face a critical trade-off: integer accumulation requires expensive conversions from narrow floating-point products, while FP32 accumulation suffers from quantization losses and costly normalization. To address these limitations, we propose a hybrid precision-scalable reduction tree for MX MACs that combines the benefits of both approaches, enabling efficient mixed-precision accumulation with controlled accuracy relaxation. Moreover, we integrate an 8x8 array of these MACs into the state-of-the-art (SotA) NPU integration platform, SNAX, to provide efficient control and data transfer to our optimized precision-scalable MX datapath. We evaluate our design both on MAC and system level and compare it to the SotA. Our integrated system achieves an energy efficiency of 657, 1438-1675, and 4065 GOPS/W, respectively, for MXINT8, MXFP8/6, and MXFP4, with a throughput of 64, 256, and 512 GOPS.
Stef Cuyckens, Xiaoling Yi, Robin Geens, Joren Dumoulin, Martin Wiesner, Chao Fang 0005, Marian Verhelst
ASP-DAC2
2026 The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
abstract
Contemporary compute platforms increasingly offload compute kernels from CPU to integrated hardware accelerators to reach maximum performance per Watt. Unfortunately, the time the CPU spends on setup control and synchronization has increased with growing accelerator complexity. For systems with complex accelerators, this means that performance can be configuration-bound. Faster accelerators are more severely impacted by this overlooked performance drop, which we call the configuration wall. Prior work evidences this wall and proposes ad-hoc solutions to reduce configuration overhead. However, these solutions are not universally applicable, nor do they offer comprehensive insights into the underlying causes of performance degradation. In this work, we first introduce a widely-applicable variant of the well-known roofline model to quantify when system performance is configuration-bound. To move systems out of the performance-bound region, we subsequently propose a domain-specific compiler abstraction and associated optimization passes. We implement the abstraction and passes in the MLIR compiler framework to run optimized binaries on open-source architectures to prove its effectiveness and generality. Experiments demonstrate a geomean performance boost of 2x on the open-source OpenGeMM system, by eliminating redundant configuration cycles and by automatically hiding the remaining configuration cycles. Our work provides key insights in how accelerator performance is affected by setup mechanisms, thereby facilitating automatic code generation for circumventing the configuration wall.
Josse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols, Xiaoling Yi, Ryan Antonio, Jackson Woodruff, Tobias Grosser, Marian Verhelst
ASPLOS (1)5
2026 Torrent : A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
abstract
The growing disparity between computational power and on-chip communication bandwidth is a critical bottleneck in modern Systems-on-Chip (SoCs), especially for data-parallel workloads like AI. Efficient point-to-multipoint (P2MP) data movement, such as multicast, is essential for high performance. However, native multicast support is lacking in standard inter-connect protocols. Existing P2MP solutions, such as multicast- capable Network-on-Chip (NoC), impose additional overhead to the network hardware and require modifications to the interconnect protocol, compromising scalability and compatibility.This paper introduces Torrent, a novel distributed DMA architecture that enables efficient P2MP data transfers without modifying NoC hardware and interconnect protocol. Torrent conducts P2MP data transfers by forming logical chains over the NoC, where the data traverses through targeted destinations resembling a linked list. This Chainwrite mechanism preserves the P2P nature of every data transfer while enabling flexible data transfers to an unlimited number of destinations. To optimize the performance and energy consumption of Chainwrite, two scheduling algorithms are developed to determine the optimal chain order based on NoC topology.Our RTL and FPGA prototype evaluations using both synthetic and real workloads demonstrate significant advantages in performance, flexibility, and scalability over network-layer multicast. Compared to the unicast baseline, Torrent achieves up to a 7.88 × speedup. ASIC synthesis on 16nm technology confirms the architecture’s minimal footprint in area (1.2%) and power (2.3%). Thanks to the Chainwrite, Torrent delivers scalable P2MP data transfers with a small cycle overhead of 82CC and area overhead of 207 μm2per destination.
Yunhao Deng, Fanchen Kong, Xiaoling Yi, Ryan Antonio, Marian Verhelst
DATE3
2026 HDStream: An Energy-efficient 7.98 TBOPS/W Hyperdimensional Computing Streaming Processor
abstract
Binary hyperdimensional computing (HDC) is a brain-inspired framework that enables energy-efficient classification through simple bitwise operations on high-dimensional binary vectors. Existing accelerators face a fundamental compute-efficiency-density gap: encoding-specific designs achieve high efficiency at the cost of flexibility, while programmable processors sacrifice area and energy efficiency for generality. This work presents HDStream, a streaming HDC processor that closes this gap through (1) a wide-vector microarchitecture with HDC-customized multi-operation instructions and (2) autonomous streaming modules with hardware instruction loops. HDStream achieves up to 5.85 × speedup over single-operation-per-cycle execution with 99% compute utilization. Fabricated in 16 nm CMOS, HDStream achieves a peak 0.870 TBOPS and 7.98 TBOPS/W, with an effective 0.637 TBOPS and 7.65 TBOPS/W across diverse HDC workloads. Compared to prior programmable HDC accelerators, HDStream delivers up to 3.18 × higher compute performance density (TBOPS/mm2) and up to 2.99 × higher energy density (TBOPS/W/mm2). This work demonstrates that encoding flexibility and silicon area efficiency are not mutually exclusive.
Ryan Antonio, Xiaoling Yi, Yunhao Deng, Fanchen Kong, Jun Yin 0001, Marian Verhelst
ACM Great Lakes Symposium on VLSI2
2026 A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
Xiaoling Yi, Ryan Antonio, Yunhao Deng, Fanchen Kong, Joren Dumoulin, Jun Yin 0001, Marian Verhelst
ISCAS1
2026 An Analytical Model for Performance-Carbon Co-Optimization of Edge AI Accelerators
abstract
ASIC accelerators have emerged as vital solutions to deploy various AI workloads, especially in resource-constrained edge scenarios. Existing analytical models for guiding the architectural exploration of these accelerators focus solely on performance and energy efficiency metrics while overlooking the carbon cost. In contrast, existing carbon cost models focus on analyzing specific hardware configurations but lack the support for exploration of the optimal architecture choices and the trade-off analysis between performance and carbon cost. To fill this gap, this paper aims to analyze the impact of different architecture configurations from both the performance and carbon perspectives, exploring the trade-off between the performance and carbon cost for AI accelerator design. For this purpose, we first built an analytical model, namedCarbonSpot, capable of modeling and estimating both performance and carbon cost for any accelerator architecture in the design space. Then, by benchmarking the overhead of these AI accelerators under MLPerf-Tiny and MLPerf-Mobile workloads, we show that architectures solely optimized for performance and energy efficiency produce 58× more carbon emissions than designs designed for the highest carbon efficiency. Importantly, co-optimized architectural choices exist, with only <20% drop in performance and <6% overhead in carbon costs when compared to the respective best cases optimized for either maximum performance or minimal carbon impact. The model is open-sourced at: https://github.com/KULeuven-MICAS/carbonspot.
Jiacong Sun, Xiaoling Yi, Arne Symons, Georges Gielen, Lieven Eeckhout, Marian Verhelst
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 FlexiGen: An Automated AI Accelerator Generation Framework With Decoupled-Access-Execute and Dynamic Dataflows
abstract
Modern tensor applications, especially artificial intelligence (AI) applications, are evolving rapidly, posing a significant demand for agile hardware design. While numerous hardware generators have been developed, they suffer from three significant limitations: 1) they are either limited to a single dataflow/data type generation, failing to cater to the computational requirements of diverse workloads; 2) or focus only on the array level optimization, omitting system-level effects, such as the influence of on-chip memory bandwidth and contention; and 3) customized workload mapping/configuration is needed, resulting in increased programming complexity. To address these challenges, we proposeFlexiGen, a flexible and extensible hardware generation framework, which targets diverse deep neural networks (DNN) tensor applications and can generate a complete synthesizable acceleration system at the RTL level with arbitrary dataflow and its combinations. Our key contributions are threefold: 1) we incorporate decoupled-access-execute architecture insideFlexiGen, enabling full system generation while maintaining flexibility and efficiency; 2) we propose a versatile spatial core generator that supports dynamic spatial dataflows and multiple data precisions in the same array and a compatible data streaming engine generator that can support arbitrary temporal dataflows and$N$-dimensional data access patterns; and 3) we leverage a uniform programming interface and provide a customized kernel library, enabling agile configuration programming. We conduct an intensive evaluation to demonstrate the versatility ofFlexiGenin dataflow accelerator generation and show the trade-offs of performance, area, and power across a wide range of dataflows and workloads at both the array level and system level. Our case study experiment showsFlexiGen’s usefulness as a hardware generator to rapidly generate desired dataflow acceleration systems. Compared with the state-of-the-art (SotA) hardware generation framework LEGO,FlexiGenachieves 36.79% and 57.16% less area and power when generating the same dual spatial dataflow design.FlexiGenis open-source and available athttps://github.com/KULeuven-MICAS/snax_cluster
Xiaoling Yi, Man Shi, Joren Dumoulin, Jiacong Sun, Yunhao Deng, Ryan Antonio, Fanchen Kong, Marian Verhelst
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 OpenGeMM: A Highly-Efficient GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
abstract
Deep neural networks (DNNs) face significant challenges when deployed on resource-constrained extreme edge devices due to their computational and data-intensive nature. While standalone accelerators tailored for specific application scenarios suffer from inflexible control and limited programmability, generic hardware acceleration platforms coupled with RISC-V CPUs can enable high reusability and flexibility, yet typically at the expense of system-level efficiency and low utilization.
Xiaoling Yi, Ryan Antonio, Joren Dumoulin, Jiacong Sun, Josse Van Delm, Guilherme Paim, Marian Verhelst
ASP-DAC1
2025 DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
abstract
Deep Neural Networks (DNNs) have achieved remarkable success across various intelligent tasks but encounter performance and energy challenges in inference execution due to data movement bottlenecks. We introduce DataMaestro, a versatile and efficient data streaming unit that brings the decoupled access/execute architecture to DNN dataflow accelerators to address this issue. DataMaestro supports flexible and programmable access patterns to accommodate diverse workload types and dataflows, incorporates fine-grained prefetch and addressing mode switching to mitigate bank conflicts, and enables customizable on-the-fly data manipulation to reduce memory footprints and access counts. We integrate five DataMaestros with a Tensor Core-like GeMM accelerator and a Quantization accelerator into a RISC-V host system for evaluation. The FPGA prototype and VLSI synthesis results demonstrate that DataMaestro helps the GeMM core achieve nearly 100% utilization, which is 1.05 $21.39 \times$ better than state-of-the-art solutions, while minimizing area and energy consumption to merely 6.43% and 15.06% of the total system.
Xiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong, Guilherme Paim, Marian Verhelst
DAC1
2025 XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
abstract
As modern AI workloads increasingly rely on heterogeneous accelerators, ensuring high-bandwidth and layout-flexible data movements between accelerator memories has become a pressing challenge. Direct Memory Access (DMA) engines promise high bandwidth utilization for data movements but are typically optimal only for contiguous memory access, thus requiring additional software loops for data layout transformations. This, in turn, leads to excessive control overhead and underutilized on-chip interconnects. To overcome this inefficiency, we present XDMA, a distributed and extensible DMA architecture that enables layout-flexible data movements with high link utilization. We introduce three key innovations: (1) a data streaming engine as XDMA Frontend, replacing software address generators with hardware ones; (2) a distributed DMA architecture that maximizes link utilization and separates configuration from data transfer; (3) flexible plugins for XDMA enabling on-the-fly data manipulation during data transfers. XDMA demonstrates up to$151.2 \times / 8.2 \times$higher link utilization than software-based implementations in synthetic workloads and achieves$2.3 \times$average speedup over accelerators with SoTA DMA in real-world applications. Our design incurs$<2 \%$area overhead over SoTA DMA solutions while consuming 17% of system power. XDMA proves that co-optimizing memory access, layout transformation, and interconnect protocols is key to unlocking heterogeneous multi-accelerator SoC performance.
Fanchen Kong, Yunhao Deng, Xiaoling Yi, Ryan Antonio, Marian Verhelst
ICCD3
2025 An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
abstract
Heterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movement efficiency and encounter compatibility issues in hardware and software. This prevents a unified approach that balances performance and ease of use. To this end, we present SNAX, an open-source integrated HW-SW framework enabling efficient multi-accelerator platforms through a novel hybrid-coupling scheme, consisting of loosely coupled asynchronous control and tightly coupled data access. SNAX brings reusable hardware modules designed to enhance compute accelerator utilization, and its customizable MLIR-based compiler to automate key system management tasks, jointly enabling rapid development and deployment of customized multi-accelerator compute clusters. Through extensive experimentation, we demonstrate SNAX’s efficiency and flexibility in a low-power heterogeneous SoC. Accelerators can be easily integrated and programmed to achieve >10× improvement in neural network performance compared to other accelerator systems while maintaining accelerator utilization of >90% in full system operation.
Ryan Antonio, Joren Dumoulin, Xiaoling Yi, Josse Van Delm, Yunhao Deng, Guilherme Paim, Marian Verhelst
ISLPED3
2025 Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
abstract
Autonomous robots require efficient on-device learning to adapt to new environments without cloud dependency. For this edge training, Microscaling (MX) data types offer a promising solution by combining integer and floating-point representations with shared exponents, reducing energy consumption while maintaining accuracy. However, the state-of-the-art continuous learning processor, namely Dacapo, faces limitations with its MXINT-only support and inefficient vector-based grouping during backpropagation. In this paper, we present, to the best of our knowledge, the first work that addresses these limitations with two key innovations: (1) a precision-scalable arithmetic unit that supports all six MX data types by exploiting sub-word parallelism and unified integer and floating-point processing; and (2) support for square shared exponent groups to enable efficient weight handling during backpropagation, removing storage redundancy and quantization overhead.We evaluate our design against Dacapo under iso-peak-throughput on four robotics workloads in TSMC 16nm FinFET technology at 400MHz, reaching a 51% lower memory footprint, and 4× higher effective training throughput, while achieving comparable energy efficiency, enabling efficient robotics continual learning at the edge.
Stef Cuyckens, Xiaoling Yi, Nitish Satya Murthy, Chao Fang 0005, Marian Verhelst
ISLPED2
2025 Prior-Boosted GRL: Microarchitecture Design Space Exploration via Graph Representation Learning
abstract
The design space exploration (DSE) of contemporary microprocessors faces a significant challenge of high-computational cost. In this context, we introduce Prior-boosted graph representation learning (GRL), a novel framework for the DSE of the microarchitectures the microprocessors underpinned by graph embeddings. Using GRL, Prior-boosted GRL constructs a compact and continuous vector space for design representation. This framework is further boosted by an efficient sampling algorithm informed by prior knowledge, which is instrumental in generating a superior set of initial designs to accelerate the exploration process. A well-designed ensemble surrogate model is combined with the multiobjective Bayesian optimization to explore the design space holistically within this graph-embedding domain. Rigorous experimental evaluations conducted on the RISC-V Berkeley-Out-of-Order Machine (BOOM) platform demonstrate that Prior-boosted GRL substantially surpasses preceding methods, achieving a 107.79% enhancement in Pareto front quality compared to the state-of-the-art DSE algorithm. It also outstrips manual designs on performance, power, and area metrics. As of this writing, Prior-boosted GRL holds the first place in the ICCAD 2022 CAD Contest evaluation platform.
Jinyi Shen, Xiaoling Yi, Fan Yang 0001, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 SenseDSE: Sensitivity-Based Performance Evaluation for Design Space Exploration of Microarchitecture
abstract
The design of modern processors is driven by plenty of benchmarks. As processors evolve and applications expand, the complexity of benchmark programs grows, which increases the computational cost of architecture design space exploration (DSE). To accelerate performance evaluations of processors in DSE, we developed a sensitivity-based framework for performance evaluation of a large set of benchmarks. The framework avoids simulating the insensitive benchmarks to the adjusted parameters during the exploration of designs. We developed a sampling algorithm based on evolutionary strategies to provide learning data for the sensitivity analysis and enhance the performance of the fast performance evaluation algorithm. We integrated this framework into a RISe-V processor architecture exploration framework. Our experiments revealed that we could achieve a significant acceleration in runtime with negligible accuracy loss in DSE.
Xiaoling Yi, Fan Yang 0001
DATE2
2023 Graph Representation Learning for Microarchitecture Design Space Exploration
abstract
Design optimization of modern microprocessors is a complex task due to the exponential growth of the design space. This work presents GRL-DSE, an automatic microarchitecture search framework based on graph embeddings. GRL-DSE uses graph representation learning to build a compact and continuous embedding space. Multi-objective Bayesian optimization using an ensemble surrogate model conducts microarchitecture design space exploration in the graph embedding space to efficiently and holistically optimize performance-power-area (PPA) objectives. Experimental studies on RISC-V BOOM show that GRLDSE outperforms previous techniques by 74.59% on Pareto front quality and outperforms manual designs in terms of PPA.
Xiaoling Yi, Jialin Lu, Xiankui Xiong, Dong Xu 0015, Fan Yang 0001
DAC1
2023 TPNoC: An Efficient Topology Reconfigurable NoC Generator
abstract
With the core count increasing in Chip to support various data-intensive workloads, Network-on-chip (NoC) has become the better solution for addressing on-chip interconnection. Various data-intensive workloads have different traffic patterns that require NoC with different topologies and microarchitectures. On the one hand, topology type selection has a great influence on the final performance, area, and energy. However, it is difficult to change the topology type in the traditional NoC RTL design process once it is determined. On the other hand, NoC platforms have many tunable micro-architecture design parameters, which require careful design space exploration to trade off performance advantages and overhead. Designing and validating each microarchitecture of NoCs to account for various trade-offs will greatly exacerbate the design cost issue.
Jiangnan Yu, Fan Yang 0001, Xiaoling Yi, Chixiao Chen, Jun Tao 0001, Dong Xu 0015, Xiankui Xiong
ACM Great Lakes Symposium on VLSI3
2022 An Automated Compiler for RISC-V Based DNN Accelerator
abstract
Multifarious hardware accelerators are developed for the widely used Deep neural networks (DNN). Nowadays the SoCs composed of a general processor and a coupled accelerator are becoming prevalent. Compared to the specialized DNN accelerator for one specific DNN, this kind of coupled architecture is programmable and supports diverse DNNs. However, for the low-level programming interface of the co-processor-like accelerator and the multi-hierarchy memory structure, programming for the DNN accelerator is not easy work. Meanwhile, there are a couple of tensor compilers that deploy the DNN on various hardware. In this work, we combine the flexibility of the tensor compiler and the high efficiency of the hardware accelerator by proposing an automated compiler that can compile tensor programs and generate high-performance programs for programmable DNN accelerators. Our compiler is based on TVM [1] and target at Rocket Chip Coprocessor (RoCC) [2]. The compiler is flexible and supports many kinds of RISC-V instructions. The programmer can define the hardware constraints in the proposed compiler which makes the generated code more efficient. Our compiler can lower the program with the ping-pong strategy and the generated code can achieve 26% speed up compared to the baseline.
Wuzhen Xie, Xiaoling Yi, Ruiyao Pu, Xiankui Xiong, Haidong Yao, Chixiao Chen, Jun Tao 0001, Fan Yang 0001
ISCAS3
2022 NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area Models
abstract
In this paper, we propose NNASIM, an efficient timing and area accurate event-driven simulator for custom DNN accelerators. NNASIM is a highly-modular and highly parameterized modeling framework. We build accurate timing and area models for common accelerator modules like GEMM, ALU array, and crossbar using ASIC synthesis flows. These models are fed into the event-driven simulator for fast simulation. NNASIM is integrated with a RISC-V simulator. This approach guarantees the functional correctness of the accelerator simulation at the instruction level. The experimental results show that our model evaluates the performance and area of DNN accelerators with less than 0.76% and 2.83% error, respectively, compared to RTL implementations. NNASIM allows designers to model the performance and area of the accelerator at a high level, and thus enables the systematic microarchitecture design space exploration of the custom accelerators. Index Terms accelerators.
Xiaoling Yi, Jiangnan Yu, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Jun Tao 0001, Fan Yang 0001
ISCAS1