EDBT 2026 Demo / reviewers in the wild / expert
Di Wu 0016
dblp:52/328-16
· DBLP profile ↗
18ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0001-9775-8026ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 9 first-author · 12 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mugi: Value Level Parallelism For Efficient LLMsabstractValue level parallelism (VLP) has been proposed to improve the efficiency of large-batch, low-precision general matrix multiply (GEMM) between symmetric activations and weights. In transformer based large language models (LLMs), there exist more sophisticated operations beyond activation-weight GEMM. In this paper, we explore how VLP benefits LLMs. First, we generalize VLP for nonlinear approximations, outperforming existing nonlinear approximations in end-to-end LLM accuracy, performance, and efficiency. Our VLP approximation follows a value-centric approach, where important values are assigned with greater accuracy. Second, we optimize VLP for small-batch GEMMs with asymmetric inputs efficiently, which leverages timely LLM optimizations, including weight-only quantization, key-value (KV) cache quantization, and group query attention. Finally, we design a new VLP architecture, Mugi, to encapsulate the innovations above and support full LLM workloads, while providing better performance, efficiency and sustainability. Our experimental results show that Mugi can offer significant improvements on throughput and energy efficiency, up to $45\times$ and $668\times$ for nonlinear softmax operations, and $2.07\times$ and $3.11\times$ for LLMs, and also decrease operational carbon for LLM operation by $1.45\times$ and embodied carbon by $1.48\times$. Daniel Price, Prabhu Vellaisamy, John Paul Shen, Di Wu 0016 |
ASPLOS (2) | 4 |
| 2026 | Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUsabstractGPU systems are increasingly powering modern datacenters at scale. Despite being highly performant, GPU systems can exhibit performance variation at the node and cluster levels. Such performance variation can significantly impact both high-performance computing and artificial intelligence workloads, such as cutting-edge large language models (LLMs). In this work, we analyze the performance of a single-node multi-GPU system running LLM training, and observe that the kernel-level performance variation is highly correlated with concurrent computation and communication (C3), a technique to overlap computation and communication across GPUs for performance gains. We then take a further step to reason that thermally induced straggling coupled with C3 impacts performance variation, which we coin the Lit Silicon effect. More specifically, Lit Silicon describes that in a multi-GPU node, thermal imbalance across GPUs can introduce node-level straggler GPUs (hotter and slower), which in turn slow down the leader GPUs (cooler and faster). Lit Silicon can lead to node-level performance variation and inefficiency, potentially impacting the entire datacenter. We propose analytical performance and power models for Lit Silicon, to understand the potential system-level gains. We further design simple detection and mitigation techniques to effectively address the Lit Silicon problem, and evaluate three different power management solutions, including (1) power optimization under GPU thermal design power, (2) performance optimization under node-level GPU power capping, and (3) performance optimization under node-level CPU power sloshing. We conduct experiments on two workloads on two AMD InstinctTM MI300X GPU systems under two LLM training frameworks, and observe up to 6% performance and 4% power improvements, potentially saving several tens of millions of dollars in electricity costs in datacenters. Marco Kurzynski, Shaizeen Aga, Di Wu 0016 |
ISCA | 3 |
| 2026 | CryptOracle: A Modular Framework to Characterize FHEabstractPrivacy-preserving machine learning has become an important long-term pursuit in this era of artificial intelligence (AI). Fully Homomorphic Encryption (FHE) is a uniquely promising solution, offering provable privacy and security guarantees. Unfortunately, computational cost is impeding its mass adoption. Modern solutions are up to six orders of magnitude slower than plaintext execution. Understanding and reducing this overhead is essential to the advancement of FHE, particularly as the underlying algorithms evolve rapidly. This paper presents a detailed characterization of OpenFHE, a comprehensive open-source library for FHE, with a particular focus on the CKKS scheme due to its significant potential for AI and machine learning applications. We introduce CryptOracle, a modular evaluation framework comprising (1) a benchmark suite, (2) a hardware profiler, and (3) a predictive performance model. The benchmark suite encompasses OpenFHE kernels at three abstraction levels: workloads, microbenchmarks, and primitives. The profiler is compatible with standard and user-specified security parameters. CryptOracle monitors application performance, captures microarchitectural events, and logs power and energy usage for AMD and Intel systems. These metrics are consumed by a modeling engine to estimate runtime and energy efficiency across different configuration scenarios, with prediction error ranging from $-7.02 \% \sim 8.40 \%$ for runtime and $-9.74 \% \sim 15.67 \%$ for energy (geomean). CryptOracle is open source, fully modular, and serves as a shared platform to facilitate the collaborative advancements of applications, algorithms, software, and hardware in FHE. The CryptOracle code can be accessed at https://github.com/UnaryLab/CryptOracle. Cory Brynds, Parker McLeod, Lauren Caccamise, Asmita Pal, Dewan Saiham, Sazadur Rahman, Joshua San Miguel, Di Wu 0016 |
ISPASS | 8 |
| 2025 | PIM-SUM: Fast and Reliable In-Memory Summation for Recommendation SystemsabstractEmbedding aggregation in large-scale recommendation systems creates severe memory bandwidth bottlenecks, as each query sums many high-dimensional vectors. Bitwise-operation-based PIM can exploit subarray bandwidth, but traditional designs struggle with summation because long carry propagation limits parallelism. We propose PIM-SUM, an in-DRAM summation primitive for Sparse Length Sum (SLS) in recommendation workloads. PIM-SUM reformulates vector summation as column-wise accumulation via popcount, truncating carry propagation and avoiding redundant in-DRAM computation. This enables highthroughput summation using native DRAM bitwise primitives. PIM-SUM integrates a Reed-Solomon-inspired error correction at the DRAM row level. Its linearity supports parity propagation, enabling integrity checking and multi-bit error correction with low overhead. On DLRM workloads, PIM-SUM achieves up to$5.14 \times$speedup, logarithmic I/O reduction, and large energy savings. It also reduces silent data corruption by$1778 \times$and improves detection by over$170 \times$. These results show PIM-SUM is a scalable, fault-tolerant, and energy-efficient solution for memory-bound inference at data center scale. Ruizhi Zhu, Huize Li, Di Wu 0016, Xin Xin 0008 |
ICCD | 4 |
| 2025 | Can Photonic Interconnects be used for High-Throughput Memory Access in FHE Accelerators?abstractFully Homomorphic Encryption (FHE) allows computations over encrypted data without sacrificing confidentiality, but its practicality is hindered by high computational demands and memory access constraints. While existing FHE accelerators focus on improving computational efficiency, they are often limited by the insufficient memory bandwidth and inefficient data transfer schemes, leading to significant bottlenecks, especially for processing large amounts of data. In this work, we evaluate whether OptoLink, a photonic interconnect architecture, is scalable and capable of providing high bandwidth to overcome these limitations. Leveraging Wavelength Division Multiplexing (WDM) with Space Division Multiplexing (SDM), OptoLink achieves an impressive bandwidth of 1.6 TB/s over 128 channels—a 300x improvement over traditional electronic network. Additionally, its ability to efficiently broadcast data and support parallel processing further enhances performance. The broadcasting capability not only enables parallelism but also reduces power consumption in earlier NTT stages, improving overall energy efficiency. With its improved data throughput, scalability, and lower latency, OptoLink offers a robust solution capable of satisfying the high data transfer and memory demands of current FHE accelerators. Dewan Saiham, Mariam Rabadi, Di Wu 0016, Sazadur Rahman |
ISLPED | 3 |
| 2024 | Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMsabstractIn recent years, hardware architectures optimized for general matrix multiplication (GEMM) have been well studied to deliver better performance and efficiency for deep neural networks. With trends towards batched, low-precision data, e.g., FP8 format in this work, we observe that there is growing untapped potential for value reuse. We propose a novel computing paradigm, value-level parallelism, whereby unique products are computed only once, and different inputs subscribe to (select) their products via temporal coding. Our architecture, Carat, employs value-level parallelism and transforms multiplication into accumulation, performing GEMMs with efficient multiplier-free hardware. Experiments show that, on average, Carat improves iso-area throughput and energy efficiency by 1.02× and 1.06× over a systolic array and 3.2× and 4.3× when scaled up to multiple nodes. Zhewen Pan 0001, Joshua San Miguel, Di Wu 0016 |
ASPLOS (2) | 3 |
| 2024 | ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV CachingabstractThe Transformer architecture has significantly advanced natural language processing (NLP) and has been foundational in developing large language models (LLMs) such as LLaMA and OPT, which have come to dominate a broad range of NLP tasks. Despite their superior accuracy, LLMs present unique challenges in practical inference, concerning the compute and memory-intensive nature. Thanks to the autoregressive characteristic of LLM inference, KV caching for the attention layers in Transformers can effectively accelerate LLM inference by substituting quadratic-complexity computation with linear-complexity memory accesses. Yet, this approach requires increasing memory as demand grows for processing longer sequences. The overhead leads to reduced throughput due to I/O bottlenecks and even out-of-memory errors, particularly on resource-constrained systems like a single commodity GPU. In this paper, we propose ALISA, a novel algorithm-system co-design solution to address the challenges imposed by KV caching. On the algorithm level, ALISA prioritizes tokens that are most important in generating a new token via a Sparse Window Attention (SWA) algorithm. SWA introduces high sparsity in attention layers and reduces the memory footprint of KV caching at negligible accuracy loss. On the system level, ALISA employs three-phase token-level dynamical scheduling and optimizes the trade-off between caching and recomputation, thus maximizing the overall performance in resource-constrained systems. In a single GPU-CPU system, we demonstrate that under varying workloads, ALISA improves the throughput of baseline systems such as FlexGen and vLLM by up to $3 \times$ and $1.9 \times$, respectively. Youpeng Zhao 0002, Di Wu 0016, Jun Wang 0001 |
ISCA | 2 |
| 2024 | LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have gained significant research attention over the past decade due to their potential for enabling resource-constrained edge devices. While existing SNN accelerators efficiently process sparse spikes with dense weights, the opportunities for accelerating SNNs with sparse weights, referred to as dual-sparsity, remain underexplored. In this work, we focus on accelerating dual-sparse SNNs, particularly on their core operation: sparse-matrix-sparse-matrix multiplication (spMspM). Our observations reveal that executing a dual-sparse SNN on existing spMspM accelerators designed for dual-sparse Artificial Neural Networks (ANNs) results in sub-optimal efficiency. The main challenge is that SNNs, which naturally processes multiple timesteps, introducing an additional loop in ANN spMspM, leading to longer latency and more memory traffic. To address this issue, we propose a fully temporal-parallel (FTP) dataflow that minimizes data movement across timesteps and reduces the end-to-end latency of dual-sparse SNNs. To enhance the efficiency of the FTP dataflow, we introduce an FTP-friendly spike compression mechanism that efficiently compresses single-bit spikes and ensures contiguous memory access. Additionally, we propose an FTP-friendly inner-join circuit that reduces the cost of expensive prefix-sum circuits with negligible throughput penalties. These innovations are encapsulated in LoAS, a Low-latency inference Accelerator for dual-sparse SNNs. Running dual-sparse SNN workloads on LoAS demonstrates significant speedup (up to$8.51 \times$) and energy reduction (up to$3.68\times$) compared to prior dual-sparse accelerators. Ruokai Yin, Youngeun Kim, Di Wu 0016, Priyadarshini Panda |
MICRO | 3 |
| 2022 | uSystolic: Byte-Crawling Unary Systolic ArrayabstractGeneral matrix multiply (GEMM) is an important operation in broad applications, especially the thriving deep neural networks. To achieve low power consumption for GEMM, researchers have already leveraged unary computing, which manipulates bitstreams with extremely simple logic. However, existing unary architectures are not well generalizable to varying GEMM configurations in versatile applications and incompatible to the binary computing stack, imposing challenges to execute unary GEMM effortlessly. In this work, we address the problem by architecting a hybrid unary-binary systolic array, uSystolic, to inherit the legacy-binary data scheduling with slow (thus power-efficient) data movement, i.e., data bytes are crawling out from memory to drive uSystolic. uSystolic exhibits tremendous area and power improvements as a joint effect of 1) low-power computing kernel, 2) spatial-temporal bitstream reuse, and 3) on-chip SRAM elimination. For the evaluated edge computing scenario, compared with the binary parallel design, the rated-coded uSystolic reduces the systolic array area and total on-chip area by 59.0% and 91.3%, with the on-chip energy and power efficiency improved by up to 112.2× and 44.8× for AlexNet. Di Wu 0016, Joshua San Miguel |
HPCA | 1 |
| 2022 | uBrain: a unary brain computer interfaceabstractBrain computer interfaces (BCIs) have been widely adopted to enhance human perception via brain signals with abundant spatial-temporal dynamics, such as electroencephalogram (EEG). In recent years, BCI algorithms are moving from classical feature engineering to emerging deep neural networks (DNNs), allowing to identify the spatial-temporal dynamics with improved accuracy. However, existing BCI architectures are not leveraging such dynamics for hardware efficiency. In this work, we present uBrain, a unary computing BCI architecture for DNN models with cascaded convolutional and recurrent neural networks to achieve high task capability and hardware efficiency. uBrain co-designs the algorithm and hardware: the DNN architecture and the hardware architecture are optimized with customized unary operations and immediate signal processing after sensing, respectively. Experiments show that uBrain, with negligible accuracy loss, surpasses the CPU, systolic array and stochastic computing baselines in on-chip power efficiency by 9.0×, 6.2× and 2.0×. Di Wu 0016, Zhewen Pan 0001, Younghyun Kim 0001, Joshua San Miguel |
ISCA | 1 |
| 2021 | Normalized Stability: A Cross-Level Design Metric for Early Termination in Stochastic ComputingabstractStochastic computing is a statistical computing scheme that represents data as serial bit streams to greatly reduce hardware complexity. The key trade-off is that processing more bits in the streams yields higher computation accuracy at the cost of more latency and energy consumption. To maximize efficiency, it is desirable to account for the error tolerance of applications and terminate stochastic computations early when the result is acceptably accurate. Currently, the stochastic computing community lacks a standard means of measuring a circuit's potential for early termination and predicting at what cycle it would be safe to terminate. To fill this gap, we propose normalized stability, a metric that measures how fast a bit stream converges under a given accuracy budget. Our unit-level experiments show that normalized stability accurately reflects and contrasts the early-termination capabilities of varying stochastic computing units. Furthermore, our application-level experiments on low-density parity-check decoding, machine learning and image processing show that normalized stability can reduce the design space and predict the timing to terminate early. Di Wu 0016, Ruokai Yin, Joshua San Miguel |
ASP-DAC | 1 |
| 2021 | Special Session: When Dataflows Converge: Reconfigurable and Approximate Computing for Emerging Neural NetworksabstractDeep Neural Networks (DNNs) have gained significant attention in both academia and industry due to the superior application-level accuracy. As DNNs rely on compute- or memory-intensive general matrix multiply (GEMM) operations, approximate computing has been widely explored across the computing stack to mitigate the hardware overheads. However, better-performing DNNs are emerging with growing complexity in their use of nonlinear operations, which incurs even more hardware cost. In this work, we address this challenge by proposing a reconfigurable systolic array to execute both GEMM and nonlinear operations via approximation with distinguished dataflows. Experiments demonstrate that such converging of dataflows significantly saves the hardware cost of emerging DNN inference. Di Wu 0016, Joshua San Miguel |
ICCD | 1 |
| 2021 | UNO: Virtualizing and Unifying Nonlinear Operations for Emerging Neural NetworksabstractLinear multiply-accumulate (MAC) operations have been the main focus of prior efforts in improving the energy efficiency of neural network inference due to their dominant contribution to energy consumption in traditional models. On the other hand, nonlinear operations, such as division, exponentiation, and logarithm, that are becoming increasingly significant in emerging neural network models, have been largely underexplored. In this paper, we propose UNO, a low-area, low-energy processing element that virtualizes the Taylor approximation of nonlinear operations on top of off-the-shelf linear MAC units already present in inference hardware. Such virtualization approximates multiple nonlinear operations in a unified, MAC-compatible manner to achieve dynamic run-time accuracy-energy scaling. Compared to the baseline, our scheme reduces the energy consumption by up to 38.4% for individual operations and increases the energy efficiency by up to 274.5% for emerging neural network models with negligible inference loss Di Wu 0016, Setareh Behroozi, Younghyun Kim 0001, Joshua San Miguel |
ISLPED | 1 |
| 2020 | UGEMM: Unary Computing Architecture for GEMM ApplicationsabstractGeneral matrix multiplication (GEMM) is universal in various applications, such as signal processing, machine learning, and computer vision. Conventional GEMM hardware architectures based on binary computing exhibit low area and energy efficiency as they scale due to the spatial nature of number representation and computing. Unary computing, on the other hand, can be performed with extremely simple processing units, often just with a single logic gate. But currently there exist no efficient architectures for unary GEMM. In this paper, we present uGEMM, an area- and energy-efficient unary GEMM architecture enabled by novel arithmetic units. The proposed design relaxes previously-imposed constraints on input bit streams-low correlation and long stream length- and achieves superior area and energy efficiency over existing unary systems. Furthermore, uGEMM's output bit streams exhibit higher accuracy and faster convergence, enabling dynamic energy-accuracy scaling on resource-constrained systems. Di Wu 0016, Ruokai Yin, Hsuan Hsiao, Younghyun Kim 0001, Joshua San Miguel |
ISCA | 1 |
| 2019 | In-Stream Stochastic Division and Square Root via CorrelationabstractStochastic Computing (SC) is designed to minimize hardware area and power consumption compared to traditional binary-encoded computation, stemming from the bit-serial data representation and extremely straightforward logic. Though existing Stochastic Computing Units mostly assume uncorrelated bit streams, recent works find that correlation can be exploited for higher accuracy. We propose novel architectures for SC division and square root, which leverage correlation via low-cost in-stream mechanisms that eliminate expensive bit stream regeneration. We also introduce new metrics to better evaluate SC circuits relying on equilibrium via feedback loops. Experiments indicate that our division converges 46.3% faster with both 43.3% lower error and 45.6% less area. Di Wu 0016, Joshua San Miguel |
DAC | 1 |
| 2019 | SECO: A Scalable Accuracy Approximate Exponential Function Via Cross-Layer OptimizationabstractFrom signal processing to emerging deep neural networks, a range of applications exhibit intrinsic error resilience. For such applications, approximate computing opens up new possibilities for energy-efficient computing by producing slightly inaccurate results using greatly simplified hardware. Adopting this approach, a variety of basic arithmetic units, such as adders and multipliers, have been effectively redesigned to generate approximate results for many error-resilient applications.In this work, we propose SECO, an approximate exponential function unit (EFU). Exponentiation is a key operation in many signal processing applications and more importantly in spiking neuron models, but its energy-efficient implementation has been inadequately explored. We also introduce a cross-layer design method for SECO to optimize the energy-accuracy trade-off. At the algorithm level, SECO offers runtime scaling between energy efficiency and accuracy based on approximate Taylor expansion, where the error is minimized by optimizing parameters using discrete gradient descent at design time. At the circuit level, our error analysis method efficiently explores the design space to select the energy-accuracy-optimal approximate multiplier at design time. In tandem, the cross-layer design and runtime optimization method are able to generate energy-efficient and accurate approximate EFU designs that are up to 99.7% accurate at a power consumption of 3.73 pJ per exponential operation. SECO is also evaluated on the adaptive exponential integrate-and-fire neuron model, yielding only 0.002% timing error and 0.067% value error compared to the precise neuron model. Di Wu 0016, Tian'en Chen, Chien-Fu Chen, Oghenefego Ahia, Joshua San Miguel, Mikko H. Lipasti, Younghyun Kim 0001 |
ISLPED | 1 |
| 2016 | Convergence-optimized variable node structure for stochastic LDPC decoderabstractBy using stochastic computation, a fully-parallel low-density parity-check (LDPC) decoder can be implemented using a lower wire complexity. In order to enhance the decoder performance, probability tracers, such as up/down counters, are added at each edge between variable nodes and check nodes, as described in previous literature. However, this causes a large decoding latency and a high number of decoding failures. In this paper, a convergence-optimized structure for variable nodes is proposed that is able to overcome these issues. As a result, the throughput for the proposed decoder is 20.5Gb/s, which is 101% higher than the original counter-based decoder presented in the previous literature. Qichen Zhang, Yun Chen 0001, Di Wu 0016, Xiaoyang Zeng, Yeong-Luh Ueng |
ICASSP | 3 |
| 2015 | Latency-optimized stochastic LDPC decoder for high-throughput applicationsabstractStochastic decoding can be applied to Low-Density Parity-Check codes in order to achieve high throughput with less area. However, most architectures suffer from large decoding latencies, due to the mechanism of stochastic computation. In this paper, three novel strategies, including the LUT-based initialization, the posterior-information-based hard decision and the Bit-Flipping-based post processing, are proposed in order to reduce decoding latency and hence improve throughput. For the standard IEEE 802.3an (2048, 1723) code, simulation indicates 75.7% reduction in average decoding cycles at 4.5 dB with satisfied bit error rate. Moreover, hardware implementation shows that the area of variable node units is reduced significantly in SMIC 65 nm technology. Di Wu 0016, Yun Chen 0001, Qichen Zhang, Lirong Zheng 0001, Xiaoyang Zeng, Yeong-Luh Ueng |
ISCAS | 1 |