EDBT 2026 Demo / reviewers in the wild / expert
Paul Xuanyuanliang Huang
dblp:332/1853
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0001-8650-1222ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 62% High-performance computing · 19% Memory systems · 19% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
1.0 | 1 | 2026 | On-the-Fly Memory Decompression in GPU Featuring In-Partition Compressed Block Prefetching and Gzip Decompression for Scientific Computing · IEEE Trans. Computers 2026 |
Memory systems
memory-bound computation |
0.3 | 1 | 2026 | On-the-Fly Memory Decompression in GPU Featuring In-Partition Compressed Block Prefetching and Gzip Decompression for Scientific Computing · IEEE Trans. Computers 2026 |
High-performance computing
scientific computing |
0.3 | 1 | 2026 | On-the-Fly Memory Decompression in GPU Featuring In-Partition Compressed Block Prefetching and Gzip Decompression for Scientific Computing · IEEE Trans. Computers 2026 |
Methods — techniques the papers use, named apart from their topics
gzip decompression · 1.0compressed block prefetching · 1.0LLC prefetching · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On-the-Fly Memory Decompression in GPU Featuring In-Partition Compressed Block Prefetching and Gzip Decompression for Scientific ComputingabstractIn modern graphics processing units (GPUs), scientific computing workloads are memory-bound due to the everincreasing data size and the limited off-chip memory bandwidth. The problem still remains with the adoption of high-bandwidth memory (HBM). To further increase memory bandwidth, researchers have proposed on-the-fly memory decompression (OTFMD) techniques that aim to store compressed data off-chip while decompressing the data on-chip. However, for scientific computing applications, a critical challenge of prior approaches is their low compression ratios, as the existing works target integer applications and thus employ simple compression algorithms. Employing a compression algorithm with a higher compression ratio, such as Gzip, for floating-point applications could be promising; however, it is not straightforward to adopt such an approach in the OTFMD scheme. For example, Gzip’s compressed block size is much larger than the GPU’s cache line. Additionally, after global-to-local memory mapping, the data in a Gzip block can be physically distributed across multiple DRAM channels, which complicates the decompression process. In this work, to address those challenges, we propose a Gzip-based OTFMD technique for GPUs for scientific computing.We propose an in-partition compressed block prefetching scheme, where data is compressed within each memory partition separately. Then, upon a request, we prefetch the compressed data into the last-level cache (LLC) while fully preserving memory-level parallelism. We also propose a new Gzip decompression accelerator that offers 5.4× higher throughput per unit silicon area than the state-of-the-art design. Lastly, we propose modifications to the GPU’s microarchitecture to fully leverage the proposed OTFMD technique. Our experimental evaluations show average performance improvements of 38% (up to 63%) for CuBLAS and 15% (up to 22%) for CuSPARSE applications compared to the baseline GPU model. Paul Xuanyuanliang Huang, Mingoo Seok |
IEEE Trans. Computers | 1 |
| 2025 | A 4.2-to-0.5-V, 0.8-μA-0.8-mA, Power-Efficient Three-Level SIMO Buck Converter for a Quad-Voltage RISC-V MicroprocessorabstractThis article presents a Li-ion battery-compatible single-inductor-multiple-output (SIMO) buck converter that fulfills the power management need of an integrated sub-mW RISC-V microprocessor. The proposed converter can directly take a 4.2-V battery voltage and produce four power rails ranging from 1.8 V for I/O to 0.5 V for the processor core. The three-level input stage is chosen to reduce the inductor ripple size and switching loss, thus increasing power conversion efficiency (PCE). In addition, the fully digital implementation using novel domino flash analog-digital converters (ADCs) enables low static current. Also, pulse frequency modulation (PFM) results in a wide dynamic range. The proposed three-level SIMO converter has been prototyped in a 65-nm CMOS technology with the 32-bit RISC-V processor. Measurement results show that the converter achieves a$1000\times $load current range ($0.8~\mu $A–0.8 mA) to support the active or sleep modes of the processor. The converter marks the PCE of 56.2%–72.8%. Compared to the ideal buck-low-dropout voltage regulator (LDO) architecture (LDO-only), it improves the PCE by 23.8% (46.4%). Dongkwun Kim, Zhaoqing Wang, Paul Xuanyuanliang Huang, Pavan Kumar Chundi, Suhwan Kim 0002, Andres A. Blanco, Ram Krishnamurthy 0001, Mingoo Seok |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | SPADES: A 0.54-GFLOPS/W Sparse Matrix Vector Multiplication Accelerator Featuring On-the-Fly GZIP Decompression for 3.36X Reduction in Off-Chip Data MovementabstractWe propose SPADES, the first sparse matrix vector multiplication (SpMV) accelerator incorporating online GZIP decompression. The goal of the decompression is to reduce off-chip data movement, which has become a major source of energy consumption in SpMV accelerators. Our proposed GZIP-SpMV flow achieves an average compression ratio of 3.36 for the sparse matrix CSC data, reducing the off-chip data traffic by 70%. We fabricated SPADES in TSMC 28-nm with a die area of 2.47 mm2. Compared with the prior SpMV accelerator, SPADES achieves 2.32X better energy efficiency. Paul Xuanyuanliang Huang, Yannis P. Tsividis, Mingoo Seok |
ISLPED | 1 |