Zhewen Pan 0001

dblp:321/4386-1 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0009-5707-1137ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 The XOR Cache: A Catalyst for Compression
abstract
Modern computing systems allocate significant amounts of resources for caching, especially for the last level cache (LLC).We observe that there is untapped potential for compression by leveraging redundancy due to private caching and inclusion that are common in today's systems.We introduce the XOR Cache to exploit this redundancy via XOR compression.Unlike conventional cache architectures, XOR Cache stores bitwise XOR values of line pairs, halving the number of stored lines via a form of inter-line compression.When combined with other compression schemes, XOR Cache can further boost intra-line compression ratios by XORing lines of similar value, reducing the entropy of the data prior to compression.Evaluation results show that the XOR Cache can save LLC area by 1.93× and power by 1.92× at a cost of 2.06% performance overhead compared to a larger uncompressed cache, reducing energy-delay product by 26.3%.
Zhewen Pan 0001, Joshua San Miguel
ISCA1
2024 Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMs
abstract
In recent years, hardware architectures optimized for general matrix multiplication (GEMM) have been well studied to deliver better performance and efficiency for deep neural networks. With trends towards batched, low-precision data, e.g., FP8 format in this work, we observe that there is growing untapped potential for value reuse. We propose a novel computing paradigm, value-level parallelism, whereby unique products are computed only once, and different inputs subscribe to (select) their products via temporal coding. Our architecture, Carat, employs value-level parallelism and transforms multiplication into accumulation, performing GEMMs with efficient multiplier-free hardware. Experiments show that, on average, Carat improves iso-area throughput and energy efficiency by 1.02× and 1.06× over a systolic array and 3.2× and 4.3× when scaled up to multiple nodes.
Zhewen Pan 0001, Joshua San Miguel, Di Wu 0016
ASPLOS (2)1
2022 uBrain: a unary brain computer interface
abstract
Brain computer interfaces (BCIs) have been widely adopted to enhance human perception via brain signals with abundant spatial-temporal dynamics, such as electroencephalogram (EEG). In recent years, BCI algorithms are moving from classical feature engineering to emerging deep neural networks (DNNs), allowing to identify the spatial-temporal dynamics with improved accuracy. However, existing BCI architectures are not leveraging such dynamics for hardware efficiency. In this work, we present uBrain, a unary computing BCI architecture for DNN models with cascaded convolutional and recurrent neural networks to achieve high task capability and hardware efficiency. uBrain co-designs the algorithm and hardware: the DNN architecture and the hardware architecture are optimized with customized unary operations and immediate signal processing after sensing, respectively. Experiments show that uBrain, with negligible accuracy loss, surpasses the CPU, systolic array and stochastic computing baselines in on-chip power efficiency by 9.0×, 6.2× and 2.0×.
Di Wu 0016, Zhewen Pan 0001, Younghyun Kim 0001, Joshua San Miguel
ISCA3