EDBT 2026 Demo / reviewers in the wild / expert
Guowei Yang 0005
dblp:43/3366-5
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0005-7991-401XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCARLET: A Scalable OPCM-Based Accelerator for Transformer Inference with Tiled CrossbarsabstractWhile transformer-based large language models (LLMs) have achieved state-of-the-art performance on a wide range of natural language processing tasks, their massive computational demands, especially during inference, pose a significant challenge. Photonic accelerators offer a promising solution, but existing designs struggle with the precision, dynamism, and storage requirements of modern LLMs. This paper introduces SCARLET, a hybrid photonic architecture that addresses these limitations through two key components. First, we design a high-density optical phase-change memory (OPCM) crossbar for static matrix multiplications, achieving 5.6× higher bit density and 86.43% lower energy compared to previous OPCM crossbar designs. Second, we introduce an approximate photonic floating-point multiplier to handle dynamic matrix multiplications and quantization steps by approximating floating-point computations with weighted integer sums, thus, eliminating the need for frequent memory reprogramming. Our evaluation on models with up to 13 billion parameters demonstrates significant performance improvements, including up to 17.15× and 8.45× lower latency during prefill and generation phases, respectively. Sina Karimi, Guowei Yang 0005, Carlos A. Ríos Ocampo, Ajay Joshi, Ayse K. Coskun |
DATE | 2 |
| 2024 | Mirage: An RNS-Based Photonic Accelerator for DNN TrainingabstractPhotonic computing is a compelling avenue for performing highly efficient matrix multiplication, a crucial operation in Deep Neural Networks (DNNs). While this method has shown great success in DNN inference, meeting the high precision demands of DNN training proves challenging due to the precision limitations imposed by costly data converters and the analog noise inherent in photonic hardware. This paper proposes Mirage, a photonic DNN training accelerator that overcomes the precision challenges in photonic hardware using the Residue Number System (RNS). RNS is a numeral system based on modular arithmetic-allowing us to perform high-precision operations via multiple low-precision modular operations. In this work, we present a novel micro-architecture and dataflow for an RNS-based photonic tensor core performing modular arithmetic in the analog domain. By combining RNS and photonics, Mirage provides high energy efficiency without compromising precision and can successfully train state-of-the-art DNNs achieving accuracy comparable to FP32 training. Our study shows that on average across several DNNs when compared to systolic arrays, Mirage achieves more than $23.8 \times$ faster training and $32.1 \times$ lower EDP in an iso-energy scenario and consumes $42.8 \times$ lower power with comparable or better EDP in an iso-area scenario. Cansu Demirkiran, Guowei Yang 0005, Darius Bunandar, Ajay Joshi |
ISCA | 2 |
| 2024 | SOPHIE: A Scalable Recurrent Ising Machine Using Optically Addressed Phase Change MemoryabstractIsing problems are nondeterministic-polynomial-hard (NP-hard) problems prevalent in various domains, such as statistical physics, circuit design, and machine learning. They pose significant challenges for traditional algorithms and architectures. Researchers have recently developed nature-inspired Ising machines to tackle these optimization problems efficiently. Many optimization problems can be mapped to the Ising model, and physical laws will drive the Ising machine towards the solution. However, existing Ising machines suffer from scalability issues, i.e., performance drops when problem sizes exceed their physical capacity. In this paper, we propose SOPHIE, a Scalable Optical PHase-change memory (OPCM) based Ising Engine. SOPHIE integrates architectural, algorithmic, and device optimizations to address scalability challenges in Ising machines. We architect SOPHIE using 2.5D integration, where we integrate a controller chiplet, a DRAM chiplet, laser sources, and multiple OPCM chiplets. SOPHIE utilizes OPCMs to perform matrix-vector multiplications efficiently. Our symmetric tile mapping at the architecture level reduces approximately half of the OPCM array area, enhancing the scalability of SOPHIE. We use algorithmic optimizations to efficiently handle large problems that cannot fit within hardware constraints. Specifically, we adopt a symmetric local update technique and a stochastic global synchronization strategy. These two algorithmic approaches decompose large problems into isolated tiles, reduce computation requirements, and minimize communication in SOPHIE. We apply device-level optimizations to adopt the modified algorithm. These device-level optimizations include employing bi-directional OPCM arrays and dual-precision analog-to-digital converters. SOPHIE is 3 x faster than the state-of-the-art photonic Ising machines on small graphs and 125x faster than the FPGA-based designs on large problems. SOPHIE alleviates the hardware capacity constraints, offering a scalable and efficient alternative for solving Ising problems. Guowei Yang 0005, Sina Karimi, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi |
MICRO | 1 |
| 2023 | FAB: An FPGA-based Accelerator for Bootstrappable Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) offers protection to private data on third-party cloud servers by allowing computations on the data in encrypted form. To support general-purpose encrypted computations, all existing FHE schemes require an expensive operation known as "bootstrapping". Unfortunately, the computation cost and the memory bandwidth required for bootstrapping add significant overhead to FHE-based computations, limiting the practical use of FHE.In this work, we propose FAB, an FPGA-based accelerator for bootstrappable FHE. Prior FPGA-based FHE accelerators have proposed hardware acceleration of basic FHE primitives for impractical parameter sets without support for bootstrapping. FAB, for the first time ever, accelerates bootstrapping (along with basic FHE primitives) on an FPGA for a secure and practical parameter set. The key contribution of this work is the architecture of a balanced FAB design, which is not memory bound. In our design, we leverage recent algorithms for bootstrapping while being cognizant of the compute and memory constraints of our FPGA. In addition, we use a minimal number of functional units for computing, operate at a low frequency, leverage high data rates to and from main memory, utilize the limited on-chip memory effectively, and perform careful operation scheduling.We evaluate FAB using a single Xilinx Alveo U280 FPGA and by scaling it to a multi-FPGA system consisting of eight such FPGAs. For bootstrapping a fully-packed ciphertext, while operating at 300MHz, FAB outperforms existing state-of-the-art CPU and GPU implementations by 213× and 1.5× respectively. Our target FHE application is training a logistic regression model over encrypted data. For logistic regression model training scaled to 8 FPGAs on the cloud, FAB outperforms a CPU and GPU by 456× and 9.5× respectively, providing practical performance at a fraction of the ASIC design cost. Rashmi S. Agrawal 0001, Leo de Castro, Guowei Yang 0005, Chiraag Juvekar, Rabia Tugce Yazicigil, Anantha P. Chandrakasan, Vinod Vaikuntanathan, Ajay Joshi |
HPCA | 3 |
| 2023 | Processing-in-Memory Using Optically-Addressed Phase Change MemoryabstractToday's Deep Neural Network (DNN) inference systems contain hundreds of billions of parameters, resulting in significant latency and energy overheads during inference due to frequent data transfers between compute and memory units. Processing-in-Memory (PiM) has emerged as a viable solution to tackle this problem by avoiding the expensive data movement. PiM approaches based on electrical devices suffer from throughput and energy efficiency issues. In contrast, Optically-addressed Phase Change Memory (OPCM) operates with light and achieves much higher throughput and energy efficiency compared to its electrical counterparts. This paper introduces a system-level design that takes the OPCM programming overhead into consideration, and identifies that the programming cost dominates the DNN inference on OPCM-based PiM architectures. We explore the design space of this system and identify the most energy-efficient OPCM array size and batch size. We propose a novel thresholding and reordering technique on the weight blocks to further reduce the programming overhead. Combining these optimizations, our approach achieves up to 65.2 × higher throughput than existing photonic accelerators for practical DNN workloads. Guowei Yang 0005, Cansu Demirkiran, Zeynep Ece Kizilates, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi |
ISLPED | 1 |
| 2023 | RISE: RISC-V SoC for En/Decryption Acceleration on the Edge for Homomorphic EncryptionabstractToday, edge devices commonly connect to the cloud to use its storage and computing capabilities. This leads to security and privacy concerns about user data. Homomorphic encryption (HE) is a promising solution to address the data privacy problem as it allows arbitrarily complex computations on encrypted data without ever needing to decrypt it. While there has been a lot of work on accelerating HE computations in the cloud, small attention has been paid to the message-to-ciphertext and ciphertext-to-message conversion operations on the edge. In this work, we profile the edge-side conversion operations, and our analysis shows that during conversion error sampling, encryption and decryption operations are the bottlenecks. To overcome these bottlenecks, we present RISE, an area and energy-efficient RISC-V system-on-chip (SoC). RISE leverages an efficient and lightweight pseudorandom number generator (PRNG) core and combines it with fast sampling techniques to accelerate the error sampling operations. To accelerate the encryption and decryption operations, RISE uses scalable data-level parallelism to implement the number theoretic transform (NTT) operation, the main bottleneck within the encryption and decryption operations. In addition, RISE saves area by implementing a unified en/decryption datapath, and efficiently exploits techniques like memory reuse and data reordering to utilize a minimal amount of ON-chip memory. We evaluate RISE using a complete RTL design containing a RISC-V processor interfaced with our accelerator. Our analysis reveals that for message-to-ciphertext and ciphertext-to-message conversions, using RISE leads up to$5986.99\times $and$1164.1\times $more energy-efficient solution, respectively, than when using just the RISC-V processor. Zahra Azad, Guowei Yang 0005, Rashmi S. Agrawal 0001, Daniel Ruelas-Petrisko, Michael B. Taylor, Ajay Joshi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | RACE: RISC-V SoC for En/decryption Acceleration on the Edge for Homomorphic ComputationabstractAs more and more edge devices connect to the cloud to use its storage and compute capabilities, they bring in security and data privacy concerns. Homomorphic Encryption (HE) is a promising solution to maintain data privacy by enabling computations on the encrypted user data in the cloud. While there has been a lot of work on accelerating HE computation in the cloud, little attention has been paid to optimize the en/decryption on the edge. Therefore, in this paper, we present RACE, a custom-designed area- and energy-efficient SoC for en/decryption of data for HE. Owing to similar operations in en/decryption, RACE unifies the en/decryption datapath to save area. RACE efficiently exploits techniques like memory reuse and data reordering to utilize minimal amount of on-chip memory. We evaluate RACE using a complete RTL design containing a RISC-V processor and our unified accelerator. Our analysis shows that, for the end-to-end en/decryption, using RACE leads to, on average, 48 × to 39729 × (for a wide range of security parameters) more energy-efficient solution than purely using a processor. Zahra Azad, Guowei Yang 0005, Rashmi S. Agrawal 0001, Daniel Ruelas-Petrisko, Michael B. Taylor, Ajay Joshi |
ISLPED | 2 |