EDBT 2026 Demo / reviewers in the wild / expert
Xuan Wang 0040
dblp:34/4799-40
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0008-0967-3582ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FHEIns: Fully Homomorphic Encryption Acceleration for Large Data Applications with In-Storage ProcessingabstractRecently, the significance of data privacy protection has been growing rapidly. Homomorphic encryption (HE) enables computation directly on ciphertexts, making it attractive for privacy-sensitive databases in cloud datacenters. Although FHE enables privacy-preserving compute, ciphertext expansion and long-latency primitives drive up memory footprint and delay, worsening compute and memory pressure for database search. In practice, encrypted databases span hundreds of gigabytes to terabytes, making the storage I/O the dominant bottleneck. However, most prior FHE accelerators optimize on-chip computation and the main memory traffic while assuming working sets fit in HBM. Therefore, in this work, we present FHEIns, an in-storage processing architecture that executes FHE kernels close to data inside the NAND flash-based solid-state drives (SSDs) to exploit the internal bandwidth of the SSD. FHEIns achieves up to 24.7× and 2.67× speedup compared to the state-of-the-art FHE ASIC accelerators on trending FHE-based database benchmarks. Xuan Wang 0040, Keming Fan, Augusto Vega, Minxuan Zhou, Tajana Rosing |
DATE | 1 |
| 2026 | NOVA-PIM: Noise-Aware Hyperdimensional Processing in Memory with Optimized Vector Allocation and Minimal ADCsabstractHyperdimensional computing (HDC) is an emerging brain-inspired paradigm that enables highly efficient and robust inference and learning. Analog processing in memory (PIM) has become a promising solution to accelerate HDC by processing lengthy hypervectors (HVs) directly in memory, thereby reducing costly data movement and leveraging massive parallelism. Despite its efficiency, analog PIM suffers from non-idealities that reduce reliability and accuracy. Although the similarity search stage in HDC is inherently error-tolerant given the high dimensionality of HVs, the encoding stage, which transforms raw input data into HVs, remains sensitive to analog noise. Moreover, encoding accounts for a dominant portion of energy consumption, creating a long-standing bottleneck that limits the overall efficiency of analog PIM-based HDC systems. To overcome this challenge, we propose a noise-aware partitioning scheme that improves HDC inference accuracy by processing a critical subset of HV dimensions digitally, while offloading most of the non-critical dimensions to analog PIM. To further synergize the PIM operations across the two consecutive stages, we eliminate the analog-to-digital converters (ADCs) overhead for encoding by employing pulse width modulation (PWM), allowing direct interfacing with the subsequent similarity search stage. The proposed system achieves a 2.6 × reduction in area, 1.5 × –10.3 × lower energy consumption, and 4.2 × –6.5 × speedup compared with state-of-the-art, while maintaining inference accuracy. Keming Fan, Chang Eun Song, Xuan Wang 0040, Tajana Rosing, Mingu Kang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Rhychee-FL: Robust and Efficient Hyperdimensional Federated Learning with Homomorphic Encryption
Yujin Nam, Abhishek Moitra, Yeshwanth Venkatesha, Xiaofan Yu 0001, Gabrielle De Micheli, Xuan Wang 0040, Minxuan Zhou, Augusto Vega, Priyadarshini Panda, Tajana Rosing |
DATE | 6 |
| 2025 | PATHE: A Privacy-Preserving Database Pattern Search Platform with Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables secure computation on encrypted data without decryption, allowing a great opportunity for privacy-preserving computation. Many companies maintain extensive, high-quality databases to deliver services, making preserving data privacy during the database pattern searches crucial. With FHE, the server can take encrypted queries from clients and search through the reference database on the server without decryption, thus guaranteeing data security for all parties. While FHE provides a promising solution to data privacy, it has severe drawbacks of explosive memory requirements and excessive latency, which amplify the computational and memory inefficiencies for database search applications.To address these, we propose PATHE that exploits FHE and hyperdimensional computing (HDC), which provides high parallelism, excellent robustness to errors, for high-performance privacy-preserving database search. On the software side, we propose an FHE-friendly PATHE algorithm that leverages efficient FHE-HDC search and a scheme-switching-based argmax to support database search and maintain comparable accuracy to the state-of-the-art. On the hardware side, PATHE proposes an efficient and scalable FHE accelerator system using Compute Express Link (CXL) for large-scale FHE database search, along with a novel, storage-aware dataflow designed to optimize memory and storage transfers for large database workloads. We evaluate PATHE on the large-scale encrypted database of protein mass spectra, PATHE achieves 2.1× speedup and 1.7× better energy efficiency compared to the baseline system. Xuan Wang 0040, Minxuan Zhou, Gabrielle De Micheli, Yujin Nam, Sumukh Pinge, Augusto Vega, Tajana Rosing |
ICCAD | 1 |
| 2025 | Fast-OverlaPIM: A Fast Overlap-Driven Mapping Framework for Processing In-Memory Neural Network AccelerationabstractProcessing in-memory (PIM) is promising to accelerate neural networks (NNs) because it minimizes data movement and provides large computational parallelism. Similar to machine learning accelerators, application mapping, which determines the operation scheduling and data layout, plays a critical role in the NN acceleration on PIM. The mapping optimization of the previous NN accelerators focused on optimizing the latency of sequential execution. However, PIM accelerators feature a distinct design space of application mapping from conventional NN accelerators, due to the spatial execution of NN layers across different memory locations. This enables opportunities for overlapping execution of consecutive NN layers to improve the latency, where the succeeding layer can start execution before the preceding layer fully completes the computation. In this article, we propose Fast-OverlaPIM framework that incorporates computational overlapping optimization into the deep neural network mapping exploration process on PIM architectures. Fast-OverlaPIM includes analytical algorithms for fast and accurate overlap analysis. Furthermore, it proposes a novel mapping search strategy and a transformation mechanism to enable efficient design space exploration on the overlap-based mapping for the whole network. Our framework demonstrates a significant improvement in runtime performance from$3.4\times $to$323.1\times $compared to the previous state-of-the-art overlap-based framework. Our experiments show that Fast-OverlaPIM can efficiently produce mappings that are$4.6\times $to$18.1\times $faster than the state-of-the-art mapping optimization framework under the same architecture constraints. Xuan Wang 0040, Minxuan Zhou, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | UFC: A Unified Accelerator for Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) is crucial for post-quantum privacy-preserving computing. Researchers have proposed various FHE schemes that excel at different encrypted computations, such as single-instruction multiple-data (SIMD) arithmetic or arbitrary single-data functions. Hybrid-scheme FHE, which exploits appropriate schemes for specific tasks, is essential for real-world applications requiring optimal performance and accuracy. However, existing FHE accelerators only adopt scheme-specific custom designs, leading to inefficiency or lack of capability to support applications in hybrid FHE settings. In this work, we propose a Unified FHE aCcelerator (UFC) that provides better performance and cost-efficiency than prior scheme-specific accelerators on hybrid FHE applications. Our design process involves a comprehensive analysis of processing flows to abstract the primitives covering all operations in hybrid FHE applications. The UFC architecture primarily comprises hardware function units for these primitives, diverging from the deeply pipelined units in previous designs. This approach enables high hardware utilization across different FHE schemes. Further-more, we propose several algorithm-hardware co-optimizations to minimize the hardware cost of supporting various data shuffling patterns in FHE. This enables high-throughput implementation of function units that provide good cost efficiency. We also propose several compiler-level optimizations to achieve high hardware utilization of the unified architecture for computing FHE data in various algorithmic parameter settings. We evaluate the performance of UFC on different FHE programs, including scheme-specific and hybrid-scheme workloads. Our experiments show that UFC provides up to 6.0 × speedup and 1.6 × delay-energy-area efficiency improvement over state-of-the-art FHE accelerators. Minxuan Zhou, Yujin Nam, Xuan Wang 0040, Youhak Lee, Chris Wilkerson, Raghavan Kumar, Sachin Taneja, Sanu Mathew, Rosario Cammarota, Tajana Rosing |
MICRO | 3 |
| 2023 | OverlaPIM: Overlap Optimization for Processing In-Memory Neural Network AccelerationabstractProcessing in-memory (PIM) can accelerate neural networks (NNs) for its extensive parallelism and data movement minimization. The performance of NN acceleration on PIM heavily depends on software-to-hardware mapping, which indicates the order and distribution of operations across the hardware resources. Previous works optimize the mapping problem by exploring the design space of per-layer and cross-layer data layout, achieving speedup over manually designed mappings. However, previous works do not consider computation overlapping across consecutive layers. By overlapping computation, we can process a layer before its preceding layer fully completes, decreasing the execution latency of the whole network. The mapping optimization without overlap analysis can result in sub-optimal performance. In this work, we propose OverlaPIM, a new framework that integrates the overlap analysis with the DNN mapping optimization on PIM architectures. OverlaPIM adopts several techniques to enable efficient overlap analysis and optimization for the whole network mapping on PIM architectures. We test OverlaPIM on popular DNN networks and compare the results to non-overlap optimization. Our experiments show that OverlaPIM can efficiently produce mappings that are 2.10 x to 4.11 x faster than the state-of-the-art mapping optimization framework. Minxuan Zhou, Xuan Wang 0040, Tajana Rosing |
DATE | 2 |