Simei Yang

dblp:249/9646 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-0130-8176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 pHNSW: PCA-Based Filtering to Accelerate HNSW Approximate Nearest Neighbor Search
abstract
Hierarchical Navigable Small World (HNSW) has demonstrated impressive accuracy and low latency for high-dimensional nearest neighbor searches. However, its high computational demands and irregular, large-volume data access patterns present significant challenges to search efficiency. To address these challenges, we introduce pHNSW, an algorithm-hardware co-optimized solution that accelerates HNSW through Principal Component Analysis (PCA) filtering. On the algorithm side, we apply PCA filtering to reduce the dimensionality of the dataset, thereby lowering the volume of neighbor access and decreasing the computational load for distance calculations. On the hardware side, we design the pHNSW processor with custom instructions to optimize search throughput and energy efficiency. In the experiments, we synthesized the pHNSW processor RTL design with a 65nm technology node and evaluated it using DDR4 and HBM1.0 DRAM standards. The results show that pHNSW boosts Queries per Second (QPS) by $14.47 \times \sim 21.37 \times$ on a CPU and $5.37 \times \sim 8.46 \times$ on a GPU, while reducing energy consumption by up to $57.4 \%$ compared to standard HNSW implementation.
Guangyi Zeng, Paul Delestrac, Enyi Yao, Simei Yang
ASP-DAC5
2025 Buffer Size Impact on End-to-End DNN Acceleration in Near-Bank DRAM-PIM Systems
abstract
Processing-in-Memory (PIM) helps reduce DRAM access bottlenecks in accelerating Deep Neural Networks (DNNs). Industrial solutions like SK Hynix GDDR6-AiM [1] place processing cores (PIMcores) near DRAM banks to lower data movement and latency. However, current near-bank PIM designs face two key limitations: (1) PIMcores mainly support multiply-accumulate (MAC) operations, while other tasks are handled by the host CPU or GPU, causing heavy host-PIM traffic. (2) The buffer capacity is limited (e.g., 2KB global buffer per channel [1]), reducing opportunities for data reuse.
Yunyu Ling, Simei Yang
SYSTOR3
2025 A Spin Scale-Aware Self-Adaptive Ising Annealing Processing Architecture for Combinatorial Optimization Problems
abstract
The Ising annealing processor has emerged as a promising approach to accelerate the discovery of the optimal solutions for a wide range of combinatorial optimization problems (COPs), by mapping various COPs into a unified Ising model. However, fixed computational strategies and inflexible architectures make previous designs suffer from a low hardware resource utilization rate when the numbers of the total required and real-time flipped spins vary across different COPs and iteration steps. In this paper, a novel spin scale-aware self-adaptive Ising annealing processing architecture (AIAPA) is proposed to address this problem, with an adaptive computational strategy, a custom instruction set, multi-traffic mode routers, and a fully-pipelined computing array. It can dynamically adapt to the varying scenarios during the Ising annealing process to maximize the performance of limited hardware resources. Its prototype, supporting 65k fully-connected spins, is implemented on an FPGA platform, operates at a clock frequency of 188 MHz. The AIAPA achieves up to a 24.22 times faster annealing speed compared to the state-of-the-art FPGA design on the max-cut optimization problem while maintaining a high convergence accuracy.
Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Longyuan Kang, Simei Yang, Enyi Yao
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Multi-Level Analysis of GPU Utilization in ML Training Workloads
abstract
Training time has become a critical bottleneck due to the recent proliferation of large-parameter ML models. GPUs continue to be the prevailing architecture for training ML models. However, the complex execution flow of ML frameworks makes it difficult to understand GPU computing resource utilization. Our main goal is to provide a better understanding of how efficiently ML training workloads use the computing resources of modern GPUs. To this end, we first describe an ideal reference execution of a GPU-accelerated ML training loop and identify relevant metrics that can be measured using existing profiling tools. Second, we produce a coherent integration of the traces obtained from each profiling tool. Third, we leverage the metrics within our integrated trace to analyze the impact of different software optimizations (e.g., mixed-precision, various ML frameworks, and execution modes) on the throughput and the associated utilization at multiple levels of hardware abstraction (i.e., whole GPU, SM subpartitions, issue slots, and tensor cores). In our results on two modern GPUs, we present seven takeaways and show that although close to 100% utilization is generally achieved at the GPU level, average utilization of the issue slots and tensor cores always remains below 50% and 5.2%, respectively.
Paul Delestrac, Debjyoti Bhattacharjee, Simei Yang, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo
DATE3
2021 0-1 ILP-based run-time hierarchical energy optimization for heterogeneous cluster-based multi/many-core systems
Simei Yang, Sébastien Le Nours, Maria Mendez Real, Sébastien Pillement
J. Syst. Archit.1