Wei Yan 0005

dblp:45/4440-5 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAMI: A Context-Aware Isolation Architecture for GPU Memories
abstract
The widespread use of GPUs in cloud and high-performance computing makes memory isolation a critical security requirement. While the programming model assumes that each thread local memory is private, the underlying hardware does not always enforce this guarantee. Weaknesses in address translation can allow one thread to access another local memory, creating a semantic gap that enables cross-thread corruption and exploitation. To address these challenges, we propose CAMI, a hardware-level framework that integrates fine-grained execution context into the memory translation pipeline. CAMI enforces a binding between the execution context of each memory access and the ownership of its target memory page, ensuring that even subtle inconsistencies in translation cannot be exploited. By introducing an efficient hardware enforcement unit within the MMU and extending page table entries with ownership metadata, CAMI achieves strong, fine-grained isolation while maintaining low performance overhead. We implement CAMI in a cycle-accurate GPU simulator and conduct comprehensive evaluations. Results show that CAMI effectively eliminates cross-thread memory access vulnerabilities with minimal runtime cost, offering a practical path toward secure and high-performance GPU architectures.
Wei Yan 0005, Qinfen Hao, Xiaochun Ye, Yier Jin, Ninghui Sun
DATE2
2025 Accelerating Oblivious Transfer with a Pipelined Architecture
abstract
With the rapid development of machine learning and big data technologies, ensuring user privacy has become a pressing challenge. Secure multi-party computation offers a solution to this challenge by enabling privacy-preserving computations, but it also incurs significant performance overhead, thus limiting its further application. Our analysis reveals that the oblivious transfer protocol accounts for up to 96.64% of execution time. To address these challenges, we propose POTA, a high-performance pipelined OT hardware acceleration architecture supporting the silent OT protocol. Finally, we implement a POTA prototype on Xilinx VCU129 FPGAs. Experimental results demonstrate that under various network settings, POTA achieves significant speedups, with maximum improvements of 22.67x for OT efficiency and 192.57x for basic operations in MPC applications.
Wei Yan 0005, Qinfen Hao, Ninghui Sun
DATE2
2025 SSMDVFS: Microsecond-Scale DVFS on GPGPUs with Supervised and Self-Calibrated ML
abstract
Over the past decade, as GPUs have evolved to achieve higher computational performance, their power density has also accelerated. Consequently, improving energy efficiency and reducing power consumption has become critically important. Dynamic voltage and frequency scaling (DVFS) is an effective technique for enhancing energy efficiency. With the advent of integrated voltage regulators, DVFS can now operate on microsecond$(\boldsymbol{\mu}\mathbf{s})$timescales. However, developing a practical and effective strategy to guide rapid DVFS remains a significant challenge. This paper proposes a supervised and self-calibrated machine learning framework (SSMDVFS) to guide microsecond-scale GPU voltage and frequency scaling. This framework features an end-to-end design that encompasses data generation, neural network model design, training, compression, and final runtime calibration. Unlike analytical models, which struggle to accurately represent GPU architectures, and reinforcement learning approaches, which can be challenging to converge during runtime, the SSMDVFS offers a practical solution for guiding microsecond-scale voltage and frequency scaling. Experimental results demonstrate that the proposed framework improves energy-delay product (EDP) by 11.09% and outperforms analytical models and reinforcement learning approaches by 13.17% and 36.80 %, respectively.
Minqing Sun, Yingtao Shen, Wei Yan 0005, Qinfen Hao, An Zou
DATE4
2025 ParTEE: A Framework for Secure Parallel Computing of RISC-V Trusted Execution Environments
Ziang Zhou, Wei Yan 0005, Qinfen Hao, Xiaochun Ye, Ninghui Sun
Euro-Par (2)4
2025 CacheGuardian: A Timing Side-Channel Resilient LLC Design
abstract
In cloud computing environments, the last-level cache (LLC) shared by multiple tenants is frequently exploited through timing side-channel attacks, enabling unauthorized data leakage. To address this issue, various defense mechanisms have been proposed. However, existing works exhibit deficiencies in terms of performance overhead, coverage of attacks, and detection accuracy. In response to these challenges, we propose CacheGuardian, a hardware-based LLC protection design which aims to provide stronger, broader, and more accurate protection against timing side-channel attacks with low performance overhead. It includes: (1) A behavior-based, generic attack detector capable of identifying multiple timing side-channel attacks in real time; (2) A cache-set-level access control mechanism that strictly restricts cache usage exclusively for the identified attackers instead of influencing all security domains.We implement our design in a gem5 simulator to evaluate both its security and performance. Our proof-of-concept attacks and SPEC 2017 benchmarks show that our design is effective against a wide range of timing side-channel attacks, reducing attack success rates by up to 256×, including camouflaged variants. Moreover, it improves the performance of benign workloads by an average of 2.26% with only 2.4% storage overhead.
Ziang Zhou, Huifeng Zhu, Wei Yan 0005, Chenglu Jin, Xuejun An, Xiaochun Ye
ICCAD5
2025 Revisiting Edge Perturbation for Graph Neural Network in Graph Data Augmentation and Attack
abstract
Edge perturbation is a basic method to modify graph structures. It can be categorized into two veins based on their effects on the performance of graph neural networks (GNNs), i.e., graph data augmentation and attack. Surprisingly, both veins of edge perturbation methods employ the same operations, yet yield opposite effects on GNNs' accuracy. A distinct boundary between these methods in using edge perturbation has never been clearly defined. Consequently, inappropriate perturbations may lead to undesirable outcomes, necessitating precise adjustments to achieve desired effects. Therefore, questions of “why edge perturbation has a two-faced effect?” and “what makes edge perturbation flexible and effective?” still remain unanswered. In this paper, we will answer these questions by proposing a unified formulation and establishing a quantizable boundary between two categories of edge perturbation methods. Specifically, we conduct experiments to elucidate the differences and similarities between these methods and theoretically unify the workflow of these methods by casting it to one optimization problem. Then, we devise Edge Priority Detector (EPD) to generate a novel priority metric, bridging these methods up in the workflow. Experiments show that EPD can make augmentation or attack flexibly and achieve comparable or superior performance to other counterparts with less time overhead.
Xin Liu 0073, Yuxiang Zhang 0011, Meng Wu 0006, Mingyu Yan, Wei Yan 0005, Shirui Pan, Xiaochun Ye, Dongrui Fan
IEEE Trans. Knowl. Data Eng.6
2024 Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
Yuxiang Zhang 0011, Xin Liu 0073, Meng Wu 0006, Wei Yan 0005, Mingyu Yan, Xiaochun Ye, Dongrui Fan
Euro-Par (2)4
2024 MPC-PAT: A Pipeline Architecture for Beaver Triple Generation in Secure Multi-party Computation
abstract
Secure Multi-Party Computation (MPC) is proposed to protect the data privacy from a group of parties, enabling collaborative computation of correct results for target functions. SPDZ, a set of mature MPC protocols widely used in machine learning and other scenarios, requires a significant number of Beaver triples for secure multiplications among parties. Given no Trusted Third Party (TTP) participated, the generation time constitutes over 92% of the total running time. This paper introduces MPC-PAT, a high-performance pipeline architecture designed for efficient Beaver triple generation. MPC-PAT accelerates random number generation, hash function, and modular multiplication(MM) in two finite fields. The evaluation results from its FPGA implementation demonstrate 99× speed-up for basic operations and 136× speed-up for various convolutional networks compared to the existing SPDZ works.
Wei Yan 0005, Qinfen Hao, Ninghui Sun
ITC-Asia2
2020 FLASH: FPGA Locality-Aware Sensitive Hash for Nearest Neighbor Search and Clustering Application
abstract
A locality sensitive hash (LSH) is a function to identify similar items in data sets. However, traditional LSH based algorithms are rarely implemented on hardware due to the high demand of computation, which limits its usage. In this paper, we propose a novel LSH design and hardware implementation called FPGA locality-aware sensitive hash (FLASH). With the unique hardware delay generated during the fabrication process, a FLASH can reduce the dimensionality of coordinate distance calculation and improve the efficiency of nearest neighbor search (NNS). It is also applied to 2-D image clustering. The experimental results show the practical value of our FLASH applications.
Wei Yan 0005, Sara Tehranipoor, Xuan Zhang 0001, John A. Chandy
FPL1
2020 PCBChain: Lightweight Reconfigurable Blockchain Primitives for Secure IoT Applications
abstract
In the era of ubiquitous intelligence, the Internet of Things (IoT) holds the promise as a breakthrough technology to enable diverse applications that benefit societal problems. Yet interconnecting myriad heterogeneous IoT devices across various application domains remain a security challenge. Decentralized technology has recently emerged as a powerful primitive in building distributed applications to facilitate secure transactions between mutually distrustful parties in a trustworthy manner. Unfortunately, these decentralized protocols demand computing resources and power far beyond the reach of resource-constrained IoT devices, preventing the full adoption of distributed consensus platform in the IoT setting. In this article, we address the key bottleneck to enable blockchain in resource-constrained IoT devices. We propose a lightweight implementation of proof-of-work (PoW) mining with reconfigurable hardware primitives. By replacing the hash and cryptographic functions in classic blockchain protocol with secure and efficient hardware implementations, our proposed solution can significantly reduce hardware resources and power overheads of PoW mining, while improving the transaction speed of large-scale IoT systems. Finally, we demonstrate the algorithm by proposing an antispoofing solution for GPS navigation among lightweight IoT devices. As a replacement for position computation, a mining process generates the expected coordinates with the correct initial value and function configuration.
Wei Yan 0005, Ning Zhang 0017, Laurent Njilla, Xuan Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2018 DVFT: A Lightweight Solution for Power-Supply Noise-Based TRNG Using Dynamic Voltage Feedback Tuning System
Sara Tehranipoor, Paul A. Wortman, Nima Karimian, Wei Yan 0005, John A. Chandy
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Phase calibrated ring oscillator PUF design and implementation on FPGAs
abstract
A ring oscillator physical unclonable function (RO PUF) is an application-constrained hardware security primitive that can be used for authentication and key generation. PUFs depend on variability during the fabrication process to produce random outputs that are nevertheless stable across multiple measurements. Unfortunately, RO PUFs are known to be unstable especially when implemented on an Field Programmable Gate Array (FPGA). In this work, we comprehensively evaluate the RO PUF's stability on FPGAs, and we propose a phase calibration process to improve the stability of RO PUFs. The results show that the bit errors in our PUFs are reduced to less than 1%.
Wei Yan 0005, Chenglu Jin, Sara Tehranipoor, John A. Chandy
FPL1
2017 Investigation of DRAM PUFs reliability under device accelerated aging effects
abstract
Physical Unclonable Functions are promising candidates for lightweight authentication applications as they are hard to predict and clone. PUFs are dependent on process variations that occurs during silicon chip fabrication. As the CMOS technology scales down towards nanoscale dimensions, there are increasing transistor reliability challenges which impact the lifetime of integrated circuits. These issues are known as aging effects, which result in degradation of the performance of circuits. In this paper, we analyze the effects of aging on the reliability of intrinsic DRAM PUFs. We present accelerated aging experimental results over 18 months (from Sep. 2014 to Feb. 2016) on 3 DRAM PUFs. Based on our observations, DRAM PUFs maintain their reliability over time, and thus, validate the use of DRAM PUFs in a number of applications such as system authentications.
Sara Tehranipoor, Nima Karimian, Wei Yan 0005, John A. Chandy
ISCAS3
2017 PUF-Based Fuzzy Authentication Without Error Correcting Codes
abstract
Counterfeit integrated circuits (IC) can be very harmful to the security and reliability of critical applications. Physical unclonable functions (PUFs) have been proposed as a mechanism for uniquely identifying ICs and thus reducing the prevalence of counterfeits. However, maintaining large databases of PUF challenge response pairs (CRPs) and dealing with PUF errors make it difficult to use PUFs reliably. This paper presents an innovative approach to authenticate CRPs on PUF-based ICs. The proposed method can tolerate considerable bit errors from responses of PUFs without the use of error correcting codes. Different types of optimization methods are applied to improve the overall performance. The simulation shows that it is successful in authenticating 99.96% authorized chips and filtering out 99.92% cloned chips by tolerating 12 errors in 128 bits. The results are verified with ring oscillator PUF and arbiter PUF implementations on Kintex-7 FPGA. The approach saves hardware and software resources significantly, compared to those of other authentication solutions.
Wei Yan 0005, Sara Tehranipoor, John A. Chandy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 DRAM-Based Intrinsic Physically Unclonable Functions for System-Level Security and Authentication
abstract
A physically unclonable function (PUF) is an irreversible probabilistic function that produces a random bit string. It is simple to implement but hard to predict and emulate. PUFs have been widely proposed as security primitives to provide device identification and authentication. In this paper, we propose a novel dynamic-memory-based PUF [dynamic RAM PUF (DRAM PUF)] for the authentication of electronic hardware systems. The DRAM PUF relies on the fact that the capacitor in the DRAM initializes to random values at startup time. Most PUF designs require custom circuits to convert unique analog characteristics into digital bits, but using our method, no extra circuitry is required to achieve a reliable 128-bit PUF. The results show that the proposed DRAM PUF provides a large number of input patterns (challenges) compared with other memory-based PUF circuits such as static RAM PUFs. Our DRAM PUFs provide highly unique PUFs with a 0.4937 average interdie Hamming distance. We also propose an enrollment algorithm to achieve highly reliable results to generate PUF identifications for system-level security. This algorithm has been validated on real DRAMs with an experimental setup to test different operating conditions.
Sara Tehranipoor, Nima Karimian, Wei Yan 0005, John A. Chandy
IEEE Trans. Very Large Scale Integr. Syst.3
2015 A Novel Way to Authenticate Untrusted Integrated Circuits
abstract
Counterfeit Integrated Circuits (IC) can be very harmful to the security and reliability of critical applications. Physical Unclonable Functions (PUF) have been proposed as a mechanism for uniquely identifying ICs and thus reducing the prevalence of counterfeits. However, maintaining large databases of PUF challenge response pairs and dealing with PUF errors makes it difficult to use PUFs reliably. This paper presents an innovative approach to authenticate PUF challenge response pairs on IC chips. The proposed method can tolerate considerable bit errors from responses of PUFs without the use of error correcting codes. It is successful in authenticating 99.96% authorized chips and filtering out 99.92% cloned chips. The overhead is reduced by 65.62% compared to that of other authenticating solutions.
Wei Yan 0005, Sara Tehranipoor, John A. Chandy
ICCAD1