Yi Estelle Wang

dblp:67/6649-16 · also Yi Wang 0016 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0003-2957-4203ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 3 since 2021Security and privacy · 3Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Defense Against Adversarial Samples for Replay Attacks in Automotive AVTP Streams via a Novel Bidirectional Mamba-Based Intrusion Detection System
abstract
As the core protocol for transmitting Advanced Driver Assistance Systems (ADAS) perception video streams over automotive Ethernet, Audio Video Transport Protocol (AVTP) is vulnerable to replay attacks [1] due to its lack of authentication and integrity protection mechanisms. The existing supervised deep learning methods primarily target conventional replay attack detection, but their large model sizes make them unsuitable for resource-constrained and real-time deployment in embedded automotive systems. To address these issues, we propose a novel intrusion detection system (IDS) against Adversarial Samples for Replay Attacks (AdvRA) in Automotive AVTP Streams, which is based on Frame-Difference, Time-Shared Convolution and Bidirectional Mamba (FD-TSConv-BiMamba). Specifically, we first construct an efficient feature engineering pipeline that reduces the number of dimensional features from 438 to 40. Second, we propose a lightweight detection algorithm based on an improved Mamba structure. The parameter complexity of the proposed algorithm isO(d·m+n·d), which is significantly lower than the parameter complexity of GRU (O(3(nd+n2))), LSTM (O(4(nd+n2))) and CNN (O(f2·cin·cout)). Finally, to validate the effectiveness of our model in detecting AdvRA, we construct a conditional Generative Adversarial Network (cGAN) to generate the AdvRA dataset. Experimental results show that the proposed model achieves the parameters of the accuracy and F1-score with 0.9974 and 0.9970 under the indoor scenario using the window size of 114, and with 1.0000 and 0.9999 under the driving scenario using the window size of 108, respectively. In addition, the model size is only 1.17 MB. Compared to advanced methods such as TOW-IDS [2] at 11.30 MB and CNN-IDS [3] at 3.58 MB, the model size is reduced by approximately 89.6% and 67.3%, respectively. Furthermore, for the AdvRA detection task, our model achieves an F1-score of 0.9917, outperforming GRU and LSTM by 7.15% and 6.27%, respectively, with a detection speed of 3155 samples/s. Our code is available at: https://github.com/Haihang-Zhao/FD_TSConv_Mamba.
Haihang Zhao, Yi Estelle Wang, Anyu Cheng
IEEE Internet Things J.2
2025 Safeguarding GRU-Based Intrusion Detection Systems From Adversarial Attacks With Dynamic Label Watermark in CAN Bus Communication
abstract
Intrusion detection systems (IDS) for control area network (CAN) bus communication using deep learning models face threats from adversarial closed-box. attacks in the Internet of Vehicles (IoVs). Although watermark techniques are proposed as defences, they lack concealment and are vulnerable. Current watermark methods for time-series data-based applications need cloud-based verification and terminal-based generation, and they cannot meet real-time requirements with large resources. To address these issues, we propose a real-time gated recurrent units (GRUs) based IDS with for CAN bus communication via a novel dynamic label watermark (DLW) method. In detail, we design a multitask learning structure at the terminal side only to detect conventional intrusion attacks. At the same time, we propose a novel DLW method applied to time-series data to defend against adversarial closed-box. attacks. Experimental results show that for the detection of Denial of Service (DoS), revolutions per minute (RPM) spoofing, and fuzzing attacks, our model achieves 1.00000, 1.00000, and close to 1.00000 with the recall, accuracy, F1 score, and precision, respectively. For detection of gear spoofing, our model with the same metrics achieves 1.00000, which are 0.0882, 0.0001, 0.0459, and 0.0208 better than CANLite and the same as ConvLSTM-GNB. Finally, we construct a new adversarial closed-box. attack embedded with four attacks above to validate the resistance and performance of our model (achieving 116 KB code size), which is 58% smaller, 0.9%–35.7% faster, and 1.52%–10.5% improvement of same metrics compared to the baseline model (LSTM).
Haihang Zhao, Yi Estelle Wang, Anyu Cheng, Hongrong Wang
IEEE Internet Things J.2
2025 A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 A 2.793µW Near-Threshold Neuronal Population Dynamics Simulator for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for a component within the back-end of simultaneous localization and mapping. A custom discretized procedural algorithm including injection, finite difference update, activation, and inhibition to approximate neuronal population dynamics is developed for digital implementation. Fabricated using a 40nm technology, the test chip features a scalable neuron 22 × 22 array with 0.1358mm2core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming 2.793µW of power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0016, Tony Tae-Hyoung Kim, Yuanjin Zheng
ISCAS4
2022 A3GAN: Attribute-Aware Anonymization Networks for Face De-identification
abstract
Face de-identification (De-ID) removes face identity information in face images to avoid personal privacy leakage. Existing face De-ID breaks the raw identity by cutting out the face regions and recovering the corrupted regions via deep generators, which inevitably affect the generation quality and cannot control generation results according to subsequent intelligent tasks (eg., facial expression recognition). In this work, for the first attempt, we think the face De-ID from the perspective of attribute editing and propose an attribute-aware anonymization network (A3GAN) by formulating face De-ID as a joint task of semantic suppression and controllable attribute injection. Intuitively, the semantic suppression removes the identity-sensitive information in embeddings while the controllable attribute injection automatically edits the raw face along the attributes that benefit De-ID. To this end, we first design a multi-scale semantic suppression network with a novel suppressive convolution unit (SCU), which can remove the face identity along multi-level deep features progressively. Then, we propose an attribute-aware injective network (AINet) that can generate De-ID-sensitive attributes in a controllable way (i.e., specifying which attributes can be changed and which cannot) and inject them into the latent code of the raw face. Moreover, to enable effective training, we design a new anonymization loss to let the injected attributes shift far away from the original ones. We perform comprehensive experiments on four datasets covering four different intelligent tasks including face verification, face detection, facial expression recognition, and fatigue detection, all of which demonstrate the superiority of our face De-ID over state-of-the-art methods.
Liming Zhai, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yi Estelle Wang, Yang Liu 0003
ACM Multimedia5
2021 A Fault Resistant AES via Input-Output Differential Tables with DPA Awareness
abstract
Nowadays, hardware-based AES faces more than one type of Side-Channel-Attacks (SCA), such as the Differential Power Analysis (DPA) attack and the Differential Fault Analysis (DFA) attack. However, most of the current DFA-resistant implementations of AES only focus on resisting the DFA attack but with minimal concurrent considerations on their DPA resistance capability. In this paper, we propose a fault resistant AES with DPA awareness. First, we evaluate the DPA resistance for different implementation architectures of AES before the implementation with S-Box over GF(24)2is selected. Second, we evaluate the DPA resistance for different fault detection architectures before the concurrent fault detection method is selected. Third, we propose a novel fault resistant technique for AES using input-output differential tables over GF(24)2. We have performed tests based on Partial Guessing Entropy (PGE) to evaluate the DPA resistance for the existing DFA-resistant designs and our proposed design. Experimental results prove that our design has a slower convergence speed (around 33%) with a 100% fault coverage rate and less area than the existing countermeasure designs for fault injection. Results also show that the fault detection designs weaken their DPA resistance, which indicates the importance of co-design of DFA and DPA to achieve less power information leakage.
Yi Estelle Wang, Marc Stöttinger, Yajun Ha
ISCAS1
2020 Ori: A Greybox Fuzzer for SOME/IP Protocols in Automotive Ethernet
abstract
With the emergence of smart automotive devices, the data communication between these devices gains increasing importance. SOME/IP is a light-weight protocol to facilitate inter- process/device communication, which supports both procedural calls and event notifications. Because of its simplicity and capability, SOME/IP is getting adopted by more and more automotive devices. Subsequently, the security of SOME/IP applications becomes crucial. However, previous security testing techniques cannot fit the scenario of vulnerability detection SOME/IP applications due to miscellaneous challenges such as the difficulty of server-side testing programs in parallel, etc. By addressing these challenges, we propose Ori - a greybox fuzzer for SOME/IP applications, which features two key innovations: the attach fuzzing mode and structural mutation. The attach fuzzing mode enables Ori to test server programs efficiently, and the structural mutation allows Ori to generate valid SOME/IP packets to reach deep paths of the target program effectively. Our evaluation shows that Ori can detect vulnerabilities in SOME/IP applications effectively and efficiently.
Yuekang Li, Hongxu Chen 0001, Cen Zhang, Siyang Xiong, Chaoyi Liu, Yi Estelle Wang
APSEC6
2019 A Highly Efficient Side Channel Attack with Profiling through Relevance-Learning on Physical Leakage Information
abstract
We propose a Profiling through Relevance-Learning (PRL) technique on Physical Leakage Information (PLI) to extract highly correlated PLI with processed data, as to achieve a highly efficient yet robust Side Channel Attack (SCA). There are four key features in our proposed PRL. First, variance analysis on PLI is implemented to determine the boundary of the clusters and objects of the clusters. Second, the nearest-neighbor k-NN variance clustering is used to reduce the sampling points of PLI by clustering the high variance sampling points and discarding the low variance sampling points of PLI measurements (traces). These clustered sampling points, which are highly correlated with the processed data, contain pertinent leakage information related to the secret key. Third, the information associated with the secret key is spread in several neighboring sampling points with different degrees of leakages. We analytically derive the Key-leakage relevance factor for each clustered sampling point to quantify the degree of leakage associated with the secret key. Fourth, by means of Hebbian learning, a weight proportional to the Key-leakage relevance factor is updated iteratively based on the values of relevance factor and traces of the sampling points. The converged weights which are being assigned to clustered sampling points are linked to their associated PLI to further increase the correlation of the PLI with the processed data. Therefore, the required number of PLI measurements, to reveal the secret key, can be reduced significantly. In addition, we analytically show that the computational complexity of our proposed PRL is O(n) when compared to the reported profiling techniques having O(n2) and O(n3) computational complexities. Based on the experiments of our proposed PRL performed on the PLI of AES-128 algorithm, the results depicting that the sampling points of PLI are reduced 87 percent after the k-NN variance clustering. The converged weight with learning error rate106traces, our proposed PRL is ~2; 000x more efficient in performing SCA.
Ali Akbar Pammu, Kwen-Siong Chong, Yi Estelle Wang, Bah-Hwee Gwee
IEEE Trans. Dependable Secur. Comput.3
2017 On the Security of In-Vehicle Hybrid Network: Status and Challenges
Tianxiang Huang, Jianying Zhou 0001, Yi Estelle Wang, Anyu Cheng
ISPEC3
2017 A DFA-Resistant and Masked PRESENT with Area Optimization for RFID Applications
abstract
Radio-Frequency Identification (RFID) tag-based applications are usually resource constrained and security sensitive. However, only about 2,000 gate equivalents in a tag can be budgeted for implementing security components [27]. This requires not only lightweight cryptographic algorithms such as PRESENT (around 1,000 gate equivalents) but also lightweight protections against modern Side Channel Attacks (SCAs). With this budget, the first-order masking and fault detection are two suitable countermeasures to be developed for PRESENT. However, if both countermeasures are applied without any optimization, it will significantly exceed the given area budget. In this work, we optimize area to include both countermeasures to maximize the security for PRESENT within this RFID area budget. The most area-consuming parts of the proposed design are the masked S-boxes and the inverse masked S-boxes. To optimize the area, we have deduced a computational relationship between these two parts, which enables us to reuse the hardware resource of the masked S-boxes to implement the inverse masked S-boxes. The proposed design takes up only 2,376 gates with UMC 65nm CMOS technology. Compared with the unoptimized design, our implementation reduces the overall area by 28.45%. We have tested the effectiveness of the first-order Differential Power Analysis (DPA) and Differential Fault Analysis (DFA) -resistant countermeasures. Experimental results show that we have enhanced the SCA resistance of our PRESENT implementation.
Yi Estelle Wang, Yajun Ha
ACM Trans. Embed. Comput. Syst.1
2016 High throughput and resource efficient AES encryption/decryption for SANs
abstract
To secure the data stored in large-scale Storage Area Network (SAN) applications, high throughput Advanced Encryption Standard (AES) encryption and decryption are required. However, this solution may take up more hardware resources, which leads to unscalability for future needs. To solve this problem, we develop a high throughput and resource efficient AES encryption/decryption based on FPGA, which fully exploiting the dedicated resources of modern FPGAs, such as Block RAM (BRAM) and Digital Signal Processing (DSP) slices. We also propose a unified architecture for AES encryption and decryption. Furthermore, we move the map and the inverse map functions outside the AES encryption/decryption round. In order to shorten the critical path, we optimized the transformation matrix of the map function and its inverse transformation matrix. We use the same hardware resource to perform computations of both SubBytes and InvSubBytes, as well as computations of MixColumns and InvMixColumns. Finally, proper pipelined registers and DSP slices have been inserted into the proposed unrolling architecture to achieve high throughput. Experimental results show that our designs can achieve 78.22 Gbits/s using 5613 slices, 144 DSP slices without BRAM; or 68.44 Gbits/s using 4345 slices, 171 DSP slices with 400X36K BRAMs on XC6VLX240T FPGA.
Yi Estelle Wang, Yajun Ha
ISCAS1
2016 Efficient differential fault analysis attacks to AES decryption for low cost sensors in IoTs
abstract
Robust sensor system plays an important role in Internet of Things (IoTs). These intelligent sensors are required to be low cost and reliable, which provides confidentiality for private sensitive data. However, this protected system is still under the risk of Differential Fault Analysis (DFA) attacks. In this paper, we focus on DFA attacks to AES decryption as decryption receives the equalling importance as encryption. First, we induce a fault at the input of the third round in the procedure of AES decryption, in which w e successfully break it using one pair of fault-free and faulty plaintexts within 232 searching space. Then, we improve this attack by use of S-Box distribution table, which reduces the computational time from 853 ms to 70 ms on a dual Intel(R) Pentium(R) E6700 core (3.20 GHz). Compared to the existing work, our proposed attack reduces 79.5% computational time when both methods employ two pairs of fault-free and faulty ciphertexts/plaintexts.
Yi Estelle Wang, Renfa Li
ISCAS2
2015 Reconfiguring Three-Dimensional Processor Arrays for Fault-Tolerance: Hardness and Heuristic Algorithms
abstract
With the increased density of three-dimensional (3D) processor arrays, faults can potentially occur quite often due to power overheating during massively parallel computing. In order to achieve fault-tolerance under such a scenario, an effective way is to find an as large as possible logical fault-free subarray of m' × n' × h' from a faulty array of m × n × h (m' ≤ m, n' ≤ n, h' ≤ h), such that an original application can still work on the m' × n' × h' subarray. This paper investigates the problem of constructing maximum fault-free subarrays with minimum interconnection length from 3D arrays with faults. First, we prove that constructing maximum logical array (MLA) is NP-complete. We propose a linear-time algorithm which is capable of producing an MLA for the problem with the constraint of selected indexes. Second, we prove that minimizing the interconnection length (inter-length) of the MLA is NP-hard. We propose an efficient heuristic which significantly reduces the inter-length by revising each logical plane of the MLA. This leads to the reduction of communication cost, capacitance and dynamic power dissipation. In addition, we propose a lower bound for the inter-length of the MLA to evaluate the proposed algorithms. Simulation results show that, the size of logical array can be improved up to 62.6 percent in average, and the inter-length redundancy can be reduced by 22.7 percent in average, compared to the state-of-the-art, for all cases considered.
Guiyuan Jiang, Jigang Wu, Yajun Ha, Yi Estelle Wang
IEEE Trans. Computers4
2014 FPGA-based high throughput XTS-AES encryption/decryption for storage area network
abstract
The key issue to improve the performance for secure large-scale Storage Area Network (SAN) applications lies in the speed of its encryption/decryption module. Software-based encryption/decryption cannot meet throughput requirements. To solve this problem, we propose a FPGA-based XTS-AES encryption/decryption to suit the needs for secure SAN applications with high throughput requirements. Besides throughput, area optimization is also considered in this proposed design. First, we reuse the same AES encryption to produce the tweak value and unify the operations of AES encryption/decryption in XTS-AES encryption/decryption. Second, we transfer the computations of AES encryption/decryption from GF(28) to GF(24)2, which enables us move the map and the inverse map functions outside the AES round. Third, we propose to support the SubBytes and the inverse SubBytes by the same hardware component. Finally, pipelined registers have been inserted into the proposed unrolled architecture for XTS-AES encryption/decryption. The experiments show that the proposed design achieves 36.2 Gbits/s throughput using 6784 slices on XC6VLX240T FPGA.
Yi Estelle Wang, Akash Kumar 0001, Yajun Ha
FPT1
2013 FPGA based Rekeying for cryptographic key management in Storage Area Network
abstract
Rekeying process plays an important role in secure large-scale Storage Area Network (SAN) applications. Software based Rekeying management could not completely prevent sensitive information leakage from theoretical and physical attacks. Traditional Rekeying process will suffer from decrypting the large data using the old key and encrypting it with the new key. In order to solve these problems, we proposed a FPGA based flexible and low-cost rekeying management to improve the security and reduce the processing time. In the proposed method, enveloping key is defined and added into the rekeying process to protect the real private key and the user's access key. During the rekeying process, the user's access key is substituted and send back to the user instead of real private key. In order to save the transformation time between the Policies Key Control (software) and key management (hardware), we proposed index extraction solution to shorten bit width of transformation from 256-bit to only 32-bit. Experimental results show that our proposed method only takes up 1.099 ms for rekeying process compared with the existing design with 3.91 ms execution time.
Yi Estelle Wang, Yajun Ha
FPL1
2013 sAES: A high throughput and low latency secure cloud storage with pipelined DMA based PCIe interface
abstract
Modern cloud storage requires a high throughput and low latency data protection system, which is usually implemented with an Advanced Encryption Standard (AES) hardware accelerator connected with CPU through PCI Express (PCIe). However, most existing systems cannot simultaneously achieve high throughput and low latency, as they impose conflicting requirements to the block size of packets used in PCIe. High throughput requires the block size to be larger, while low latency requires the block size to be smaller. To provide both high throughput and low latency, we have developed an FPGA based data protection system called sAES. It uses a highly pipelined Direct Memory Access (DMA) based PCIe interface. It can achieve 10.4 Gbps throughput when the block size is 512 bytes, which is 51 times higher than the state-of-the-art Speedy PCIe interface [1]. The worst latency of sAES is only 4.368 μs when its block size is 512 bytes.
Yongzhen Chen, Miguel Rodel Felipe, Yi Estelle Wang, Yajun Ha, Shu Qin Ren, Khin Mi Mi Aung
FPT3
2013 An area-efficient shuffling scheme for AES implementation on FPGA
abstract
Power analysis attack is an efficient way to retrieve the sensitive information from the hardware implementation of modern cryptographic algorithms, such as Advance Encryption Standard (AES). First-order masking could defend against Differential Power Analysis (DPA) attack without extra hardware support. However, it is vulnerable to Higher-Order Differential Power Analysis (HODPA) attack. HODPA attack could be avoided using a higher order masking scheme, but it takes up huge hardware resources. In this paper, we propose a low cost shuffling scheme for FPGA based AES implementations, which is able to efficiently resist against HODPA attack. We reuse our previous masked S-box proposed in [20-21] to reduce hardware resources and defend against glitch attacks. Also, we reorder the executing sequence of the MixColumns and the AddRoundKey transformations in the first-second, the last and the second to last rounds. It is difficult for the attackers to find the “real” attacking points in our proposed design. The experimental results show that our proposed design is only 5.6% larger than the masking only scheme.
Yi Estelle Wang, Yajun Ha
ISCAS1
2013 FPGA based unified architecture for public key and private key cryptosystems
Yi Estelle Wang, Renfa Li
Frontiers Comput. Sci.1
2013 An improved ridge features extraction algorithm for distorted fingerprints matching
Thi Hanh Nguyen, Yi Estelle Wang, Renfa Li
J. Inf. Secur. Appl.2
2008 A unified architecture for a public key cryptographic coprocessor
Yi Estelle Wang, Douglas L. Maskell, Jussipekka Leiwo
J. Syst. Archit.1