VLDB 2026 Research / reviewers in the wild / expert
Yiwen Gao 0001
dblp:08/7458-1
· DBLP profile ↗
34ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0001-5446-2014ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 16 · 2 first-author · 15 since 2021Systems, architecture and hardware · 12 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DR-Arena: an Automated Evaluation Framework for Deep Research AgentsabstractAs Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task performance has become a critical bottleneck.Current benchmarks predominantly rely on static datasets, which suffer from several limitations: limited task generality, temporal misalignment, and data contamination.To address these, we introduce DR-Arena, a fully automated evaluation framework that pushes DR agents to their capability limits through dynamic investigation.DR-Arena constructs real-time Information Trees from fresh web trends to ensure the evaluation rubric is synchronized with the live world state, and employs an automated Examiner to generate structured tasks testing two orthogonal capabilities: Deep reasoning and Wide coverage.DR-Arena further adopts Adaptive Evolvement Loop, a state-machine controller that dynamically escalates task complexity based on real-time performance, demanding deeper deduction or wider aggregation until a decisive capability boundary emerges.Experiments with six advanced DR agents demonstrate that DR-Arena achieves a Spearman correlation of 0.94 with the LM-SYS Search Arena leaderboard.This represents state-of-the-art alignment with human preferences without any manual efforts, validating DR-Arena as a reliable alternative for costly human adjudication. Yiwen Gao 0001, Yang Deng 0002, Wenxuan Zhang 0001 |
ACL (1) | 1 |
| 2026 | Memory-Optimized Masked CRYSTALS-Kyber Implementation on ARM Cortex-M4
Ruiqi Hou, Yiwen Gao 0001, Wei Cheng 0003, Yuejun Liu, Jingdian Ming, Yongbin Zhou |
ICDCS | 2 |
| 2026 | Ats-dta: adaptive two-stage DDoS detection with dynamic threshold adjustment in SDN networksabstractAbstract Software-Defined Networking (SDN), as a new network architecture, has brought convenience, but also suffered from the threat of Distributed Denial of Service (DDoS) attack. However, most existing DDoS attack detection schemes for SDN employ only a single detection method, leading to imbalances in detection speed, system overhead, and detection accuracy. Even though a few schemes improve detection efficiency and accuracy through two-stage detection, they still suffer from low system flexibility and do not support dynamic threshold adjustment. In order to resolve these issues, we propose an adaptive two-stage DDoS attack detection scheme with Dynamic Threshold Adjustment (ATS-DTA for short), which contains three sub-modules. More specifically, by dividing DDoS attack detection into two modules: a conditional entropy-based network traffic anomaly detection phase and a DDoS attack detection phase based on machine learning methods. Additionally, an adaptive threshold adjustment module is introduced to improve the system’s flexibility. Finally, the experimental results show that our scheme, compared to related schemes, not only significantly improves detection accuracy and speed but also supports flexible and dynamic threshold adjustment. Specifically, our method achieves an average accuracy improvement of 1.91% and a precision increase of 1.23% over baseline methods, underscoring its effectiveness in adapting to complex and evolving network environments. These advantages illustrate that our ATS-DTA scheme provides a more balanced, efficient, and reliable solution for DDoS detection in dynamic network scenarios. Tianrui Bai, Yuan Liu 0013, Yiwen Gao 0001, Yongbin Zhou |
Cybersecur. | 3 |
| 2026 | REACTS: robust encrypted search for dynamic spatial-textual data with permission controlabstractAbstract The proliferation of spatial-textual data applications has created significant challenges in securely managing such data within untrusted cloud environments. Existing encrypted spatial-textual data retrieval schemes primarily focus on static data and overlook the complexities of practical data updates, particularly lacking robustness in managing irrational updates. In this paper, we introduce a novel robust dynamic encrypted spatial-textual data search scheme, called , that enhances existing systems by addressing the challenges of security and robustness in data dynamic settings. This is the first scheme to simultaneously achieve forward security, Type-I $$^-$$ - backward security, and enhanced robustness for boolean range queries on spatial-textual data. We formally define the security model and classify three levels of robustness. Our customized Asymmetric Scalar-Product-Preserving Encryption (ASPE) design incorporates an “update check mechanism” can efficiently mitigate repeated and disruptive update attacks while supporting efficient search and update permission control. Experimental evaluations demonstrate that only maintains 100% precision and recall while showing practical search efficiency, even outperforming the existing static scheme with a similar ASPE-based approach. Jiabei Wang, Dandan Xu, Yiwen Gao 0001, Yongbin Zhou |
Cybersecur. | 4 |
| 2026 | SERP-SCA: A Strengthened Side-Channel Attack Framework for FALCON SignatureabstractNIST has selected Falcon as one of the standardized post-quantum digital signature algorithms, making the security of Falcon against side-channel attacks (SCAs) a critical area of concern. This paper introduces SERP-SCA, a strengthened SCA framework for Falcon, which exploits leakage from floating-point multiplications to recover secret floating-point values, which can subsequently be mapped to the coefficients of the secret key. The framework adopts a divide-and-conquer strategy to target the mantissa, exponent, and sign bit of floating-point values, with specialized optimizations for attacks on the mantissa and exponent. For mantissa attacks, SERP-SCA integrates bit segmentation with a sequential attack strategy to fully exploit the leakage from mantissa multiplication and incorporates an error-correction mechanism to further enhance attack efficiency. For exponent attacks, SERP-SCA leverages Fast Fourier Transform (FFT) properties to narrow the enumeration space, significantly improving time efficiency. We conduct practical attacks on Falcon-512 and Falcon-1024 implementations running on ARM Cortex-M4. Compared to state-of-the-art methods, SERP-SCA achieves an improvement of 22.11% in success rate for Falcon-512 and 18.64% for Falcon-1024, with remarkable time efficiency gains of around 195,369× and 183,175× for mantissa attacks, and 180× and 185× for exponent attacks, respectively. Honglin Shao, Jingdian Ming, Yuejun Liu, Yiwen Gao 0001, Yongbin Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Breaking the Shield: Novel Fault Attacks on CRYSTALS-Dilithium
Dixiao Du, Yuejun Liu, Yiwen Gao 0001, Jingdian Ming, Yongbin Zhou |
ACISP (2) | 3 |
| 2025 | Side-Channel Collision Attacks Against ASCONabstractSide-channel attack poses a significant threat to the security of electronic devices, particularly IoT/AIoT terminals. By leveraging side-channel leakages, collision attacks can efficiently extract the secret keys from cryptographic devices while requiring considerably less computational effort. In this paper, we investigate side-channel collision attacks against ASCON, a lightweight crypto designed for resource-constrained devices, which has been standardized by the NIST. For the first time, we propose a side-channel key recovery attack against ASCON by identifying the collisions in the linear diffusion layer. Using Pearson correlation coefficient and Euclidean distance for internal collision detections, our attack successfully recovers the secret key with approximately 5,000 power traces from an 8-bit software implementation on an AVR device. To further reduce attack complexity, we introduce a novel metric, Locally-Weighted Sum (LWS), which focuses on the most likely points of leakage, thereby decreasing the number of required power traces for successful attack. Our experiment on the same target demonstrates that the LWS-based collision attack can recover the full secret key with approximately 3,000 power traces, a reduction of 40 percent. Our study indicates that ASCON is susceptible to side-channel collision attacks, and bitslice implementations remain vulnerable to such threats. Hao Zhang 0009, Yiwen Gao 0001, Yongbin Zhou, Jingdian Ming |
DATE | 2 |
| 2025 | MEML-KEM: A Memory-Efficient Implementation of ML-KEM for IoT Devices
Ruiqi Hou, Yiwen Gao 0001, Yuejun Liu, Jingdian Ming, Yongbin Zhou |
ICA3PP (7) | 2 |
| 2025 | Parallel Cuckoo Hashing: Accelerating Secure Encrypted Data Search in Cloud EnvironmentsabstractCuckoo hashing serves as a fundamental technique in various privacy-enhancing cryptographic primitives, such as Private Information Retrieval, Symmetric Searchable Encryption, owing to its excellent performance. However, achieving space-efficient Cuckoo hashing that maintains fast insertion and query operations, while facilitating its applications with customized design, remains highly challenging. In this work, we propose a multi-segment permutation-based Cuckoo hashing (MS-PCH) that can be efficiently parallelized on multi-threaded platforms, followed by a strategy for further parallelization over the hashing within single segments (MH-MS-PCH). To demonstrate its practical utility, we then investigate its application on encrypted data search by constructing full-fledged Public Key (Authenticated) Encryption with Keyword Search schemes (PCH(-MD)-PEKS and PCH(-MD)-PAEKS). We evaluate the performance of these schemes on a public dataset, showing that, with 16 segments and 4 hash functions, both PCH(-MD)-PEKS and PCH(-MD)-PAEKS outperforms the plain PEKS and PAEKS across index generation, query, and update. Notably, our optimized Cuckoo hashing achieves up to a 10 times improvement over plain cuckoo hashing, and our enhanced P(A)EKS schemes demonstrate approximately a 6 times improvement in index generation efficiency and a 335 times acceleration in query processing. Hongyang Lin, Jiabei Wang, Tiancheng Zhu, Yiwen Gao 0001, Quan Yang, Yongbin Zhou |
ICCCN | 4 |
| 2025 | Masked Gadgets for Integer-Floating-Point Conversion with Applications to FalconabstractFalcon has been selected by the National Institute of Standards and Technology (NIST) as one of the standardized post-quantum digital signature algorithms. Unlike other candidate algorithms, Falcon relies heavily on floating-point operations, which are known to be vulnerable to side-channel attacks (SCAs). Although recent research proposes masking schemes for basic floating-point operations, secure conversion between integer and floating-point representations remains unaddressed. These conversions operate directly on the private key or other sensitive variables and therefore constitute a potential attack surface. This paper proposes new masking gadgets that securely perform conversions between integers and floating-point representations, addressing this previously unprotected surface. As part of the design, we optimize existing normalization and floating-point composition components to support the conversion process. We formally prove that our gadgets satisfy$t$-Non-Interference$(t-\text{NI})$or$t$-Strong Non-Interference ($t$-SNI) in the probing model and evaluate their side-channel resistance using Test Vector Leakage Assessment (TVLA) on an Arm Cortex-M4 processor. We also assess performance on an Intel Core CPU, where the optimized normalization and floating-point composition components demonstrate improvements of approximately 47.6 % and 10.8 %, respectively, over prior work. Jingdian Ming, Yuejun Liu, Yiwen Gao 0001, Yongbin Zhou |
ICCD | 4 |
| 2025 | A Versatile Decentralized Attribute Based Signature Scheme for IoT
Dazhi Xu, Yuejun Liu, Jiabei Wang, Yiwen Gao 0001, Yongbin Zhou |
ICICS (1) | 4 |
| 2025 | High-Performance Implementations of Classic McEliece KEM on GPUsabstractThis paper studies the high-performance implementation of Classic McEliece Key Encapsulation Mechanism (KEM), one of the candidates in the fourth round NIST Post-Quantum Cryptography competition. We focus on batch processing scenarios and improve the performance of key components in Classic McEliece through loop unrolling optimization. The study reveals that while loop unrolling is effective in improving performance, it may lead to performance degradation during the key encapsulation phase under high task loads. To address this issue, we introduce inline expansion optimization. We evaluate our approach on NVIDIA GeForce GTX 1660S and RTX 3090 platforms. The results show that loop unrolling optimization in the key encapsulation phase reduces latency by up to 54.09% (resp. 74.53%) at low task loads, while inline expansion optimization increases throughput by up to 16.17% (resp. 30.31%) at high task loads. In the key decapsulation phase, loop unrolling optimization reduces latency by up to 96.97% (resp. 96.54%) at low task loads and increases throughput by up to 58.64% (resp. 96.54%) at high task loads. Dingyan Xu, Yiwen Gao 0001, Yongbin Zhou, Jian Weng 0001 |
ISCAS | 2 |
| 2025 | CDSRNP: Cross-Domain Sequential Recommendation via Neural ProcessabstractCross-Domain Sequential Recommendation (CDSR) is a hot topic in sequence-based user interest modeling, which aims at utilizing a single model to predict the next items for different domains. To tackle the CDSR, many methods are focused on domain overlapped users’ behaviors fitting, which heavily relies on the same user’s different-domain item sequences collaborating signals to capture the synergy of cross-domain item-item correlation. Indeed, these overlapped users occupy a small fraction of the entire user set only, which introduces a strong assumption that the small group of domain overlapped users is enough to represent all domain user behavior characteristics. However, intuitively, such a suggestion is biased, and the insufficient learning paradigm in non-overlapped users will inevitably limit model performance. Further, it is not trivial to model non-overlapped user behaviors in CDSR because there are no other domain behaviors to collaborate with, which causes the observed single-domain users’ behavior sequences to be hard to contribute to cross-domain knowledge mining. Considering such a phenomenon, we raise a challenging and unexplored question: How to unleash the potential of non-overlapped users’ behaviors to empower CDSR? To this end, we propose a novel CDSR framework with Neural Processes (NP), briefly termed CDSRNP, where NP combines the advantages of meta-learning and stochastic processes. As a meta-learning based method, we first sample some observed overlapped users’ behaviors as the support set to empower query users’ prediction. Next, we employ the NP principle to align the cross-domain correlation prior/posterior distributions generated by support/query user sets, thus the query user (e.g., non-overlapped user) behaviors sequence could also establish a straight bridge to connect other domain items. Additionally, we design a fine-grained interest adaptive layer to identify the users’ interests to enhance prediction. Experimental results illustrate that CDSRNP1 outperforms state-of-the-art methods in two real-world datasets. Jiangxia Cao, Yiwen Gao 0001, Yunhuai Liu, Shuchao Pang |
SDM | 3 |
| 2025 | Multi-Channel Attacks on Dilithium: Bridging Power Gaps with EM-Guided Synthetic TracesabstractMulti-Channel Fusion Attacks (MCFAs) boosts the effectiveness of side-channel cryptanalysis by combining information leakages from multiple channels such as power and electromagnetic (EM) leakages. However, current MFCAs require the same number of traces across the channels, which poses a critical constraint that does not necessarily hold in practical scenarios. For instance, acquiring contactless EM traces is considerably easier than obtaining power traces. Consequently, an attacker can collect many more EM traces, but a considerable portion of them must be discarded due to this constraint. To address the constraint of requiring equal numbers of traces across channels while fully utilizing the otherwise discarded EM traces, we propose a novel multi-channel fusion framework. This framework leverages a machine learning model to synthesize the missing power traces, thereby aligning the number of traces in the power and EM channels. As a result, it maximizes the utilization efficiency of the available traces. Specifically, we introduce GTL, a hybrid generative adversarial network that integrates Transformer and LSTM architectures for effective trace synthesis. We present the first MCFAs against the postquantum cryptographic scheme CRYSTALS-Dilithium. Experimental results show that, compared with existing MCFAs, our approach reduces the required number of traces by more than 50%. Qikang Fan, Jingdian Ming, Yiwen Gao 0001, Yongbin Zhou |
TrustCom | 3 |
| 2025 | Fine-Grained Revocable Lattice-Based ABE: Dual User and Attribute Revocation with Low Overhead for Cloud EnvironmentsabstractAttribute-based encryption (ABE) provides fine-grained access control over encrypted data without restricting itself to a single access policy, making it applicable to diverse scenarios such as cloud environments. However, existing bilinear pairing-based ABE schemes are vulnerable to quantum attacks, while lattice-based ABE schemes typically lack flexible and efficient user or attribute revocation mechanisms. To address these challenges, this paper presents a fine-grained revocable attribute-based encryption (FR-ABE) scheme based on the ring learning with errors (RLWE) assumption. We introduce a two-dimensional attribute structure and an extended Shamir’s secret sharing method, which together support multi-valued attributes and flexible threshold access policies, while reducing storage and computational overhead. Furthermore, We have devised a fine-grained revocation mechanism that functions at both the attribute and user levels, thereby accommodating the frequent role and permission changes. The attribute authority centrally manages attribute revocation through an indirect revocation method, eliminating the need to re-run the sampling algorithm and instead relying solely on polynomial-level operations, which reduces time overhead. User revocation is implemented using a binary tree structure, which updates only the ciphertext and leaves all keys unchanged. Theoretical analysis suggests that our scheme performs well in terms of both computational and storage overhead, and experimental findings further corroborate its superiority. The security analysis establishes the scheme’s selective security under the RLWE assumption. Yuan Liu 0013, Yiwen Gao 0001, Yongbin Zhou, Licheng Wang 0004 |
TrustCom | 3 |
| 2025 | Traceable and Revocable Key-Policy Attribute-Based Encryption scheme from LatticesabstractABE as a powerful tool for secure fine-grained access control, has been widely adopted in data-sharing scenarios such as cloud computing. However, most existing traceable and revocable ABE schemes are constructed using bilinear pairings and are thus vulnerable to quantum attacks. Although lattice-based ABE schemes provide post-quantum security, they often lack critical features such as traitor tracing and timely revocation, which may lead to key abuse problem. Some revocable lattice-based schemes are unable to identify the source of leaked keys, while existing traceable and revocable lattice-based schemes also fail to resist collusion attacks between revoked and un-revoked users. In addition, the exposure of attribute values in access policies poses serious privacy risks. To solve the above issues, this paper proposes a traceable and revocable KP-ABE scheme based on RLWE assumption. The proposed scheme embeds the user identity into the key through a white-box tracing mechanism, enabling the system to detect and revoke malicious users in case of key leakage. It also protects attribute privacy by hiding attribute values in the access policy and resists collusion attacks by embedding random parameters in the key. Theoretical analysis and experimental evaluation results show that the proposed scheme ensures security and privacy while maintaining high computational efficiency and is suitable for post-quantum secure data sharing environments. Yanqi Ma, Yuan Liu 0013, Yongbin Zhou, Yiwen Gao 0001 |
TrustCom | 4 |
| 2025 | Fully-incremental public key encryption with adjustable timed-release keyword search
Tiancheng Zhu, Jiabei Wang, Yiwen Gao 0001, Yongbin Zhou, Jian Weng 0001 |
Inf. Sci. | 4 |
| 2024 | A Novel Power Analysis Attack against CRYSTALS-Dilithium ImplementationabstractPost-Quantum Cryptography (PQC) was proposed due to the potential threats quantum computer attacks against conventional public key cryptosystems, and four PQC algorithms besides CRYSTALS-Dilithium (Dilithium for short) have so far been selected for National Institute of Standards and Technology (NIST) standardization. However, the selected algorithms are still vulnerable to side-channel attacks in practice, and their physical security need to be further evaluated. This paper proposes two efficient power analysis attacks against Dilithium implementation, the optimized fast two-stage approach and the single-bit approach, aiming at reducing the key guess space. Our findings reveal that the optimized approach outperforms the conservative approach and the fast two-stage approach proposed in ICCD 2021 by factors of 338 and 49, respectively. Similarly, compared to these two approaches, the single-bit approach achieves acceleration of 367 times and 53 times, respectively. Yuejun Liu, Yongbin Zhou, Yiwen Gao 0001, Zehua Qiao, Huaxin Wang |
ETS | 4 |
| 2024 | MDTM: A Multi-dimensional Trust Management Scheme for Enhancing Security and Stability in SDNabstractAn accurate and efficient trust evaluation mechanism is the cornerstone of maintaining the stability of SDN networks, especially in face of the escalating threats like Distributed Denial of Service (DDoS) attacks. The traditional trust evaluation mechanisms in SDN often lack of adaptability and accuracy, due to that they typically rely on the direct trust or ignore the influence of indirect and historical trust factors which lead to inaccurate trust evaluation and lower efficiency. In order to solve these problems, we propose a novel multidimensional trust evaluation mechanism consisting of three sub-modules: Device Behavior Trust Evaluation Scheme (DBTES), Machine Learning-based Trust Evaluation Scheme (MLTES), and Bayesian-Based Trust Evaluation Scheme (BBTES). These modules work together to enhance the precision and adaptability of trust assessments by capturing various aspects of device behavior. Additionally, we introduce a trust fusion algorithm that combines the Analytical Hierarchy Process(AHP) with Criteria Importance Through Inter-criteria Correlation (CRITIC) to optimize weight distribution, further improving the accuracy and robustness of trust evaluations. Our approach overcomes the limitations of existing methods by providing a more comprehensive and adaptable trust assessment. Simulation results show that under varying Malicious Device Ratio conditions, the MDTM method improves the task success rate by up to 4.39% compared to traditional methods, along with an increase in average trust values by 5.16%. Additionally, with changing historical factors, MDTM further improves trust values by 5.33%. These results demonstrate the enhanced resilience and effectiveness of our approach in maintaining network stability and security within SDN environments. Tianrui Bai, Yuan Liu 0013, Yiwen Gao 0001, Yongbin Zhou |
HPCC | 3 |
| 2024 | An Efficient Flow Rule Conflict Comprehensive Detection Scheme for SDN NetworksabstractSoftware-Defined Networking (SDN) has introduced flexibility and efficiency to network management but also faces challenges from flow rule conflicts, including static, dynamic, and dependency conflicts. Existing detection algorithms often focus on a single conflict type, resulting in inefficiencies and high false positive rates. To address these issues, we propose a comprehensive flow rule conflict detection scheme that improves real-time detection of explicit conflicts (both static and dynamic) and reduces false positives in implicit (dependency) conflict detection. Specifically, we present a real-time explicit conflict detection algorithm based on the Protocol-Divided Trie (PDT), which categorizes flow rules by protocol type and uses a prefix tree for rapid matching. Experimental results show that this approach significantly reduces detection times by at least 39.2% and achieves 100% detection accuracy. Additionally, we propose a two-stage detection (TSD) algorithm that combines the precision of path-based detection (PBD) with the efficiency of alias set-based detection (ASD). Our experiments reveal a 51% reduction in false positives compared to ASD and a 48% reduction in detection time compared to PBD, while maintaining equivalent false positive rates. This approach provides a robust solution for conflict detection in SDN, improving network security and resource utilization efficiency. Yuan Liu 0013, Yongbin Zhou, Yiwen Gao 0001 |
ISPA | 4 |
| 2024 | Heterogeneous Performs Better: High Throughput Implementations of Falcon in Multi-Client ScenariosabstractThis paper investigates high-performance implementations of Falcon scheme, which is one of the four post-quantum cryptography algorithms standardized by NIST. We explore high-throughput implementation solutions by improving how critical components of Falcon execute on the GPU, with a focus on low latency requirements. Our research reveals that some components of Falcon cannot fully exploit the parallelism of the GPU. Consequently, we propose a heterogeneous parallel implementation of the Falcon signature scheme that utilizes the parallelism of both CPUs and GPUs, while mitigating the additional overhead introduced by heterogeneous computing. Finally, we conduct evaluations on three typical testbeds. The experimental results show that our improved GPU implementation is 222 percent faster for signing and 13.9 percent faster for verification in cloud scenarios compared to the state-of-the-art GPU implementations. On an embedded GPU platform (Jetson AGX Orin), our heterogeneous parallel implementation outperforms the CPU multi-threaded implementation by 46 percent and the GPU implementation by 297 percent. Quan Yang, Yiwen Gao 0001, Yuejun Liu, Jiabei Wang, Yongbin Zhou |
ISPA | 2 |
| 2024 | Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalabstractScene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP. Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015 |
ACM Multimedia | 6 |
| 2024 | Improving Interpretability: Visual Analysis of Deep Learning-Based Multi-channel Attacks
Ziyue Shen, Yiwen Gao 0001, Wei Cheng 0003, Jiabei Wang, Yongbin Zhou |
SecureComm (1) | 2 |
| 2024 | Solving ILWE Problem More Efficiently and Application to BLISS Side-Channel Attack
Yuejun Liu, Yiwen Gao 0001, Yongbin Zhou |
SecureComm (4) | 3 |
| 2024 | Towards High-Quality Electromagnetic Leakage Acquisition in Side-Channel AnalysisabstractSide-channel leakage acquisition plays a crucial role in side-channel analysis against cryptographic implementations, since it usually has a decisional impact on the security claims of the target devices. While most existing research has concentrated on power consumption acquisition settings, the exploration of electromagnetic (EM) radiation leakage acquisition remains limited and surprisingly under-discussed. In this study, we systematically investigate the parameter setting for EM leakage acquisition across two devices of different architectures. We specifically examine the effects of the amplifier, coupling mode, sampling rate, and EM probe on the quality of collected EM traces. For the first device STM32F405, our proposed optimal acquisition settings enhance the signal-to-noise ratio by a factor of 16 compared to the reference settings and reduce the number of EM traces required to achieve a success rate of 90% by a factor of 38. For the second device ATmega2560, the optimal settings improve the signal-to-noise ratio by a factor of 400 compared to the reference settings and reduce the number of EM traces required to achieve a success rate of 90% by a factor of 320. In summary, this work offers a comprehensive investigation into high-quality EM leakage acquisition. While some conclusions may be specific to certain devices, we believe that the proposed guidelines can be applied to EM trace acquisition in other devices as well. Xiaoran Huang, Yiwen Gao 0001, Wei Cheng 0003, Yuejun Liu, Jingdian Ming, Yongbin Zhou, Jian Weng 0001 |
TrustCom | 2 |
| 2024 | Attacking High-Performance SBCs: A Generic Preprocessing Framework for EMAabstractFor the high-performance single-board computers (SBCs) running an operating system, side-channel attacks usually come at high analytical costs, requiring millions of traces. This work uses the case of electromagnetic attacks on encryption services in real-world SBCs to explore how various preprocessing methods can reduce the attack costs for such devices. Specifically, we propose a general preprocessing framework that effectively combines multiple preprocessing methods based on their inherent characteristics. Utilizing this framework, we design a four-layer preprocessing scheme that significantly reduces the number of traces for key recovery. The experimental results show that, when the attack success rate reaches 80 percent, the proposed preprocessing scheme reduces the number of traces required by a factor of approximately 10 compared to the latest alignment algorithm on the Raspberry Pi 2B. Our research demonstrates the feasibility of conducting low-cost side-channel attacks on SBCs, further emphasizing the need to protect sensitive applications running on these devices. Debao Wang, Yiwen Gao 0001, Jingdian Ming, Yongbin Zhou |
TrustCom | 2 |
| 2024 | Enhancing Higher-Order Masking: A Faster and Secure Implementation to Mitigate Bit Interaction LeakageabstractHigher-order masking is an effective countermeasure against side-channel attacks but is often perceived as impractical due to its cost. Recent advancements in parallelization technologies have made higher-order masking more feasible. However, the security of these implementations can be jeopardized by the parallelization methods used, as their theoretical assurances may not hold in real-world scenarios. In this paper, we improve the security and efficiency of higher-order masking schemes for block ciphers through a refined bit-sliced implementation. Our method addresses lower-order leakage caused by bit interactions, which are prevalent in current schemes. We evaluated the efficiency of our approach across various widely used ARM and AVR platforms, demonstrating efficiency improvements across all test platforms. By optimizing the masking gadgets, our approach reduces the required clock cycles by 30% to 50% compared to the original implementation, across share numbers from 2 to 32 on the 32-bit ARM platform. We validate the enhanced security of our method through both theoretical analysis and practical leakage detection, proving its effectiveness against bit interaction leakage. Yuejun Liu, Jingdian Ming, Yiwen Gao 0001, Yongbin Zhou, Debao Wang |
TrustCom | 4 |
| 2024 | In-depth Correlation Power Analysis Attacks on a Hardware Implementation of CRYSTALS-DilithiumabstractAbstract During the standardisation process of post-quantum cryptography, NIST encourages research on side-channel analysis for candidate schemes. As the recommended lattice signature scheme, CRYSTALS-Dilithium, when implemented on hardware, has seen limited research on side-channel analysis, and current attacks are incomplete or requires a substantial quantity of traces. Therefore, we conducted a more complete analysis to investigate the leakage of an FPGA implementation of CRYSTALS-Dilithium using the Correlation Power Analysis (CPA) method, where with a minimum of 70,000 traces partial private key coefficients can be recovered. Furthermore, we optimise the attack by extracting Point-of-Interests using known information due to parallelism (named CPA-PoI) and by iteratively utilising parallel leakages (named CPA-ITR). Our experimental results show that CPA-PoI reduces the number of traces by up to 16.67%, CPA-ITR by up to 25%, and both increase the number of recovered key coefficients by up to 55.17% and 93.10% using the same number of traces. They outperfom the CPA method. As a result, it suggests that the FPGA implementation of CRYSTALS-Dilithium is more vulnerable than thought before to side-channel analysis. Huaxin Wang, Yiwen Gao 0001, Yuejun Liu, Qian Zhang 0042, Yongbin Zhou |
Cybersecur. | 2 |
| 2024 | MAFFN-SAT: 3-D Point Cloud Defense via Multiview Adaptive Feature Fusion and Smooth Adversarial TrainingabstractAdversarial attacks pose a significant threat to deep neural networks (DNNs) used for 3-D point cloud classification, especially in safety-critical applications. While previous works have proposed several defense model architectures and adversarial training strategies, they often either fall short in capturing the intricate geometric and topological aspects of point cloud data or grapple with challenges pertaining to model convergence. To solve these problems, in this article, we propose an innovative point cloud defense framework, called MAFFN-SAT, which contains a multiview adaptive feature fusion network (MAFFN) along with a smooth adversarial training (SAT) strategy. Specifically, we construct a multiview defense module to obtain multiview features in MAFFN, which uses geometric proximity and spatial queries to comprehensively explore the inherent characteristics of point cloud data. Subsequently, an adaptive feature fusion module is designed to integrate the multiview features. Furthermore, we introduce SAT, which uses an optimized regularization to measure the information divergence between two probability distributions, guiding the model to develop a smoother decision boundary, thereby more robust to adversarial attacks. Extensive experiments conducted on three benchmark datasets demonstrate the robustness of our approach against various attacks. Remarkably, our defense framework achieves 15.34% performance improvement under point dropping attacks on the ModelNet40 dataset. Our implementation:https://github.com/shenyu234/MAFFN-SAT. Anan Du, Jue Zhang 0001, Yiwen Gao 0001, Shuchao Pang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Recovering Multi-prime RSA Keys with Erasures and Errors
Yuejun Liu, Yongbin Zhou, Yiwen Gao 0001 |
ISPEC | 4 |
| 2022 | cuNH: Efficient GPU Implementations of Post-Quantum KEM NewHopeabstractPost-quantum cryptography was proposed in the past years due to the foreseeable emergence of quantum computers that are able to break the conventional public key cryptosystems at acceptable costs. However, post-quantum schemes are usually less efficient than conventional ones, which makes them less practical in scenarios with limited resources or high concurrency. Server-side applications always feature multiple users, therefore requiring efficient execution of batch tasks. GPU is intrinsically well-suited to batch tasks owing to its SIMD/SIMT execution fashion, so it naturally helps to achieve high performance. However, a naive GPU-based implementation cannot make the best use of hardware resources of the GPU regardless of task loads. In this article, we propose SIMD parallelization paradigms for fine-grained GPU implementations and then apply them to a post-quantum key encapsulation algorithm called NewHope, where we carefully design every module, especially NTT and inverse NTT, to fit into the SIMD parallelization paradigms. In addition, we employ multi-streaming to improve performance in user's perspective. Finally, our evaluations are made on two testbeds with GPU accelerators NVIDIA GeForce MX150 and GeForce GTX 1650, respectively. The experimental results show that the fine-grained implementations save up to 98 percent latency at low task loads, and their throughputs increase by up to 86 percent at high task loads, when compared with the naive ones in kernel's perspective, and the multi-streaming implementations greatly reduce the latency overhead percentage at high task loads by up to 86 percent, when compared with the fine-grained implementation in user's perspective. Moreover, our fine-grained implementation and multi-streaming implementation are respectively 51.5 and 45.5 percent faster than Gupta et al.'s implementations when compared with it under reasonable assumptions. Furthermore, as lattice-based post-quantum schemes have similar operations, our proposal also easily applies to other lattice-based post-quantum schemes. Yiwen Gao 0001, Jia Xu 0006 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Side-Channel Attacks With Multi-Thread Mixed LeakageabstractSide-channel attacks are one of the greatest practical threats to security-related applications, because they are capable of breaking ciphers that are assumed to be mathematically secure. Lots of studies have been devoted to power or electro-magnetic (EM) analysis against desktop CPUs, mobile CPUs (including ARM, MSP, AVR, etc) and FPGAs, but rarely targeted modern GPUs. Modern GPUs feature their special and specific single instruction multiple threads (SIMT) execution fashion, which makes their power/EM leakage more sophisticated in practical scenarios. In this article, we study side-channel attacks with leakage from SIMT systems, and propose leakage models suited to any SIMT systems and specifically to CUDA-enabled GPUs. Afterwards, we instantiate the models with a GPU AES implementation, which is also used for performance evaluations. In addition to the models, we provide optimizations on the attacks that are based on the models. To evaluate the models and optimizations, we run the GPU AES implementation on a CUDA-enabled GPU and, at the same time, collect its EM leakage. The experimental results show that the proposed models are more efficient and the optimizations are effective as well. Our study suggests that GPU-based cryptographic implementations may be much vulnerable to microarchitecture-based side-channel attacks. Therefore, GPU-specific countermeasures should be considered for GPU-based cryptographic implementations in practical applications. Yiwen Gao 0001, Yongbin Zhou |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Efficient electro-magnetic analysis of a GPU bitsliced AES implementationabstractAbstract The advent of CUDA-enabled GPU makes it possible to provide cloud applications with high-performance data security services. Unfortunately, recent studies have shown that GPU-based applications are also susceptible to side-channel attacks. These published work studied the side-channel vulnerabilities of GPU-based AES implementations by taking the advantage of the cache sharing among multiple threads or high parallelism of GPUs. Therefore, for GPU-based bitsliced cryptographic implementations, which are immune to the cache-based attacks referred to above, only a power analysis method based on the high-parallelism of GPUs may be effective. However, the leakage model used in the power analysis is not efficient at all in practice. In light of this, we investigate electro-magnetic (EM) side-channel vulnerabilities of a GPU-based bitsliced AES implementation from the perspective of bit-level parallelism and thread-level parallelism in order to make the best of the localization effect of EM leakage with parallelism. Specifically, we propose efficient multi-bit and multi-thread combinational analysis techniques based on the intrinsic properties of bitsliced ciphers and the effect of multi-thread parallelism of GPUs, respectively. The experimental result shows that the proposed combinational analysis methods perform better than non-combinational and intuitive ones. Our research suggests that multi-thread leakages can be used to improve attacks if the multi-thread leakages are not synchronous in the time domain. Yiwen Gao 0001, Yongbin Zhou, Wei Cheng 0003 |
Cybersecur. | 1 |
| 2018 | Electro-magnetic analysis of GPU-based AES implementationabstractIn this work, for the first time, we investigate Electro-Magnetic (EM) attacks on GPU-based AES implementation. In detail, we first sample EM traces using a delicate trigger; then, we build a heuristic leakage model and a novel leakage model to exploit the simultaneous EM leakages in parallel scenarios. After that, we evaluate the effectiveness of EM attacks on GPU-based AES implementation. Our evaluation results show that GPU-based AES implementation is vulnerable to EM attacks. This work also suggests that GPU-based AES implementation needs to be protected against EM attacks in real scenarios. Yiwen Gao 0001, Hailong Zhang 0001, Wei Cheng 0003, Yongbin Zhou, Yuchen Cao 0002 |
DAC | 1 |