Yang Li 0192

dblp:37/4190-192 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-0749-0028ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
abstract
Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of backdoor triggers has evolved from fixed triggers to dynamic or implicit triggers. This increased flexibility in trigger design makes it challenging for defenders to accurately identify their specific forms. Most existing backdoor defense methods are limited to specific types of triggers or rely on an additional clean model for support. To address this issue, we propose a backdoor detection method based on attention similarity, enabling backdoor detection without prior knowledge of the trigger. Our study reveals that models subjected to backdoor attacks exhibit unusually high similarity among attention heads when exposed to triggers. Based on this observation, we propose an attention safety alignment approach combined with head-wise fine-tuning to rectify potentially contaminated attention heads, thereby effectively mitigating the impact of backdoor attacks. Extensive experimental results demonstrate that our method significantly reduces the success rate of backdoor attacks while preserving the model’s performance on downstream tasks.
Haotian Jin, Yang Li 0192, Haihui Fan, Xiangfang Li, Bo Li 0063
AAAI2
2026 MithrilRB: Resource-Efficient Redactable Blockchain With Single-Use Authorization
abstract
Redactable blockchains preserve the integrity of hash links while enabling authorized redactions to comply with regulatory requirements. However, existing permissioned solutions suffer from three severe issues. First, fine-grained privilege control incurs significant storage overhead, especially in the case of single-use authorization. Second, the reliance on bilinear pairings in chameleon hash leads to significant performance degradation when handling large-scale redaction requests. Finally, multiple incorporated components often unconsciously introduce centralized entities, undermining the decentralized nature. In this paper, we propose MithrilRB, a resource-efficient and decentralized redactable blockchain with single-use authorization. Specifically, we introduce a privilege control mechanism with our proposed multi-authority attribute-based signature (MA-ABS) and the threshold BLS signature, achieving fine-grained single-use authorization and direct user revocation without extra ciphertext storage. We also design a pairing-free non-interactive threshold chameleon hash (PNITCH), which enhances efficiency and is better suited for large-scale redaction requests. Moreover, MithrilRB eliminates centralized trust points that hold secret information, ensuring fully decentralization-compatible functional integration in redactable blockchains. Finally, we implement MithrilRB, and the experimental results demonstrate that MithrilRB significantly outperforms existing solutions in both computational efficiency and storage requirements.
Tianming Hou, Hui Ma 0002, Jinchao Zhang 0002, Yang Li 0192, Bo Li 0063, Weiping Wang 0005
IEEE Trans. Inf. Forensics Secur.4
2025 Dangerous Language Habits! Exploiting Code-Mixing for Backdoor Attacks on NLP Models
abstract
Backdoor attacks threaten the reliability of NLP models by embedding hidden behaviors during training, which are activated by specific inputs at inference time. Traditional backdoor triggers often rely on explicit content alterations-such as token insertion or stylistic modification-which may compromise semantic coherence and be easily detected.In this work, we propose a novel backdoor attack strategy that leverages the linguistic properties of code-mixing(a language form that combines elements from two or more languages) as implicit triggers. Drawing inspiration from natural code-mixing communication, we design three types of linguistically grounded triggers: inter-word mixing, intra-sentential mixing, and inter-sentential mixing. These forms reflect realistic language usage patterns in bilingual communities, enhancing the stealthiness of the attack. The experiment results show that existing NLP models perform poorly when faced with backdoor attacks based on code-mixing triggers. We are the first to focus on code-mixing as a trigger for text backdoor attacks. We hope this research raises awareness of the vulnerability of models during training when faced with code-mixing.
Haotian Jin, Haihui Fan, Jinchao Zhang 0002, Yang Li 0192, Bo Li 0063, Junhao Zhou
CIKM4
2025 An Efficient and Privacy-Preserving Cross-Modal Retrieval Scheme for Encrypted Data in the Cloud
abstract
The rapid development and widespread adoption of cloud computing have made privacy-preserving data retrieval in the cloud a hot topic in privacy computing research. While numerous schemes have been proposed for privacy-preserving text-to-text and image-to-image retrieval, there is a lack of research on cross-modal retrieval with privacy protection in the cloud, leaving ample room for enhancing the security and efficiency of existing schemes. To address these challenges, this paper proposes a scheme for privacy-preserving cross-modal retrieval. The proposed scheme leverages a cross-modal pretrained model to extract features from both text and images, which are then mapped into hash buckets using locality-sensitive hashing (LSH), thereby enhancing retrieval efficiency. Additionally, a data structure is employed to generate secure indexes for these hash buckets via symmetric searchable encryption (SSE). To further improve retrieval accuracy, a secure k-nearest neighbor (kNN) algorithm is applied for second-round ranking of the retrieval results. Analysis and extensive experiments on widely-used real-world datasets have demonstrated the scheme's effectiveness in preserving privacy while maintaining efficiency and feasibility.
Yang Li 0192, Bo Li 0063, Jinchao Zhang 0002, Chuanrong Li
CSCWD2
2024 ELSEIR: A Privacy-Preserving Large-Scale Image Retrieval Framework for Outsourced Data Sharing
abstract
Privacy-preserving content-based image retrieval aims to safeguard the security of outsourced private images while maintaining their searchability. However, existing schemes encounter challenges in striking a balance between security, accuracy, and efficiency, as well as difficulties in scaling to large-scale image retrieval in multi-user settings. In this paper, we propose a novel Efficient Large-Scale Encrypted Image Retrieval (ELSEIR) framework for outsourced data sharing. We first utilize a deep hashing model for image feature extraction. Building upon this, we design an irreversible random hash code generation method that incorporates permutation keys for personalized access and integrates differential privacy to further enhance data security. In our multi-user implementation, we distribute the switch keys to the cloud to standardize each key, enabling the accurate search. In addition, we have theoretically proven that our ELSEIR guarantees both outsourced data security and query user privacy. Extensive experiments on real-world datasets demonstrate that our ELSEIR yields comparable accuracy to the unprotected baseline while outperforming existing methods in terms of both retrieval accuracy and efficiency.
Haihui Fan, Xiaoyan Gu 0001, Yang Li 0192, Bo Li 0063
ICMR4
2024 Toward Automated Field Semantics Inference for Binary Protocol Reverse Engineering
abstract
Network protocol reverse engineering is the basis for many security applications. A common class of protocol reverse engineering methods is based on the analysis of network message traces. After performing message field identification by segmenting messages into multiple fields, a key task is to infer the semantics of the fields. One of the limitations of existing field semantics inference methods is that they usually infer semantics for only a few fields and often require a lot of manual effort. In this paper, we propose an automated field semantics inference method for binary protocol reverse engineering (FSIBP). FSIBP aims to automatically learn semantics inference knowledge from known protocols and use it to infer the semantics of any field of an unknown protocol. To achieve this goal, we design a feature extraction method that can extract features of the field itself and of the field context. We also propose a semantic category aggregation method that abstracts the fine-grained semantics of all fields of known protocols into aggregated semantic categories. Moreover, we make FSIBP infer semantics based on the similarity of fields to semantic categories. The above design enables FSIBP to utilize the semantic knowledge of all fields of known protocols and infer the semantics of any fields of unknown protocols. The whole process of FSIBP does not require any expert knowledge or manual parameter setting. We conduct extensive experiments to demonstrate the effectiveness of FSIBP. Moreover, we find a utility for FSIBP besides field semantics inference, its output can help to detect the mis-segmented fields generated during the message field identification.
Mengqi Zhan, Yang Li 0192, Bo Li 0063, Jinchao Zhang 0002, Chuanrong Li, Weiping Wang 0005
IEEE Trans. Inf. Forensics Secur.2
2023 GuardBox: A High-Performance Middlebox Providing Confidentiality and Integrity for Packets
abstract
The deepening of digital transformation has led to an increasing amount of data from industries being transmitted over the Internet. However, packets in plaintext originally designed for transmission in private networks suffer from significant security threats on the Internet. Unfortunately, existing encryption schemes, such as the representative TLS, are difficult to be applied to these industrial protocols due to their specific requirements and conditions such as low latency requirements and restricted operating environments. In this paper, we present a high-performance encryption/decryption middlebox called GuardBox to provide confidentiality and integrity for packets. GuardBox is expected to transparently encrypt/decrypt packets sent/received by protected industrial equipment with low latency and supports almost any application-layer protocol. To do that, we design a high-performance packet I/O framework and an optimized encryption/decryption scheme for GuardBox. More importantly, we use commodity trusted hardware, Intel SGX, to ensure the security of keys and the encryption/decryption process. Our extensive evaluation demonstrates that GuardBox can provide confidentiality and integrity for packets transmitted over the Internet with low latency and a near-native throughput.
Mengqi Zhan, Yang Li 0192, Guangxi Yu, Yan Zhang 0014, Bo Li 0063, Weiping Wang 0005
IEEE Trans. Inf. Forensics Secur.2
2023 Website-Aware Protocol Confusion Network for Emergent HTTP/3 Website Fingerprinting
abstract
Website fingerprinting is exploited to analyze encrypted traffic traces and infer the visited website. Existing website fingerprinting methods can achieve satisfying performance for the HTTP traffic visiting websites over TCP. Recently, a new protocol QUIC has been proposed, and HTTP-over-QUIC has been formalized as the next generation HTTP, named HTTP/3. Thus, it is necessary to classify HTTP/3 traces. However, since HTTP/3 is newly proposed and is being deployed, it is difficult to collect a large number of HTTP/3 traces. Intuitively, we can use sufficient TCP traces to improve the performance of the QUIC trace classifier. Unfortunately, the protocol discrepancy exists between TCP and QUIC traces, which undermines the generalization ability of the classifier. In this paper, for practical website fingerprinting of HTTP/3, we propose a Website-Aware Protocol Confusion Network (WAPCN), which exploits only a few QUIC traces to train a website classifier with the help of lots of available TCP traces. It consists of four main parts: a feature extractor, a website classifier, a protocol discriminator, and a website-aware adaptor. The feature extractor aims to extract trace representations from both TCP and QUIC traces. It cooperates with the website classifier to learn the discriminative representation for the website classification. The role of the protocol discriminator is to confuse protocols and guide the feature extractor to learn protocol-invariant representations. The website-aware adaptor can enhance protocol-invariant representations to be aware of the website classification boundary. Extensive experiments are conducted on various tasks to demonstrate the effectiveness of WAPCN.
Mengqi Zhan, Yang Li 0192, Yongchun Zhu, Guangxi Yu, Yan Zhang 0014, Bo Li 0063, Weiping Wang 0005
IEEE Trans. Inf. Forensics Secur.2
2023 Coda: Runtime Detection of Application-Layer CPU-Exhaustion DoS Attacks in Containers
abstract
Denial of service (DoS) attacks have increasingly exploited vulnerabilities in algorithms or implementation methods in application-layer programs. In this type of attack, called CPU-exhaustion DoS attack, a few well-crafted requests may consume a lot of server resources, which is essentially different from traditional volumetric DoS attacks. Due to the lack of recognizable patterns, the traditional network-layer defense mechanism is usually unable to detect such sophisticated DoS attacks. In this paper, we proposeCoda, a framework for detecting application-layer CPU-exhaustion DoS attacks in containers.Codamonitors the CPU time consumed by each connection and uses statistical methods to detect attacks. It traces system calls and other related information from the container based on Linux eBPF at the host level. Some specific system calls are used to indicate the establishment and closure of the connection, which in turn indicate the start/end of the request processing. After triggering these specific system calls,Codastarts/ends monitoring the CPU time consumed by a connection. An attack can be detected when the CPU time consumed by an attack connection is statistically different from that consumed by a legitimate connection.Codahas the following key advantages. First, it works with programs built in different programming languages. Second, it remains agnostic to the source code of protected programs. Third, it supports monitoring the container and is transparent to the container. Through evaluation of real-world attacks, we demonstrate thatCodacan accurately detect ongoing application-layer CPU-exhaustion DoS attacks with low additional overhead.
Mengqi Zhan, Yang Li 0192, Huiran Yang, Guangxi Yu, Bo Li 0063, Weiping Wang 0005
IEEE Trans. Serv. Comput.2
2022 Detecting DNS over HTTPS based data exfiltration
Mengqi Zhan, Yang Li 0192, Guangxi Yu, Bo Li 0063, Weiping Wang 0005
Comput. Networks2
2020 NSAPs: A novel scheme for network security state assessment and attack prediction
Mengqi Zhan, Yang Li 0192, Xinghua Yang, Yulin Fan
Comput. Secur.2
2019 Towards Homograph-Confusable Domain Name Detection Using Dual-Channel CNN
Guangxi Yu, Xinghua Yang, Yan Zhang 0014, Huajun Cui, Huiran Yang, Yang Li 0192
ICICS6
2019 Mitigating Negative Impacts on DNS Caches Caused by Disposable Domain Names
abstract
DNS caches play an important role in DNS querying. However, the performance of DNS caches will be remarkably influenced by disposable domain names, which are generated by services of cloud storage, social networks, etc., and belong to a new class of misused case of DNS. In this paper, we proposed a novel solution named DC3(Domain Classification and Cascade Cache) to mitigate the negative impact. Domain Classification adopts a classifier which is based on a long short-term memory (LSTM) network to prevent disposable domains from being cached. Cascade Cache is a refined cascade LRU policy considering cache size allocation to process the remaining disposable domains. By querying the real DNS traces collected from a large ISP network, experiment results show that this solution can detect disposable domain names and mitigate their negative impacts on DNS caches effectively. Specifically, in our dataset, 67.4% of all distinct domain names are detected as disposable domain names. Correspondingly, when getting rid of them by using this solution, we can raise the cache hit rate more than double.
Guangxi Yu, Yan Zhang 0014, Huajun Cui, Xinghua Yang, Yang Li 0192
ISCC5
2017 Provably secure cloud storage for mobile networks with less computation and smaller overhead
Rui Zhang 0002, Hui Ma 0002, Yao Lu 0002, Yang Li 0192
Sci. China Inf. Sci.4