VLDB 2026 Research / reviewers in the wild / expert
Lihai Nie
dblp:181/9509
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-8569-3739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-TuningabstractFine-tuning-as-a-service, while commercially successful for Large Language Model (LLM) providers, exposes models to harmful finetuning attacks.As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing them from being used to perform malicious tasks.However, we highlight a critical flaw: the inherent general adaptability of LLMs allows them to easily bypass selective unlearning by rapidly relearning or repurposing their general capabilities for harmful tasks.To address this fundamental limitation, we propose a paradigm shift: instead of selective removal, we advocate for inducing model collapse, effectively forcing the model to "unlearn everything", specifically in response to updates characteristic of malicious adaptation.This collapse directly neutralizes the very general capabilities that attackers exploit, tackling the core issue unaddressed by selective unlearning.We introduce the Collapse Trap (CTRAP) as a practical mechanism to implement this concept conditionally.Embedded during alignment, CTRAP pre-configures the model's reaction to subsequent fine-tuning dynamics.If updates during fine-tuning constitute a persistent attempt to reverse safety alignment, the pre-configured trap triggers a progressive degradation of the model's core language modeling abilities, ultimately rendering it inert and useless for the attacker.Crucially, this collapse mechanism remains dormant during benign fine-tuning, ensuring the model's utility and general capabilities are preserved.1 Biao Yi, Tiansheng Huang, Baolei Zhang, Tong Li 0011, Lihai Nie, Zheli Liu, Li Shen 0008 |
ACL (1) | 5 |
| 2026 | Practical Poisoning Attacks against Retrieval-Augmented GenerationabstractLarge language models (LLMs) have demonstrated impressive natural language processing abilities but face challenges such as hallucination and outdated knowledge. Retrieval-Augmented Generation (RAG) has emerged as a state-of-the-art approach to mitigate these issues. While RAG enhances LLM outputs, it remains vulnerable to poisoning attacks. Recent studies show that injecting poisoned texts into the knowledge database can compromise RAG systems, but most existing attacks assume that the attacker can insert a sufficient number of poisoned texts per query to outnumber correct-answer texts in retrieval, an assumption that is often unrealistic. To address this limitation, we propose CorruptRAG, a practical poisoning attack against RAG systems in which the attacker injects only a single poisoned text, enhancing both feasibility and stealth. Extensive experiments conducted on multiple large-scale datasets demonstrate that CorruptRAG achieves higher attack success rates than existing baselines. Baolei Zhang, Zhuqing Liu, Lihai Nie, Tong Li 0011, Zheli Liu, Minghong Fang |
SACMAT | 4 |
| 2026 | Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) integrates external knowledge into large language models to improve response quality. However, recent work has shown that RAG systems are highly vulnerable to poisoning attacks, where malicious texts are inserted into the knowledge database to influence model outputs. While several defenses have been proposed, they are often circumvented by more adaptive or sophisticated attacks. This paper presents RAGOrigin, a black-box responsibility attribution framework designed to identify which texts in the knowledge database are responsible for misleading or incorrect generations. Our method constructs a focused attribution scope tailored to each misgeneration event and assigns a responsibility score to each candidate text by evaluating its retrieval ranking, semantic relevance, and influence on the generated response. The system then isolates poisoned texts using an unsupervised clustering method. We evaluate RAGOrigin across seven datasets and fifteen poisoning attacks, including newly developed adaptive poisoning strategies and multi-attacker scenarios. Our approach outperforms existing baselines in identifying poisoned content and remains robust under dynamic and noisy conditions. These results suggest that RAGOrigin provides a practical and effective solution for tracing the origins of corrupted knowledge in RAG systems. Our code is available at: https://github.com/zhangbl6618/RAG-Responsibility-Attribution Baolei Zhang, Haoran Xin 0002, Zhuqing Liu, Biao Yi, Tong Li 0011, Lihai Nie, Zheli Liu, Minghong Fang |
SP | 7 |
| 2026 | SIsomap: Secure Collaborative Manifold Learning with Reducing Communication CostsabstractSecure manifold learning on datasets distributed among multiple data owners can benefit or even spawn many applications. For example, multiple service providers can jointly fit low-dimensional embeddings of their users' network behavior data to improve the accuracy of anomaly detection while addressing their privacy concerns about the datasets. In this paper, we focus on a classic manifold learning technique, known as isometric mapping (Isomap), and propose SIsomap, the first secure, distributed manifold learning system. We construct SIsomap based on secret sharing techniques and introduce careful optimizations. In particular, we propose two communication-efficient secure building blocks that focus on top-k and all-pairs shortest paths computation, respectively, and reduce secure operations by leveraging the characteristics of Isomap. Experimental results on both synthetic and real-world datasets demonstrate that our secure top-k and all-pairs shortest paths protocols are respectively up to 13.6× and 1818.5× faster than the state-of-the-art methods, and SIsomap as a whole is 11.1× to 28.8× faster than the baseline solution. Peizhao Zhou, Xiaojie Guo 0004, Pinzhi Chen, Ranyang Liu, Lihai Nie, Tong Li 0011, Zheli Liu |
WWW | 5 |
| 2026 | Viper: Priority-Based High-Visibility Per-Flow Packet Sampling for SDNsabstractPacket sampling is crucial for managing datacenter networks, serving fault diagnosis, traffic measurement, and intrusion detection functions. However, traditional sampling techniques, such as those based on sketches or ports, either lack packet–level granularity or provide insufficient visibility, leading to functional performance degradation. Recent research has employed the software-defined networking (SDN) model to enable flow-based packet sampling. However, these approaches often introduce substantial control and computation overhead, limiting their scalability. This paper presents Viper, a novel priority-based, high-visibility per-flow packet sampling mechanism tailored to address these challenges. Specifically, Viper leverages existing priority-based traffic scheduling mechanisms to prioritize shorter flows over longer ones. Then, a logical centralized controller orchestrates sampling policies for packets of different priorities. In-depth analysis indicates that the orchestration performed by the controller significantly impacts Viper’s performance. Consequently, we model this process as a nonlinear optimization problem, seeking to maximize the utility of sampling. Then, we propose an online primal–dual interior–point algorithm to address this optimization problem and prove the algorithm’s convergence, optimality, and efficiency. Experimental results show that Viper increases visibility by 3.83% to 8.3%, with negligible control overhead and a substantial reduction in sampling load by at least 20.51%. Xiaodong Dong, Xiulong Liu 0001, Lihai Nie, Jiuwu Zhang, Yinglong Wang 0001 |
IEEE Trans. Computers | 3 |
| 2026 | Stinger: A Light-Weight Website Fingerprinting Defense Through Poisoning Packet SequencesabstractWebsite Fingerprinting (WF) attack can be mitigated throughrandom camouflageorpair camouflage.Random camouflageinserts random dummy packets into the traces according to pre-defined rules. It can be compromised easily by machine learning-based WF attacks.Pair camouflageobfuscates the distinguishing features of paired websites by inserting elaborated perturbations into raw traces, thereby misleading the attacker. It is costly in maintaining a perturbation generator for each pair of websites. Based on these insights, we proposeStinger, a novel data poisoning based WF defense, which enables effective defense against WF attacks with low bandwidth overhead and only maintains one generator for all websites.Stingerexploits the idea of poisoning by contaminating the model directly in such a way that the WF attacks only classify based on the inserted poison sequences, thus being low overhead and website independent. We experimentally evaluateStingerusing the DF and AWF datasets. The results show that Stinger improves the successful defending rate by an average of 20.37% and 22.83% while reducing overhead by 85.88% and 81.35%, respectively. Lihai Nie, Xiaodong Dong, Lili Shi, Laiping Zhao, Zheli Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Scalable Private k-Nearest Neighbors Search Based on Secret Sharingabstractk-nearest neighbors search (kNN) is a fundamental algorithm widely used in signal and image processing, recommendation systems, and pattern recognition. Due to the sensitivity of data privacy and laws, directly disclosing raw data for kNN search between entities is often unacceptable in many cases. In this paper, we propose a kNN search scheme based on secret sharing techniques. Our scheme preserves the privacy of both the data owner’s datasets and the user’s queries, except that the user learns the query results. We introduce a communication-efficient secure top-k protocol and construct two secure kNN search protocols: a protocol that searches kNN results exactly, and a more efficient approximate protocol that leverages clustering-based preprocessing. For the clustering-based protocol, we also design an efficient cluster retrieval approach and a compact kNN search phase. Experiments on three large real-world datasets containing 1M/10M samples with 96/128 dimensions demonstrate that the proposed kNN search protocols achieve a speedup of 2.77× to 10.67× compared to the state-of-the-art scheme. Peizhao Zhou, Lihai Nie, Zheli Liu |
TrustCom | 2 |
| 2025 | Slark: A Performance Robust Decentralized Inter-Datacenter Deadline-Aware Coflows Scheduling Framework With Local InformationabstractInter-datacenter network applications generate massive coflows for purposes, e.g., backup, synchronization, and analytics, with deadline requirements. Decentralized coflow scheduling frameworks are desirable for their scalability in cross-domain deployment but grappling with the challenge of information agnosticism for lack of cross-domain privileges. Current information-agnostic coflow scheduling methods are incompatible with decentralized frameworks for relying on centralized controllers to continuously monitor and learn from coflow global transmission states to infer global coflow information. Alternative methods propose mechanisms for decentralized global coflow information gathering and synchronization. However, they require dedicated physical hardware or control logic, which could be impractical for incremental deployment. This article proposes Slark, a decentralized deadline-aware coflow scheduling framework, which meets coflows’ soft and hard deadline requirements using only local traffic information. It eschews requiring global coflow transmission states and dedicated hardware or control logic by leveraging multiple software-implemented scheduling agents working independently on each node and integrating such information agnosticism into node-specific bandwidth allocation by modeling it as a robust optimization problem with flow information on the other nodes represented as uncertain parameters. Subsequently, we validate the performance robustness of Slark by investigating how perturbations in the optimal objective function value and the associated optimal solution are affected by uncertain parameters. Finally, we propose a firebug-swarm-optimization-based heuristic algorithm to tackle the non-convexity in our problem. Experimental results demonstrate that Slark can significantly enhance transmission revenue and increase soft and hard deadline guarantee ratios by 10.52% and 7.99% on average. Xiaodong Dong, Lihai Nie, Zheli Liu, Yang Xiang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | PKDGA: A Partial Knowledge-Based Domain Generation Algorithm for BotnetsabstractDomain generation algorithms (DGAs) can be categorized into three types:zero-knowledge,partial-knowledge, andfull-knowledge. While prior research merely focused onzero-knowledgeandfull-knowledgetypes, we characterize their anti-detection ability and practicality and find thatzero-knowledgeDGAs present low anti-detection ability againstdetectors, andfull-knowledgeDGAs suffer from low practicality due to the strong assumption that they are fullydetector-aware. Given these observations, we proposePKDGA, a partial knowledge-based domain generation algorithm with high anti-detection ability and high practicality.PKDGAemploys the reinforcement learning architecture, which makes it evolve automatically based only on the easily-observable feedback from detectors. We evaluatePKDGAusing a comprehensive set of real-world datasets, and the results demonstrate that it reduces the detection performance of existingdetectorsfrom 91.7% to 52.5%. We further applyPKDGAto theMiraimalware, and the evaluations show that the proposed method is quite lightweight and time-efficient. Lihai Nie, Xiaoyang Shan, Laiping Zhao, Keqiu Li |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Robust website fingerprinting through resource loading sequence
Changzhi Li, Lihai Nie, Laiping Zhao, Keqiu Li |
World Wide Web (WWW) | 2 |
| 2021 | RLTree: Website Fingerprinting Through Resource Loading Tree
Changzhi Li, Lihai Nie, Laiping Zhao |
NSS | 2 |
| 2021 | Robust Anomaly Detection Using Reconstructive Adversarial NetworkabstractDetecting abnormal service performance is significant for Internet-based service management and operation. Recent advances in anomaly detection methods prefer unsupervised learning algorithms since they can work without manually labelled data. However, existing unsupervised methods converge into suboptimal solutions due to their heuristic-based objectives. Moreover, they frequently rely on the strong assumption that noise follows a Gaussian distribution, and their detection accuracy is also highly sensitive to threshold settings. To detect anomalies precisely and robustly, we presentAdran, an unsupervised anomaly detection model that introduces adversarial learning into a reconstructive model, generating a reconstructive adversarial network with an anomaly detection-based training objective. It tolerates non-Gaussian noise by activating the discriminator with a non-smooth function. Our experimental results demonstrate thatAdranachieves an improvement of$\geq 32\%$over the state-of-the-art methods in terms ofF-score. Moreover, the robustness analysis demonstrates that it is reasonably easy and straightforward to set an appropriate threshold usingAdran. Lihai Nie, Laiping Zhao, Keqiu Li |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | Glad: Global And Local Anomaly DetectionabstractDetecting anomaly in images is challenging due to the high dimension nature of image data. While the previous learning-based anomaly detection approaches can detect a particular type of anomaly precisely, they often fail in detecting multiple types of abnormal samples simultaneously.We identify the two specific types of anomalies that can be precisely detected by either compress-based or reconstruction-based anomaly detection approaches, named global anomaly and local anomaly. We then propose Glad, an anomaly detector that can precisely detect both of them at the same time. Glad adopts a joint approach combining the density estimation and auto-encoder. Firstly, it designs a multimodal density estimation model to derive the latent representation probability for identifying the global anomaly. Then, it uses structural similarity to measure the reconstruction loss for characterizing local anomaly. Finally, both anomalies can be diagnosed according to the joint density of latent representation and reconstruction loss. Experimental results on public benchmark datasets demonstrate that Glad outperforms the state-of-the-art methods significantly. Lihai Nie, Laiping Zhao, Keqiu Li |
ICME | 1 |
| 2020 | XShot: Light-weight Link Failure Localization using Crossed Probing Cycles in SDNabstractAccurate and quick failure localization is critical for automatic network troubleshooting. While it is particularly difficult to solve the problem in the traditional network due to the uncertain routing, Software Defined Networking (SDN) enables the deterministic routing for packet transmission through the traffic engineering algorithm in the centralized controller. Hongyun Gao 0002, Laiping Zhao, Huanbin Wang, Lihai Nie, Keqiu Li |
ICPP | 5 |
| 2019 | Deeplive: QoE Optimization for Live Video Streaming through Deep Reinforcement LearningabstractA new broad of video services that support live streaming has become tremendously popular in recent years. Compared with traditional video-on-demand (VOD) services, live video streaming has much higher requirements on Quality-of-Experience (QoE), including low rebuffering, high definition, low latency and low bitrate oscillations. While previous adaptive bitrate algorithms (ABR) solely optimize bitrate for ensuring QoE of VOD, live video streaming has a larger decision space, making the optimization problem more difficult to solve. We propose Deeplive, which maximizes QoE through deep reinforcement learning (DRL), so it does not rely on fixed rules. To accelerate the training process of Deeplive, we further propose optimization including window completion with historical data and quick-start with rate-based algorithm. We compare Deeplive with other advanced ABR algorithms in a frame-level dynamic adaptive video streaming simulator using different network traces, QoE definitions, and video categories. In all experiments, we find that Deeplive not only has significant improvement in training time, but also shows an average of 15-55% improvement on QoE than the state-of-the-art ABR algorithms. Laiping Zhao, Lihai Nie, Peiqi Chen |
ICPADS | 3 |