VLDB 2026 Research / reviewers in the wild / expert
Hailun Ding
dblp:230/7753
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2024
0009-0003-5190-339XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Madeline: Continuous and Low-cost Monitoring with Graph-free Representations to Combat Cyber ThreatsabstractAdvanced persistent threats (APTs) have caused significant financial losses for enterprises, making the development of effective detection systems a critical priority. While existing provenance graph-based APT defenses demonstrate high accuracy, the high complexity and cost of graph operations (e.g., construction, iteration) require extensive processing time and computational resources, making them impractical for real-time detection. To address this challenge, we introduce Madeline, a graph-free, lightweight APT detection system that leverages historical system statistics. Through a multi-step state score calculation for a set of behavioral attributes, MADE-LINE meticulously captures subtle, gradual system changes indicative of stealthy APT activities. Using an LSTM autoencoder, Madeline performs anomaly detection effectively without the need for prior knowledge or manual labeling of attacks. Additionally, Madeline supports continuous monitoring, enhancing the assessment of ongoing risks. This feature helps reduce the need for extensive investigative resources by prioritizing genuine high-risk periods. Our experiments show that Madeline achieves comparable detection accuracy with the state-of-the-art APT detection, with 0.996 recall and 0.011 false positive rate on average. Madeline also significantly reduces the computational overhead, with over 1000x reduction in processing time and 5x in memory utilization. Wenjia Song, Hailun Ding, Na Meng 0001, Danfeng Yao |
ACSAC | 2 |
| 2024 | Merlin: Multi-tier Optimization of eBPF Code for Performance and CompactnessabstracteBPF (extended Berkeley Packet Filter) significantly enhances observability, performance, and security within the Linux kernel, playing a pivotal role in various real-world applications. Implemented as a register-based kernel virtual machine, eBPF features a customized Instruction Set Architecture (ISA) with stringent kernel safety requirements, e.g., a limited number of instructions. This constraint necessitates substantial optimization efforts for eBPF programs to meet performance objectives. Despite the availability of compilers supporting eBPF program compilation, existing tools often overlook key optimization opportunities, resulting in suboptimal performance. In response, this paper introduces Merlin, an optimization framework leveraging customized LLVM passes and bytecode rewriting for Instruction Representation (IR) transformation and bytecode refinement. Merlin employs two primary optimization strategies, i.e., instruction merging and strength reduction. These optimizations are deployed before eBPF verification. We evaluate Merlin across 19 XDP programs (drawn from the Linux kernel, Meta, hXDP, and Cilium) and three eBPF-based systems (Sysdig, Tetragon, and Tracee, each comprising several hundred eBPF programs). The results show that all optimized programs pass the kernel verification. Meanwhile, Merlin can reduce number of instructions by 73% and runtime overhead by 60% compared with the original programs. Merlin can also improve the throughput by 0.59% and reduce the latency by 5.31%, compared to state-of-the-art technique K2, while being 106 times faster and more scalable to larger and more complex programs without additional manual efforts. Jinsong Mao, Hailun Ding, Juan Zhai, Shiqing Ma |
ASPLOS (3) | 2 |
| 2023 | The Case for Learned Provenance Graph Storage Systems
Hailun Ding, Juan Zhai, Dong Deng 0001, Shiqing Ma |
USENIX Security Symposium | 1 |
| 2023 | AIRTAG: Towards Automated Attack Investigation by Unsupervised Learning with Log Texts
Hailun Ding, Juan Zhai, Yuhong Nan, Shiqing Ma |
USENIX Security Symposium | 1 |
| 2022 | Training with More Confidence: Mitigating Injected and Natural Backdoors During TrainingabstractThe backdoor or Trojan attack is a severe threat to deep neural networks (DNNs). Researchers find that DNNs trained on benign data and settings can also learn backdoor behaviors, which is known as the natural backdoor. Existing works on anti-backdoor learning are based on weak observations that the backdoor and benign behaviors can differentiate during training. An adaptive attack with slow poisoning can bypass such defenses. Moreover, these methods cannot defend natural backdoors. We found the fundamental differences between backdoor-related neurons and benign neurons: backdoor-related neurons form a hyperplane as the classification surface across input domains of all affected labels. By further analyzing the training process and model architectures, we found that piece-wise linear functions cause this hyperplane surface. In this paper, we design a novel training method that forces the training to avoid generating such hyperplanes and thus remove the injected backdoors. Our extensive experiments on five datasets against five state-of-the-art attacks and also benign training show that our method can outperform existing state-of-the-art defenses. On average, the ASR (attack success rate) of the models trained with NONE is 54.83 times lower than undefended models under standard poisoning backdoor attack and 1.75 times lower under the natural backdoor attack. Our code is available at https://github.com/RU-System-Software-and-Security/NONE. Zhenting Wang, Hailun Ding, Juan Zhai, Shiqing Ma |
NeurIPS | 2 |
| 2022 | Rethinking the Reverse-engineering of Trojan TriggersabstractDeep Neural Networks are vulnerable to Trojan (or backdoor) attacks. Reverse-engineering methods can reconstruct the trigger and thus identify affected models. Existing reverse-engineering methods only consider input space constraints, e.g., trigger size in the input space.Expressly, they assume the triggers are static patterns in the input space and fail to detect models with feature space triggers such as image style transformations. We observe that both input-space and feature-space Trojans are associated with feature space hyperplanes.Based on this observation, we design a novel reverse-engineering method that exploits the feature space constraint to reverse-engineer Trojan triggers. Results on four datasets and seven different attacks demonstrate that our solution effectively defends both input-space and feature-space Trojans. It outperforms state-of-the-art reverse-engineering methods and other types of defenses in both Trojaned model detection and mitigation tasks. On average, the detection accuracy of our method is 93%. For Trojan mitigation, our method can reduce the ASR (attack success rate) to only 0.26% with the BA (benign accuracy) remaining nearly unchanged. Our code can be found at https://github.com/RU-System-Software-and-Security/FeatureRE. Zhenting Wang, Kai Mei, Hailun Ding, Juan Zhai, Shiqing Ma |
NeurIPS | 3 |
| 2021 | ELISE: A Storage Efficient Logging System Powered by Redundancy Reduction and Representation Learning
Hailun Ding, Shenao Yan, Juan Zhai, Shiqing Ma |
USENIX Security Symposium | 1 |