Yunlong Xing

dblp:300/5803 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Empirical Study of Multi-language Security Patches in Open Source Software
Yunlong Xing, Grant Zou, Xinda Wang 0001, Kun Sun 0001
DIMVA (2)2
2025 Semantics-Guided Dynamic Hypergraph Network for Human Mobility Nowcasting in Disaster
abstract
Human mobility nowcasting is crucial for public safety, especially during disasters when human mobility significantly differs from normal patterns, posing unique challenges. Recent studies have shown a correlation between disaster-related social media information and abnormal patterns in human mobility. However, these studies mainly focus on text counts while neglecting semantic text, which limits the effective use of social media data and reduces model prediction performance. The social text semantics reveal inherent non-pairwise relationships between regions in human mobility, posing a challenge to traditional graph neural network approaches. Thus, we propose a Semantics-Guided Dynamic Hypergraph Convolutional Network (SG-DyHGCN) for human mobility nowcasting in disaster. The model leverages semantic information to guide dynamic hyper-graph construction, enabling flexible adjustments to the hyper-graph structure, effectively capturing non-pairwise relationships between regions, and enhancing prediction performance. Experimental results validate the effectiveness of our method.
Bowen Zhang 0005, Yunlong Xing, Zinao Su, Jinzhou Cao, Tianhong Zhao, Genan Dai
ICASSP2
2025 DISPATCH: Unraveling Security Patches from Entangled Code Changes
Yunlong Xing, Xinda Wang 0001, Shu Wang 0004, Qi Li 0002, Kun Sun 0001
USENIX Security Symposium2
2024 Poster: Repairing Bugs with the Introduction of New Variables: A Multi-Agent Large Language Model
abstract
Trained on billions of tokens, large language models (LLMs) have a broad range of empirical knowledge which enables them to generate software patches with complex repair patterns. We leverage the powerful code-fixing capabilities of LLMs and propose VarPatch, a multi-agent conversational automated program repair (APR) technique that iteratively queries the LLM to generate software patches by providing various prompts and context information. VarPatch focuses on the variable addition repair pattern, as previous APR tools struggle to introduce and use new variables to fix buggy code. Additionally, we summarize commonly used APIs and identify four repair patterns involving new variable addition. Our evaluation on the Defects4J 1.2 dataset shows that VarPatch can repair 69% more bugs than baseline tools and over 8 times more bugs than GPT-4.
Elisa Zhang, Yunlong Xing, Kun Sun 0001
CCS3
2024 What IF Is Not Enough? Fixing Null Pointer Dereference With Contextual Check
Yunlong Xing, Shu Wang 0004, Kun Sun 0001, Qi Li 0002
USENIX Security Symposium1
2024 A Hybrid System Call Profiling Approach for Container Protection
abstract
Over-privileged Linux containers might put the underlying OS at risk by permitting pointless system calls that could be exploited as entry points to the kernel. However, finding such security profiles is a difficult task as it demands examining the implementation/operation of containers in the absence of knowledge regarding its required system calls. In this article, we propose a hybrid approach to limit the system call usage during the execution of containers. Specifically, given an application container, we maintain an initial fine-grained whitelist by dynamic tracking to control the run-time security along with a complementary whitelist extracted via static analysis to maintain container's functionality while addressing the coverage limitation of dynamic analysis. Our method automatically analyzes the container behavior to identify three execution phases and dynamically enforce the corresponding fine-grained system call whitelists. The invoked system call will be compared with both whitelists to decide if it should be killed to guarantee the container security or logged for further analysis. Our evaluation results with 193 Docker images demonstrate the effectiveness of our approach in significantly reducing the required system calls during the applications' life-cycle. Furthermore, we discuss the reduced attack surface and demonstrate the efficiency of our approach through empirical analysis results.
Yunlong Xing, Xinda Wang 0001, Sadegh Torabi, Lingguang Lei, Kun Sun 0001
IEEE Trans. Dependable Secur. Comput.1
2023 Exploring Security Commits in Python
abstract
Python has become the most popular programming language as it is friendly to work with for beginners. However, a recent study has found that most security issues in Python have not been indexed by CVE and may only be fixed by "silent" security commits, which pose a threat to software security and hinder the security fixes to downstream software. It is critical to identify the hidden security commits; however, the existing datasets and methods are insufficient for security commit detection in Python, due to the limited data variety, non-comprehensive code semantics, and uninterpretable learned features. In this paper, we construct the first security commit dataset in Python, namely PySecDB, which consists of three subsets including a base dataset, a pilot dataset, and an augmented dataset. The base dataset contains the security commits associated with CVE records provided by MITRE. To increase the variety of security commits, we build the pilot dataset from GitHub by filtering keywords within the commit messages. Since not all commits provide commit messages, we further construct the augmented dataset by understanding the semantics of code changes. To build the augmented dataset, we propose a new graph representation named CommitCPG and a multi-attributed graph learning model named SCOPY to identify the security commit candidates through both sequential and structural code semantics. The evaluation shows our proposed algorithms can improve the data collection efficiency by up to 40 percentage points. After manual verification by three security experts, PySecDB consists of 1,258 security commits and 2,791 non-security commits. Furthermore, we conduct an extensive case study on PySecDB and discover four common security fix patterns that cover over 85% of security commits in Python, providing insight into secure software maintenance, vulnerability detection, and automated program repair.
Shu Wang 0004, Xinda Wang 0001, Yunlong Xing, Elisa Zhang, Kun Sun 0001
ICSME4
2023 Cross Container Attacks: The Bewildered eBPF on Clouds
Yi He 0020, Roland Guo, Yunlong Xing, Xijia Che, Kun Sun 0001, Zhuotao Liu, Ke Xu 0002, Qi Li 0002
USENIX Security Symposium3
2022 BinProv: Binary Code Provenance Identification without Disassembly
abstract
Provenance identification, which is essential for binary analysis, aims to uncover the specific compiler and configuration used for generating the executable. Traditionally, the existing solutions extract syntactic, structural, and semantic features from disassembled programs and employ machine learning techniques to identify the compilation provenance of binaries. However, their effectiveness heavily relies on disassembly tools (e.g., IDA Pro) and tedious feature engineering, since it is challenging to obtain accurate assembly code, particularly, from the stripped or obfuscated binaries. In addition, the features in machine learning approaches are manually selected based on the domain knowledge of one specific architecture, which cannot be applied to other architectures. In this paper, we develop an end-to-end provenance identification system BinProv, which leverages a BERT (Bidirectional Encoder Representations from Transformers) based embedding model to learn and represent the context semantics and syntax directly from the binary code. Therefore, BinProv avoids the disassembling step and manual feature selection in provenance identification. Moreover, BinProv can distinguish the compilers and the four optimization levels (O0/O1/O2/O3) by fine-tuning the classifier model with the embedding inputs for specific provenance identification tasks. Experimental results show that BinProv achieves 92.14%, 99.4%, and 99.8% accuracy at byte sequence, function, and binary levels, respectively. We further demonstrate that BinProv works well on obfuscated binary code, suggesting that BinProv is a viable approach to remarkably mitigate the disassembler dependence in future provenance identification tasks. Finally, our case studies show that BinProv can better identify compiler helper functions and improve the performance of binary code similarity detection.
Shu Wang 0004, Yunlong Xing, Pengbin Feng, Haining Wang 0001, Qi Li 0002, Songqing Chen, Kun Sun 0001
RAID3
2022 The devil is in the detail: Generating system call whitelist for Linux seccomp
Yunlong Xing, Jiahao Cao 0001, Kun Sun 0001, Fei Yan 0008, Shengye Wan
Future Gener. Comput. Syst.1