Zhou Xuan

dblp:45/2569 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2023
0009-0000-5738-885XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2023 PEM: Representing Binary Program Semantics for Similarity Analysis via a Probabilistic Execution Model
abstract
Binary similarity analysis determines if two binary executables are from the same source program. Existing techniques leverage static and dynamic program features and may utilize advanced Deep Learning techniques. Although they have demonstrated great potential, the community believes that a more effective representation of program semantics can further improve similarity analysis. In this paper, we propose a new method to represent binary program semantics. It is based on a novel probabilistic execution engine that can effectively sample the input space and the program path space of subject binaries. More importantly, it ensures that the collected samples are comparable across binaries, addressing the substantial variations of input specifications. Our evaluation on 9 real-world projects with 35k functions, and comparison with 6 state-of-the-art techniques show that PEM can achieve a precision of 96% with common settings, outperforming the baselines by 10-20%.
Xiangzhe Xu, Zhou Xuan, Shiwei Feng 0002, Siyuan Cheng 0005, Yapeng Ye, Qingkai Shi, Guanhong Tao 0001, Zhuo Zhang 0002, Xiangyu Zhang 0001
ESEC/SIGSOFT FSE2
2023 Learning Approximate Execution Semantics From Traces for Binary Function Similarity
abstract
Detecting semantically similar binary functions – a crucial capability with broad security usages including vulnerability detection, malware analysis, and forensics – requires understanding function behaviors and intentions. This task is challenging as semantically similar functions can be compiled to run on different architectures and with diverse compiler optimizations or obfuscations. Most existing approaches match functions based on syntactic features without understanding the functions’ execution semantics. We presentTrex, a transfer-learning-based framework, to automate learning approximate execution semantics explicitly from functions’ traces collected via forced-execution (i.e., by violating the control flow semantics) and transfer the learned knowledge to match semantically similar functions. While it is known that forced-execution traces are too imprecise to be directly used to detect semantic similarity, our key insight is that these traces can instead be used to teach an ML model approximate execution semantics of diverse instructions and their compositions. We thus design a pretraining task, which trains the model to learn approximate execution semantics from the two modalities (i.e., forced-executed code and traces) of the function. We then finetune the pretrained model to match semantically similar functions. We evaluateTrexon 1,472,066 functions from 13 popular software projects, compiled to run on 4 architectures (x86, x64, ARM, and MIPS), and with 4 optimizations (O0-O3) and 5 obfuscations.Trexoutperforms the state-of-the-art solutions by 7.8%, 7.2%, and 14.3% in cross-architecture, optimization, and obfuscation function matching, respectively, while running 8× faster. Ablation studies suggest that the pretraining significantly boosts the function matching performance, underscoring the importance of learning execution semantics. Our case studies demonstrate the practical use-cases ofTrex– on 180 real-world firmware images,Trexuncovers 14 vulnerabilities not disclosed by previous studies. We release the code and dataset ofTrexathttps://github.com/CUMLSec/trex.
Kexin Pei, Zhou Xuan, Suman Jana, Baishakhi Ray
IEEE Trans. Software Eng.2
2022 Checkpointing and deterministic training for deep learning
abstract
Checkpointing and faithful replay are important for the training process of a Deep Learning (DL) model. It may improve productivity, model performance, robustness, and help security auditing. However, the inherent nondeterminism in training poses prominent challenges. Even with fixed random seeds, multiple runs of a same training pipeline may yield models whose performance varies by 20% percent. With existing infrastructural checkpointing support, developers cannot faithfully replay a training process. In this paper, we propose DETrain, a new solution to checkpointing and faithful execution/replay for long running DL training programs. We introduce a novel random number generation mechanism that can generate consistent random numbers in the presence of data parallelism. In addition, we devise a novel analysis that can determine a set of state variables that are necessary for faithful replay. These variables are either saved in a checkpoint or re-generated by fast forwarding, a selective execution technique. DETrain is evaluated on 13 PyTorch models and 16 Tensorflow models. It can deterministically execute these programs and replay from checkpoints with reasonable overhead. It also helps developers in diagnosing problems in training.
Xiangzhe Xu, Hongyu Liu 0005, Guanhong Tao 0001, Zhou Xuan, Xiangyu Zhang 0001
CAIN4
2022 NeuDep: neural binary memory dependence analysis
abstract
Determining whether multiple instructions can access the same memory location is a critical task in binary analysis. It is challenging as statically computing precise alias information is undecidable in theory. The problem aggravates at the binary level due to the presence of compiler optimizations and the absence of symbols and types. Existing approaches either produce significant spurious dependencies due to conservative analysis or scale poorly to complex binaries.
Kexin Pei, Dongdong She, Scott Geng, Zhou Xuan, Yaniv David, Suman Jana, Baishakhi Ray
ESEC/SIGSOFT FSE5