VLDB 2026 Research / reviewers in the wild / expert
Xuezixiang Li
dblp:243/2006
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0005-9713-3815ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Effectiveness of Custom Transformers for Binary AnalysisabstractIn recent years, there has been increasing interest in using deep learning for binary analysis tasks. Particularly, Transformer-based pre-trained language models have attracted enormous attention and obtained encouraging results. Numerous research attempts modified the Transformer network architecture and designed new pre-training tasks explicitly tailored for individual downstream binary analysis tasks, and positive results were reported. However, it remains unclear whether these architectural changes and their associated pre-training tasks are beneficial to other downstream binary analysis tasks, and whether a vanilla Transformer model can perform equally well via fine-tuning.In order to provide guidance for future explorations in this direction, in this paper, we evaluate four custom Transformer-based models (i.e. jTrans, PalmTree, StateFormer, and Trex) and their pre-training tasks on four downstream applications. According to our evaluation results, we have the following surprising observations: aside from MLM (Masked Language Model), many existing pre-training tasks seem either too noisy or too challenging for the Transformer model to learn effectively; the vanilla BERT model is comparable or superior to these custom Transformers in all the four downstream applications. Moreover, our evaluation suggests that improvements in fine-tuning are generally more beneficial than introducing new pre-training tasks or making architectural modifications. Consequently, we conclude that recent architectural modifications and additional pre-training tasks for Transformer models may offer limited impact that does not sufficiently justify their associated costs. Xuezixiang Li, Lian Gao, Yu Qu, Heng Yin 0001 |
RAID | 1 |
| 2022 | MAB-Malware: A Reinforcement Learning Framework for Blackbox Generation of Adversarial MalwareabstractModern commercial antivirus systems increasingly rely on machine learning (ML) to keep up with the rampant inflation of new malware. However, it is well-known that machine learning models are vulnerable to adversarial examples (AEs). Previous works have shown that ML malware classifiers are fragile to the white-box adversarial attacks. However, ML models used in commercial antivirus (AV) products are usually not available to attackers and only return hard classification labels. Therefore, it is more practical to evaluate the robustness of ML models and real-world AVs in a pure black-box manner. We propose a black-box Reinforcement Learning (RL) based framework to generate AEs for PE malware classifiers and AV engines. It regards the adversarial attack problem as a multi-armed bandit problem, which finds an optimal balance between exploiting the successful patterns and exploring more varieties. Compared to other frameworks, our improvements lie in three points: 1) limiting the exploration space by modeling the generation process as a stateless process to avoid combination explosions, 2) reusing the successful payload in modeling; and 3) minimizing the changes on AE samples to correctly assign the rewards in RL learning (which also helps identify the root cause of evasions). As a result, our framework has much higher evasion rates than other off-the-shelf frameworks. Results show it has over 74%--97% evasion rate for two state-of-the-art ML detectors and over 32%--48% evasion rate for commercial AVs in a pure black-box setting. We also demonstrate that the transferability of adversarial attacks among ML-based classifiers is higher than that between ML-based classifiers and commercial AVs. Xuezixiang Li, Sadia Afroz 0001, Deepali Garg, Dmitry Kuznetsov, Heng Yin 0001 |
AsiaCCS | 2 |
| 2021 | PalmTree: Learning an Assembly Language Model for Instruction EmbeddingabstractDeep learning has demonstrated its strengths in numerous binary analysis tasks, including function boundary detection, binary code search, function prototype inference, value set analysis, etc. When applying deep learning to binary analysis tasks, we need to decide what input should be fed into the neural network model. More specifically, we need to answer how to represent an instruction in a fixed-length vector. The idea of automatically learning instruction representations is intriguing, but the existing schemes fail to capture the unique characteristics of disassembly. These schemes ignore the complex intra-instruction structures and mainly rely on control flow in which the contextual information is noisy and can be influenced by compiler optimizations. In this paper, we propose to pre-train an assembly language model called PalmTree for generating general-purpose instruction embeddings by conducting self-supervised training on large-scale unlabeled binary corpora. PalmTree utilizes three pre-training tasks to capture various characteristics of assembly language. These training tasks overcome the problems in existing schemes, thus can help to generate high-quality representations. We conduct both intrinsic and extrinsic evaluations, and compare PalmTree with other instruction embedding schemes. PalmTree has the best performance for intrinsic metrics, and outperforms the other instruction embedding schemes for all downstream tasks. Xuezixiang Li, Yu Qu, Heng Yin 0001 |
CCS | 1 |
| 2020 | DeepBinDiff: Learning Program-Wide Code Representations for Binary Diffing
Yue Duan, Xuezixiang Li, Heng Yin 0001 |
NDSS | 2 |