EDBT 2026 Demo / reviewers in the wild / expert
Zhongling He
dblp:271/5984
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CATI++: empirical study and evaluation for adjacent instruction enhanced type inferenceabstractAbstract Variable-type information is fundamental, and it greatly helps in understanding the program semantics. Previous work applies rule-based and machine learning-based methods to recover variable types from commercial off-the-shelf binaries, heavily relying on the data flow or control flow. However, according to our study, about half of the variables lacked or even had no data flow; this problem has not received much attention from previous work. We empirically explore the severity of this problem to the type inference task and analyze its root causes. Based on compilation properties, we find that the instructions surrounding the instructions that operate on variables provide good contextual information that can be used for co-encoding to overcome the above problem. In this paper, we present an effective machine learning-based method to infer variable types and overcome the challenge of limited data dependency via adjacent instructions co-encoding. Therefore, we implement a system called CATI++, which locates variables from stripped binaries and infers 19 types of variables. We evaluate CATI++ on different compilation options, all of which outperforms state-of-the-art methods. The ablation experiments verify that our scheme is not sensitive to compilation conditions, while our designed method effectively alleviates the problems caused by missing data dependency. Ligeng Chen, Zhongling He, Bing Mao 0001 |
Comput. J. | 2 |
| 2022 | AVMiner: Expansible and Semantic-Preserving Anti-Virus Labels Mining MethodabstractWith the increase in the variety and quantity of malware, there is an urgent need to speed up the diagnosis and analysis of malware. Extracting the malware family-related tokens from AV (Anti-Virus) labels, provided by online antivirus engines, paves the way for pre-diagnosing the malware. Automatically extracting vital information from AV labels will greatly enhance the detection ability of security enterprises and equip the research ability of security analysts. Recent works like AVCLASS and AVCLASS2 try to extract the attributes of malware from AV labels and establish the taxonomy based on expert knowledge. However, due to the uncertain trend of complicated malicious behaviors, the system needs the following abilities to face the challenge: preserving vital semantics, being expansible, and being free from expert knowledge. In this work, we present AVMiner, an expansible malware tagging system that can mine the most vital tokens from AV labels. AVMiner adopts natural language processing techniques and clustering methods to generate a sequence of tokens without expert knowledge ranked by importance. AVMiner can self-update when new samples come. Finally, we evaluate AVMiner on over 8,000 samples from well-known datasets with manually labeled ground truth, which outperforms previous works. Ligeng Chen, Zhongling He, Hao Wu 0067, Yuhang Gong, Bing Mao 0001 |
TrustCom | 2 |
| 2022 | DIComP: Lightweight Data-Driven Inference of Binary Compiler Provenance with High AccuracyabstractBinary analysis is pervasively utilized to assess software security and test vulnerabilities without accessing source codes. The analysis validity is heavily influenced by the inferring ability of information related to the code compilation. Among the compilation information, compiler type and optimization level, as the key factors determining how binaries look like, are still difficult to be inferred efficiently with existing tools. In this paper, we conduct a thorough empirical study on the binary's appearance under various compilation settings and propose a lightweight binary analysis tool based on the simplest machine learning method, called DIComP to infer the compiler and optimization level via most relevant features according to the observation. Our comprehensive evaluations demonstrate that DIComP can fully recognize the compiler provenance, and it is effective in inferring the optimization levels with up to 90% accuracy. Also, it is efficient to infer thousands of binaries at a millisecond level with our lightweight machine learning model (1MB). Ligeng Chen, Zhongling He, Hao Wu 0067, Fengyuan Xu, Bing Mao 0001 |
SANER | 2 |
| 2021 | RoBin: Facilitating the Reproduction of Configuration-Related VulnerabilityabstractVulnerability reproduction paves a way in debugging software failures, which need intensive manual efforts. However, some key factors (e.g., software configuration, trigger method) are often missing, so we can not directly reproduce the failure without extra attempts. Even worse, highly customized configuration options of programs create a barrier for reproducing the vulnerabilities that only appear under some specific combinations of configurations. In this paper, we address the problem mentioned above - reproducing the configuration-related vulnerability. We try to solve it by proposing a binary similarity-based method to infer the specific building configurations via the binary from crash report. The main challenges are as follows: precise compilation option inference, program configuration inference, and source-code-to-binary matching. To achieve the goal, we implement RoBin, a binary similarity-based building configuration inference tool. To demonstrate the effectiveness, we test RoBin on 21 vulnerable cases upon 4 well-known open-source programs. It shows a strong ability in pinpointing the building configurations causing the vulnerability. The result can help developers reproduce and diagnose the vulnerability, and finally, patch the programs. Ligeng Chen, Zhongling He, Dongliang Mu, Bing Mao 0001 |
TrustCom | 3 |
| 2020 | CATI: Context-Assisted Type Inference from Stripped BinariesabstractCode analysis is a powerful way to eliminate vulnerabilities. Closed-source programs lack crucial information vital for code analysis because that information is stripped on compilation to achieve smaller executable size. Restoration has always been a challenge for experts. Variable type information is fundamental in this process because it helps to provide a perspective on program semantic. In this paper, we present an efficient approach for inferring types, and we overcome the challenge of scattered information provided by static analysis on stripped binaries. We discover that neighboring instructions are likely to operate the same type of variables, which are leveraged to enrich the features that we rely on. Therefore, we implement a system called CATI, which locates variables from stripped binaries and infers 19 types from variables. Experiments show that it infers variable type with 71.2% accuracy on unseen binaries. Meanwhile, it takes approximately 6 seconds to process a typical binary. Ligeng Chen, Zhongling He, Bing Mao 0001 |
DSN | 2 |