VLDB 2026 Research / reviewers in the wild / expert
Jingdong Guo
dblp:343/9945
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0005-1300-0987ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationabstractBinary Code Similarity Detection (BCSD), essential for binary-code related tasks like vulnerability detection, has attracted increasing attention in recent years. However, existing methods frequently fall short of achieving both high precision and recall at scale, and their results often lack interpretability due to the neglect of function context and reliance on purely similarity-driven outputs. Our key insights are twofold: 1) Binary functions are not self-contained; they depend on other code and data beyond their content to fulfill their functionalities. 2) Large language models (LLMs) excel not only at analyzing code but also at generating reasonable explanations. Motivated by these insights, we propose a general BCSD framework, Co2F uLL. We first systematically select stable and representative code and data features, along with their corresponding dependencies on the functions, to construct the function context. Then, by fusing function context with content similarities computed by the existing BCSD approach, we substantially narrow down the search space. Ultimately, we employ LLMs with a carefully designed prompt to verify the remaining candidates and produce clear, human-readable explanations. We conduct comprehensive experiments on a large function pool under varying compilation settings and after binary stripping. The results show that Co2F uLL based on HermesSim and DeepSeek-V3 achieves 80.5% precision and 94.4% recall, improving the baseline HermesSim by 142.5% and 42.2%, respectively, providing an accurate and interpretable solution for BCSD. Chaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li 0008, Dongliang Fang, Yang Xiao 0011, Yongle Chen, Limin Sun 0001 |
ASE | 2 |
| 2025 | Lares: LLM-driven Code Slice Semantic Search for Patch Presence TestingabstractIn modern software ecosystems, 1-day vulnerabilities pose significant security risks due to extensive code reuse. Identifying vulnerable functions in target binaries alone is insufficient; it is also crucial to determine whether these functions have been patched. Existing methods, however, suffer from limited usability and accuracy. They often depend on the compilation process to extract features, requiring substantial manual effort and failing for certain software. Moreover, they cannot reliably differentiate between code changes caused by patches or compilation variations.To overcome these limitations, we propose Lares, a scalable and accurate method for patch presence testing. Lares introduces Code Slice Semantic Search, which directly extracts features from the patch source code and identifies semantically equivalent code slices in the pseudocode of the target binary. By eliminating the need for the compilation process, Lares improves usability, while leveraging large language models (LLMs) for code analysis and SMT solvers for logical reasoning to enhance accuracy. Experimental results show that Lares achieves superior precision, recall, and usability. Furthermore, it is the first work to evaluate patch presence testing across optimization levels, architectures, and compilers. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/Lares. Siyuan Li 0014, Yaowen Zheng, Hong Li 0004, Jingdong Guo, Chaopeng Dong, Chunpeng Yan, Weijie Wang 0005, Yimo Ren, Limin Sun 0001, Hongsong Zhu |
ASE | 4 |
| 2024 | Precise and Efficient Third-party Java Libraries Identification Tool for Collaborative SoftwareabstractCollaborative systems frequently depend on various software components, like third-party libraries (TPLs), to execute their functions and expedite the development of the system. The security of an entire collaboration system can be compromised by a TPL that is vulnerable, particularly in an industrial setting. Unfortunately, current TPL detection tools encounter difficulties in precisely identifying version levels and exhibit inefficiency in detecting TPLs on a large scale.To address these challenges, we recommend JHunter, a precise and efficient tool for detecting TPL version details. Our approach involves introducing a novel concept called the attribute class dependency graph (ACDG) as a feature at the package level for TPLs. We then utilise a graph neural network-based method to compare the similarity of ACDGs and identify a list of candidate TPLs. Later, we use more detailed class-level features, such as Control Flow Graphs (CFGs), and constant features to determine version-specific information. We collected 19,095 different versions of TPLs from Maven to build our feature database. Our analysis demonstrates the effectiveness of JHunter on a real-world dataset, achieving F1 scores of 99.34% and 97.28% at the library and version levels, respectively, surpassing previous state-of-the-art (SOTA) results. Hongtu Zhang, Jingdong Guo, Laile Xi, Sidy Tambadou, Fang Zuo, Hong Li 0004 |
CSCWD | 3 |
| 2024 | Boosting Multimode Ruling in DHR Architecture With Metamorphic RelationsabstractABSTRACT The DHR architecture provides a revolutionary security defense structure for cyberspace. The multimode ruling in DHR is expected to alleviate the oracle problem, which still suffers from the existence of common model vulnerability. In this work, we design a test segmentation method to transform multimode ruling to a metamorphic testing problem. The text test input that causes inconsistency of heterogeneous executors is converted to a condition set, and we extract subsets of conditions based on its syntax tree. The original test can exploit a specific vulnerability, the follow‐up tests are composed by different subsets of conditions within the original test. We collect the execution matrix for the follow‐up tests to analyse the impact of each subset of conditions on ruling decision. Metamorphic relations are extracted based on the localization of independent condition, that is, the subsets of conditions that can impact ruling decision independently. The executors in an inconsistent ruling should be examined with metamorphic testing methods, rather than traditional majority voting mechanism. The proposed test segmentation and improved multimode ruling methods are evaluated on two DHR‐based cases, SQL injection in cyber‐range system and deserialization attack in ‐ project. The experimental results show that our test segmentation can help to locate malicious expressions and the metamorphic testing‐based multimode ruling can generate more correct results than majority voting mechanism with an average 15.8% performance loss. Ruosi Li, Xianglong Kong, Wei Guo 0018, Jingdong Guo, Hongfa Li, Fan Zhang 0044 |
Softw. Test. Verification Reliab. | 4 |